1 server tagged with "llm-routing"
Check whether a task can run on a local model instead of cloud before every inference call. Saves money on every call that does not need cloud.