LiteLLM: one gateway to talk to every LLM
Using several LLMs (Anthropic, OpenAI, a local Ollama model) in the same stack quickly multiplies glue code: each provider has its own SDK, its own response format, its own way of handling keys.
What LiteLLM does
LiteLLM is an open-source proxy that exposes a single OpenAI-compatible API in front of any provider. Your application code always calls the same endpoint; LiteLLM routes to the right model behind the scenes.
Concrete benefits for an agent stack:
- A single switch point between models: changing provider doesn't touch agent code, only the proxy config.
- Built-in response caching, useful when several agents ask similar questions.
- Centralized cost tracking, per API key or per agent, instead of piecing together the bill from several provider consoles.
Self-hosting vs SaaS
LiteLLM also ships as a managed cloud version, but the proxy self-hosts very easily as a single Docker container. For a stack already running on a VPS, self-hosting avoids exposing LLM traffic to yet another third-party service and keeps API keys on your own infrastructure.
The one caveat
LiteLLM adds an extra network hop between your agents and the final provider. For very latency-sensitive tasks, that overhead (usually a few milliseconds if the proxy runs on the same machine or private network) stays negligible compared to the observability gains.