Tools & Tutorials

LiteLLM: one gateway to talk to every LLM

Using several LLMs (Anthropic, OpenAI, a local Ollama model) in the same stack quickly multiplies glue code: each provider has its own SDK, its own response format, its own way of handling keys.

What LiteLLM does

LiteLLM is an open-source proxy that exposes a single OpenAI-compatible API in front of any provider. Your application code always calls the same endpoint; LiteLLM routes to the right model behind the scenes.

Concrete benefits for an agent stack:

  • A single switch point between models: changing provider doesn't touch agent code, only the proxy config.
  • Built-in response caching, useful when several agents ask similar questions.
  • Centralized cost tracking, per API key or per agent, instead of piecing together the bill from several provider consoles.

Self-hosting vs SaaS

LiteLLM also ships as a managed cloud version, but the proxy self-hosts very easily as a single Docker container. For a stack already running on a VPS, self-hosting avoids exposing LLM traffic to yet another third-party service and keeps API keys on your own infrastructure.

The one caveat

LiteLLM adds an extra network hop between your agents and the final provider. For very latency-sensitive tasks, that overhead (usually a few milliseconds if the proxy runs on the same machine or private network) stays negligible compared to the observability gains.