Ollama in production: what self-hosting really costs vs API
"Self-hosting is free" is a common misconception. A VPS with enough RAM/CPU to run a 7B-14B model properly has a fixed monthly cost, whether it's used 10 times or 10,000 times a day.
The two cost structures
Proprietary API: variable cost, proportional to tokens consumed. No fixed cost, but the bill grows linearly with usage.
Self-hosted Ollama: fixed cost (the VPS), close to zero beyond that. Profitable once usage volume crosses a certain threshold.
Where the break-even point sits
For a 7B-14B class model, a VPS with 16GB of RAM is largely enough and typically costs between $20 and $50/month depending on the host. That monthly budget roughly equals, in proprietary API terms, a few million tokens depending on the provider and model — the exact math depends heavily on current pricing, worth checking before deciding.
Below that volume, the proprietary API stays cheaper and avoids server maintenance. Above it, self-hosting becomes profitable and stays that way regardless of volume — the main argument for an agent stack running continuously (automated watch, bulk document processing).
The factor often overlooked: model quality
A self-hosted 7B-14B open-source model doesn't always match a top-tier proprietary model on complex reasoning tasks. Many teams adopt a hybrid approach: Ollama for high-volume, repetitive tasks (classification, extraction), proprietary API for tasks that need the best available reasoning.
To pick the VPS that will host your Ollama instance, see our article on hosting AI agents on a VPS.