2026 comparison: the best open-source LLMs to self-host
The performance gap between proprietary and open-source models keeps narrowing, especially for structured agent tasks (extraction, classification, summarization).
Selection criteria
- Model size vs. available hardware: a 7B-14B model runs comfortably on most entry-level GPU servers, or even on CPU for lighter use cases.
- Tool support (function calling): essential for agentic use cases.
- License: some licenses restrict commercial use, worth checking before any production deployment.
Our picks for self-hosting via Ollama
The Llama and Qwen model families remain solid choices for local deployment, offering a good quality-to-inference-cost ratio on reasonable hardware.
We'll keep updating this comparison as new models ship.