The Best Open Source LLMs to Self-Host in 2026

7 minUpdated:
The Best Open Source LLMs to Self-Host in 2026

For most teams the strongest self-hosted picks are the Qwen, Llama, Mistral, Gemma, DeepSeek and gpt-oss families. Choose by licence first, then by the largest size your GPU memory holds at 4-bit or 8-bit quantization, then test on your own prompts.

What does “open source LLM” actually mean?

Most models people call open source are really open-weight: you can download the trained weights and run them, but the training data and full recipe are not published. That is enough to self-host, fine-tune and ship products, which is what most builders care about.

The part that differs wildly is the licence. Some families use Apache 2.0 or MIT, some use a custom community licence with usage rules, and some add acceptable-use policies. Treat the licence as a hard filter before you compare quality.

A small number of projects, such as AI2’s OLMo, publish weights, data and training code together. Those are the ones closest to open source in the strict sense, and they matter if you need auditability.

Which open-weight model families are worth shortlisting?

The table below covers the families that show up in almost every serious self-hosting shortlist. Licences change between releases, so read the licence file of the exact checkpoint you download.

FamilyPublisherLicence familyBest forTrade-off
QwenAlibabaMostly Apache 2.0 (check per model)General chat, coding, multilingual, many sizesSome checkpoints carry different terms
LlamaMetaCustom community licenceHuge ecosystem, fine-tunes, tooling supportLicence has usage conditions and naming rules
Mistral / MixtralMistral AIApache 2.0 for many open releasesEfficient mid-size models, European vendorNot every Mistral model is open-weight
GemmaGoogleGemma terms of useSmall, capable models for single GPUsCustom terms, read before commercial use
DeepSeekDeepSeekCheck the licence file per releaseReasoning and coding at large scaleLargest versions need multi-GPU servers
gpt-ossOpenAIApache 2.0Reasoning with tool use, permissive termsFewer size options than Qwen or Llama
PhiMicrosoftCheck the licence file (often MIT)Small models for edge and laptopsLess world knowledge than bigger models
OLMoAI2Apache 2.0Research, full transparency of data and codeUsually behind the leaders on raw quality

How much hardware do you need to self-host an LLM?

Memory is the constraint, not compute. A rough rule: the weights take parameter count multiplied by bytes per parameter. At 16-bit that is two bytes per parameter, at 8-bit one byte, and at 4-bit about half a byte.

So a 7–8B model at 4-bit needs roughly 4–5 GB for weights, a 32B model around 16–20 GB, and a 70B model around 35–40 GB. Add headroom for the KV cache, which grows with context length and the number of parallel requests.

Mixture-of-experts models such as Mixtral, DeepSeek and some Qwen releases only activate part of the network per token. They run faster than their total size suggests, but you still need memory for all the weights.

  • Laptop or 8–12 GB GPU: 3–8B models at 4-bit, fine for autocomplete, extraction and simple chat.
  • 24 GB consumer GPU: up to roughly 30B models at 4-bit, the sweet spot for serious single-user work.
  • 48–80 GB workstation or single datacenter GPU: 70B-class models, or smaller models serving many users.
  • Multi-GPU node: the largest MoE releases, usually only worth it when you have steady traffic.
  • Apple Silicon with unified memory: a practical way to run large quantized models slowly but cheaply.

Which model size should you run for each job?

Bigger is not automatically better once you pay for the hardware yourself. Many production tasks, such as classifying tickets, extracting fields from emails or rewriting text into a template, work well with models in the 3–14B range when prompts are tight and outputs are structured.

Mid-size models around 30B are where general chat, coding help and multi-step reasoning start to feel close to hosted assistants for everyday work. They are also the largest size most teams can serve on one consumer GPU.

The 70B class and large mixture-of-experts releases are for difficult reasoning, long documents and agent workflows where a wrong step is expensive. Reserve them for the requests that need them rather than routing everything there.

TaskTypical size that worksWhy
Classification, routing, tagging1–8BShort outputs, easy to validate, latency matters
Structured extraction to JSON7–14BNeeds reliable format following, little world knowledge
Internal chat and summarization14–32BBetter nuance and fewer hallucinated details
Coding assistanceCoding-tuned 14–32B or largerCode-specific training beats raw size at the same budget
Agents and complex reasoning70B class or large MoEFewer compounding errors across many steps

Which runtime should you serve the model with?

The model and the serving engine are separate choices. For a single developer, Ollama or LM Studio wrap llama.cpp and get you running in minutes with quantized GGUF files.

For production with concurrent users, vLLM and Hugging Face TGI use continuous batching and paged attention to serve many requests per GPU. SGLang is another option focused on structured generation and throughput.

Most of these expose an OpenAI-compatible HTTP API, so your application code barely changes when you swap models or engines.

How do fine-tuned variants and distilled models fit in?

Every major family spawns community fine-tunes on Hugging Face: instruction-tuned, coding-tuned, uncensored, long-context and domain versions. They can be excellent, but provenance varies, and the base model’s licence still applies to the derivative.

Distilled reasoning models, where a small model is trained on outputs of a larger one, give surprising reasoning ability at low sizes. They tend to write long chains of thought, which costs tokens and latency, so measure end-to-end response time and not only answer quality.

If you plan your own fine-tune with LoRA, start from an official instruct checkpoint of a well-supported family. Tooling, chat templates and quantized builds will be available immediately, which saves days of debugging.

How to choose a model for your use case

  • Write down the licence constraints first: commercial use, redistribution, user-count clauses, attribution.
  • Measure your hardware budget in GPU memory and pick the largest size that fits at 4-bit or 8-bit with context headroom.
  • Shortlist two or three families that publish that size, including one coding-tuned variant if you generate code.
  • Build a test set of 30–100 real prompts from your product and grade outputs blind, not by leaderboard position.
  • Check context length, tool-calling format and chat template support in your chosen runtime.
  • Load-test with realistic concurrency before committing; latency under load often decides between two similar models.

Where self-hosted LLMs break

The common failure is choosing a model from a leaderboard and discovering it is weak at your language, domain or output format. Public benchmarks are a starting filter, not a decision.

The second is underestimating the KV cache. A model that fits comfortably with a short prompt can run out of memory once you allow long contexts and several users at once.

  • Mixing chat templates: a wrong template silently degrades quality and breaks tool calls.
  • Over-quantizing: 2–3 bit versions of small models lose far more quality than 4-bit versions of larger ones.
  • Ignoring licence updates between model versions of the same family.
  • No fallback: plan a hosted API route for peaks or outages if uptime matters.
  • Skipping evaluation after upgrades, so a new checkpoint quietly changes behavior in production.

When is self-hosting worth it versus an API?

Self-hosting pays off when data cannot leave your infrastructure, when you have steady high volume, or when you need to fine-tune and control the exact model version. It also removes per-token pricing surprises.

It is usually not worth it for low, spiky traffic or when you need the very best frontier quality. In those cases a hosted API is cheaper once you count GPU idle time and operations work.

A hybrid is common: a small self-hosted model handles classification, extraction and routing, while a hosted frontier model handles the hardest requests. RepoLoot’s catalog tags self-hostable projects by licence and difficulty, which helps when you assemble that stack.

Frequently asked questions

What is the best open source LLM to self-host right now?
There is no single winner. Qwen, Llama, Mistral, Gemma, DeepSeek and gpt-oss all have strong releases, and the right one depends on licence terms, the size your GPU memory allows and how the model performs on your own prompts. Test two or three candidates on real tasks before deciding.
Can I run a useful LLM without a GPU?
Yes, small models of roughly 1–8B parameters run on CPU with llama.cpp or Ollama using 4-bit quantization. Expect slow generation compared with a GPU, but it is fine for batch jobs, extraction, prototyping and private personal use. Apple Silicon machines with unified memory are a strong middle ground.
Are open-weight models safe for commercial use?
Many are, but not all under the same terms. Apache 2.0 and MIT models are the simplest. Community licences such as Llama’s and Gemma’s terms allow commercial use with conditions. Always read the licence file of the exact checkpoint you ship, since terms can differ between versions.
Should I quantize a model before self-hosting?
Usually yes. 4-bit or 8-bit quantization cuts memory needs by two to four times with a modest quality loss, which often lets you run a larger, better model on the same hardware. Avoid very aggressive 2–3 bit quantization on small models, where the quality drop becomes obvious.
Free for builders

Get a hand-picked shortlist of repos for your project

Tell us what you are building. A person — not a bot — reviews it and replies within 48 hours with the catalog projects that fit, including licence and difficulty notes.

We use your email only for this request. Privacy policy

Related guides