Open source ChatGPT alternatives: self-hosted chat UIs plus open models

The practical open source ChatGPT alternative is two layers: a self-hosted chat interface such as Open WebUI, LibreChat or LobeChat, and a model runtime such as Ollama, llama.cpp or vLLM serving open-weight models like Llama, Mistral, Qwen or DeepSeek. Desktop apps like Jan bundle both.
What is an open source ChatGPT alternative, really?
ChatGPT is a product with three parts: a chat interface, a hosted model, and a pile of extras such as file uploads, memory and web search. An open source replacement rarely comes as one package. You usually assemble it from a front end, a model server and the models themselves.
That split is good news. You can swap the model without retraining your team on a new interface, and you can point the same interface at a local model for private data and at a commercial API for hard tasks.
The question to answer first is not “which project is best” but “where must the data live and how much hardware do I have”. Those two constraints decide most of the stack.
Which self-hosted chat UIs are worth considering?
The chat front end is where users spend their time, so it matters more than people expect. The leading projects all support multiple conversations, markdown rendering, model switching and some form of document upload. They differ in how they handle users, plugins and providers.
| Project | Type | Model sources | Best for | Trade-off |
|---|---|---|---|---|
| Open WebUI | Web app (Python + Svelte) | Ollama and OpenAI-compatible APIs | Teams running Ollama on a shared server | Licence terms changed over time; check the licence file before rebranding |
| LibreChat | Web app (Node.js) | Many commercial APIs plus OpenAI-compatible endpoints | Multi-provider access behind one login | More configuration files to manage |
| LobeChat | Web app (Next.js) | Wide provider list, plugin system | Polished UI and quick personal deployments | Check the licence file for commercial use |
| Jan | Desktop app | Bundled local engine plus remote APIs | Single users who want offline chat | Not a multi-user server |
| text-generation-webui | Web app (Python) | Local model loaders | Tinkering with model formats and settings | Built for experimenters, not end users |
Which runtime should serve the models?
The runtime loads model weights and exposes an API the chat UI can call. Most of them now speak an OpenAI-compatible API, which is what makes the layers interchangeable.
Ollama is the easiest starting point: one binary, a model library and a simple pull command. llama.cpp is the engine underneath many desktop tools and runs quantized GGUF models on CPUs and consumer GPUs. vLLM targets GPU servers and focuses on throughput when many users hit the same model at once.
- One person, one laptop: Jan or Ollama with a small quantized model.
- A small team on one GPU box: Ollama or llama.cpp server behind Open WebUI.
- Many concurrent users: vLLM or a similar GPU-first server behind LibreChat or Open WebUI.
- Mixed workloads: keep a commercial API configured in the same UI for tasks local models handle poorly.
Which open models should you pair with it?
Open-weight model families move fast, so pick by licence and size class rather than by leaderboard position. Meta’s Llama models ship under Meta’s own community licence, which has conditions of its own. Several Mistral and Qwen releases have been published under Apache 2.0, but each release has its own terms, so read the model card for the exact version you download.
Size matters more than brand for a first deployment. Small models in the single-digit billions of parameters run on laptops and handle drafting, summarizing and classification. Larger models need serious GPU memory but cope better with reasoning, long documents and code.
Keep at least two models configured: a fast one for everyday chat and a stronger one for tasks where quality matters. Users will learn when to switch.
How do you set one up step by step?
Most of these projects publish a Docker Compose file, which is the fastest honest route to a working install. Resist the urge to enable every plugin on day one.
- Decide the data boundary: fully local, private cloud, or local UI with external APIs.
- Install a runtime such as Ollama on the machine with the GPU and pull one small model to test.
- Deploy the chat UI with Docker and point it at the runtime’s API address.
- Turn on authentication and create accounts before anyone else can reach the URL.
- Add document upload or retrieval only after plain chat works reliably.
- Put the service behind a reverse proxy with TLS and back up the UI’s database.
Which features matter beyond plain chat?
Once basic chat works, users ask for the extras they know from ChatGPT. Each one adds a moving part, so decide deliberately which you actually need.
Document chat relies on retrieval: files are split into chunks, embedded with an embedding model and stored in a vector index. The chat UIs ship a built-in version, which is fine for a few manuals but weak for a large archive with strict permissions.
Web search needs a search backend. Several UIs can call a self-hosted SearXNG instance or a commercial search API, then feed results to the model. Quality depends heavily on how results are cleaned before they reach the prompt.
- Start with accounts and model switching only.
- Add document chat for one team with a clear document set.
- Add web search and tools once people trust the basic answers.
| Feature | What it needs | Self-hosting effort | Watch out for |
|---|---|---|---|
| Multi-user accounts | Built-in auth or SSO via OIDC | Low | Default admin accounts left open |
| Document chat | Embedding model plus vector store | Medium | Permissions: every user may see every file |
| Web search | SearXNG or a search API | Medium | Slow answers when many pages are fetched |
| Tool calling and agents | A model trained for function calling | Medium to high | Small models call tools unreliably |
| Voice input and output | Speech-to-text and text-to-speech models | Medium | Extra GPU or CPU load |
| Image generation | A separate image model server | High | Hardware needs rise sharply |
How much does it cost to run?
The software is free to download, but the hardware and your time are not. A single user can run small models on an existing laptop at no extra cost. Shared use needs a machine with a capable GPU, either owned or rented, and GPU rental is billed by the hour whether anyone is chatting or not.
The cheaper pattern for many teams is hybrid: a self-hosted UI that keeps conversation history on your server, calling a pay-per-token API for most requests and a local model for sensitive ones. Measure your own usage for a month before buying hardware.
Where does it break? Common mistakes
Most failed pilots fail on expectations, not software. Tell users which model they are talking to and what it is good at.
- Exposing the UI to the internet without authentication, which invites strangers to use your GPU.
- Expecting a small local model to match a frontier hosted model on hard reasoning tasks.
- Ignoring context length, then wondering why long documents get cut off or forgotten.
- Treating document upload as a full search system; retrieval quality depends on chunking and embeddings.
- Skipping backups of the UI database, which holds every conversation and user account.
- Mixing model licences without reading them, especially when the output feeds a commercial product.
How to choose the right combination
Start from the user count. One user points to a desktop app; a handful points to Open WebUI or LibreChat on one server; dozens point to a GPU server with a throughput-oriented runtime.
Then check the licence of each layer against what you plan to do. Internal use is rarely a problem. Rebranding the UI or reselling access is where licence terms start to matter. RepoLoot’s catalog tags projects by licence and difficulty, which helps when you are shortlisting components for a client build.
Finally, budget for maintenance. These projects release often, and staying a few versions behind is fine, but staying a year behind usually means a painful migration.
Once the UI and runtime are running, the same stack becomes a platform. Internal assistants for support teams, drafting tools for sales and private document helpers for legal or finance are all the same chat server with different system prompts, models and document sets.
Because the runtime speaks an OpenAI-compatible API, your own scripts and apps can call it too. A small internal service can reuse the model server for classification or summarization without adding a new vendor.
Frequently asked questions
- Is there a single open source app that fully replaces ChatGPT?
- Not as one package. Desktop apps like Jan come closest for a single user because they bundle an interface and a local engine. For teams you normally combine a web chat UI with a separate model runtime and choose open-weight models to load into it.
- Can I run an open source ChatGPT alternative without a GPU?
- Yes, for small quantized models. llama.cpp and tools built on it run on CPUs and Apple silicon, and response speed is acceptable for one user. Larger models and several simultaneous users realistically need a GPU with enough memory to hold the weights.
- Are open-weight models free for commercial use?
- It depends on the model. Some releases use permissive licences such as Apache 2.0, while others use custom licences with conditions. Always read the licence or model card for the exact version you deploy rather than assuming the whole model family shares the same terms.
- Can a self-hosted chat UI still use OpenAI or Anthropic models?
- Yes. LibreChat, Open WebUI and LobeChat can call commercial APIs alongside local models. This gives you one interface, your own conversation storage and user accounts, while letting you route difficult or high-volume tasks to whichever provider suits them.