How to run an LLM locally (Ollama, llama.cpp, LM Studio) step by step

Install Ollama, run “ollama run” with a small model such as a 3B–8B instruct model, and call its local API on port 11434. For more control use llama.cpp with a quantized GGUF file; for a desktop GUI use LM Studio. Size the model to your RAM or VRAM first.
What do you need to run an LLM on your own machine?
You need three things: a runtime that executes the model, a model file in a format that runtime understands, and enough memory to hold the weights plus the context window. Everything else, from chat UIs to IDE plugins, sits on top of those three.
The memory rule is the one that decides everything. A model’s weights take roughly its parameter count multiplied by the bytes per parameter, so quantization (storing weights in 4 or 5 bits instead of 16) is what makes laptops viable. The context window adds a KV cache on top, which grows with the number of tokens you keep in play.
Apple Silicon Macs share memory between CPU and GPU, which makes them unusually good local inference machines. On Windows and Linux, an NVIDIA GPU with enough VRAM is the smooth path; CPU-only inference works but is noticeably slower for anything above small models.
Local inference also changes the economics of experimentation. There is no per-token bill, so you can run long batch jobs, retry prompts freely and test agents that make many calls without watching a meter.
Which local runtime should you choose?
The three popular options share the same engine lineage: Ollama and LM Studio both use llama.cpp (LM Studio also supports Apple’s MLX on Macs). The difference is how much they hide from you.
| Runtime | Interface | Licence | Best for | Trade-off |
|---|---|---|---|---|
| Ollama | CLI + local REST API | MIT | Developers who want a model server in one command | Less control over low-level flags; own model registry naming |
| llama.cpp | CLI binaries + llama-server | MIT | Maximum control, custom builds, embedded use | You manage builds, model files and flags yourself |
| LM Studio | Desktop GUI + local server | Proprietary app (free to use); check its terms for work use | Non-CLI users, browsing and testing models quickly | Closed source; not a fit for headless servers |
How to run a model with Ollama
Ollama is the shortest path from zero to a working local model, and it exposes an HTTP API that most AI tools already know how to talk to.
- Install it: on macOS and Windows download the installer from ollama.com; on Linux run curl -fsSL https://ollama.com/install.sh | sh
- Start a chat with a small model: ollama run llama3.2 (the model downloads on first use).
- List what you have downloaded: ollama list; remove a model with ollama rm <name>.
- Pull a model without starting a chat: ollama pull qwen2.5-coder (useful for scripts and CI images).
- Call the API from code: send a POST request to http://localhost:11434/api/chat, or point any OpenAI SDK at http://localhost:11434/v1 as the base URL.
- Customize behaviour with a Modelfile (system prompt, temperature, context length), then build it with ollama create my-model -f Modelfile.
How to run a model with llama.cpp
llama.cpp is the engine underneath much of the local ecosystem. Using it directly gives you every knob: GPU layer offload, context size, batch size, sampling settings and the exact quantization you want.
- Clone the repository: git clone https://github.com/ggml-org/llama.cpp and cd llama.cpp
- Build it: cmake -B build, then cmake --build build --config Release (add -DGGML_CUDA=ON for NVIDIA GPUs; Metal is enabled by default on Apple Silicon).
- Download a GGUF model file from Hugging Face; a Q4_K_M or Q5_K_M quantization is a sensible starting point.
- Chat in the terminal: ./build/bin/llama-cli -m path/to/model.gguf
- Serve an OpenAI-compatible API: ./build/bin/llama-server -m path/to/model.gguf --port 8080 -c 8192, then use http://localhost:8080/v1 as the base URL.
- Offload layers to the GPU with -ngl followed by a number; raise it until the model fits in VRAM.
How to run a model with LM Studio
- Download LM Studio for your OS from lmstudio.ai and open it.
- Use the model search to find an instruct model; the app shows which quantizations are likely to fit your hardware.
- Load the model and chat in the built-in window to sanity-check quality and speed.
- Open the developer or server tab and start the local server; it speaks an OpenAI-compatible API, by default on port 1234.
- Point your app or editor plugin at that base URL; keep the app running while you use it.
How big a model can your hardware handle?
Treat the following as rough planning ranges for 4-bit quantized models, not benchmarks. Leave headroom for the operating system, other apps and the context cache.
| Available RAM / VRAM | Realistic model size | Typical use |
|---|---|---|
| 8 GB | 1B–4B parameters | Autocomplete, classification, simple chat |
| 16 GB | 7B–8B, some 12B–14B | General chat, coding help, RAG answers |
| 24–32 GB | 14B–32B | Stronger reasoning and coding |
| 64 GB and up | 70B-class at low quantization | Near cloud-quality answers, slower tokens |
Common mistakes when running LLMs locally
- Picking the largest model that technically loads: if it spills into swap or CPU, tokens crawl and the model feels worse than a smaller one.
- Ignoring context length: long contexts cost memory; set it explicitly instead of assuming the maximum.
- Using base models for chat: pick instruct or chat variants, otherwise the model continues text instead of answering.
- Exposing the API to the network by accident: Ollama and llama-server should stay on localhost unless you add authentication in front.
- Comparing quality across quantizations blindly: very low-bit quants (2–3 bit) can degrade answers more than a smaller model at 4–5 bit.
- Forgetting the chat template: when running raw GGUF files, a wrong template produces rambling or broken output.
What can you build once a local model runs?
A local model becomes useful once other software talks to it. Because all three runtimes offer an OpenAI-style API, you can plug them into chat UIs, coding assistants, RAG pipelines and automation tools with a base URL change.
Good first projects are a private document Q&A tool, a coding assistant that never sends code off the machine, or batch jobs such as tagging and summarizing files overnight. RepoLoot’s catalog tags local-first projects by difficulty, which helps when choosing what to build on top of your runtime.
How do you connect a local model to your editor and apps?
The practical value of a local model shows up when your existing tools use it. Because Ollama, llama-server and LM Studio all speak an OpenAI-style API, most integrations only need a base URL, a model name and a dummy API key.
Test the connection with a plain HTTP request before debugging a plugin. If curl works and the plugin does not, the problem is plugin configuration, not the model.
Model files are large, and it is easy to fill a disk while experimenting. Ollama stores models in its own directory, while llama.cpp and LM Studio keep GGUF files wherever you download them.
Delete candidates you rejected, keep one or two models per task, and note which quantization you settled on. If you work on several machines, a shared script that pulls the exact models you use saves time and keeps everyone on the same versions.
Decide with your own tasks, not leaderboards. Write ten to twenty prompts that mirror real use, such as summarizing your documents or fixing your code, and run them against two or three candidate models.
Record speed as well as quality. A slightly weaker model that answers in two seconds often beats a stronger one that takes thirty, especially inside editor integrations and agent loops that make many calls.
Re-test when you change quantization, context length or runtime version, because each can shift results. Keep the prompt set in a file so the comparison is repeatable.
- Chat UIs: Open WebUI and similar projects can point at Ollama’s address and list your downloaded models automatically.
- Coding assistants: extensions such as Continue or Cline accept an Ollama or OpenAI-compatible provider; pick a code-tuned model for completions.
- Scripts: in the OpenAI Python or JavaScript SDK, set the base URL to the local /v1 endpoint and pass any non-empty string as the key.
- Docker: from inside a container, localhost is the container itself; use host.docker.internal or put both services on one Docker network.
- Remote access: if a teammate needs the model, place an authenticating reverse proxy in front instead of binding the raw API to 0.0.0.0.
Frequently asked questions
- Is Ollama or llama.cpp faster?
- They use the same core engine, so raw speed is usually similar for the same model and quantization. llama.cpp can be faster when you tune flags such as GPU layers, batch size and threads for your hardware, while Ollama picks reasonable defaults for you. Measure on your own machine with your real prompts.
- Can I run an LLM locally without a GPU?
- Yes. llama.cpp and Ollama run on CPU only, using system RAM. Small models of 1B to 4B parameters are comfortable on a modern laptop CPU, and 7B–8B models work at slower speeds. For interactive use of larger models, a GPU or an Apple Silicon Mac makes a large difference.
- What is a GGUF file?
- GGUF is the model file format used by llama.cpp and the tools built on it. One file bundles the quantized weights, tokenizer and metadata such as the chat template. You typically download a GGUF from Hugging Face in the quantization level that fits your memory, such as Q4_K_M.
- Are local LLMs private?
- Inference happens on your machine, so prompts and outputs are not sent to a model provider. Privacy still depends on the surrounding software: check whether your chat UI or plugin sends telemetry, keep the API bound to localhost, and review the model licence if you plan commercial use.