Ollama vs LM Studio vs llama.cpp: which should you use?
Three names come up whenever someone asks "how do I run an LLM on my own machine?" They are not exact rivals: one is an engine, the other two are friendly layers around it. Here is how to choose.
How they relate#
At a glance#
| Ollama | LM Studio | llama.cpp | |
|---|---|---|---|
| What it is | CLI + local API | Desktop app + local server | Inference engine and server |
| Best for | Developers, scripts, apps | Exploring models, chat, non-CLI users | Control, edge devices, custom builds |
| Setup | Installer, then one command | Installer, point and click | Build or download binaries |
| Model discovery | ollama pull <model> | Built-in model browser | Download GGUF files yourself |
| API | REST on localhost:11434 | Local server, OpenAI-compatible | llama-server, OpenAI-compatible |
| Tuning knobs | Modelfile, a few options | GUI sliders | Everything, via flags |
| Learning curve | Low | Lowest | Highest |
Ollama#
You install it, pull a model and run it:
ollama pull <model>
ollama run <model>It also runs a local server. Hit it from code:
const res = await fetch("http://localhost:11434/api/chat", {
method: "POST",
body: JSON.stringify({
model: "<model>",
stream: false,
messages: [{ role: "user", content: "Explain closures in one paragraph." }],
}),
});
const data = await res.json();
console.log(data.message.content);Pros: simplest developer workflow, scriptable, large library of models. Cons: fewer low-level knobs than raw llama.cpp.
LM Studio#
A desktop app. You search for a model, download it, and chat. It can also start a local server so your editor or app can use it. It is the easiest way to try models and compare them without touching a terminal. Cons: GUI-first, less natural for automation and servers.
llama.cpp#
The engine under the others. You choose the model file, the quantization, the number of layers on GPU, the context size, the threads. It runs on almost anything: NVIDIA, AMD, Apple Silicon, plain CPUs.
llama-server -m ./model-Q4_K_M.gguf -c 8192Pros: maximum control and portability, great for edge devices. Cons: you manage files and flags yourself.
Which one should you pick?#
- "I want to build something with a local model." Ollama.
- "I want to try many models and chat with them." LM Studio.
- "I need to run on a Raspberry Pi, tune every flag, or embed it." llama.cpp.
- "I need many users hitting one model at high throughput." Look at a serving engine such as vLLM, which is a different class of tool.
You can use more than one. They all read the same kind of GGUF file.
Before you start#
New to this? Follow how to run an LLM locally first.
- Estimate memory first: how much RAM do you need.
- Pick a quantization, usually 4-bit.
- Test the model on your own cases with a small eval. A local model that is fine for classification may not be fine for complex coding.
Frequently asked questions
How do I run an LLM locally?
Check that the model fits in your memory, install a runner such as Ollama, pull a model and run it. The step-by-step guide is in the post How to run an LLM locally.
What is the difference between Ollama, LM Studio and llama.cpp?
llama.cpp is the inference engine. Ollama wraps a model runner with a simple CLI and local API. LM Studio is a desktop app with a model browser, chat UI and a local server.
Which is best for developers?
Ollama is the usual pick for developers because it is scriptable, has a REST API and is easy to use in code. LM Studio is better if you prefer a GUI. llama.cpp is for when you need full control.
Can I use a local model with code written for a cloud API?
Often yes. Ollama, LM Studio and llama.cpp's server expose OpenAI-compatible endpoints, so many clients work by changing the base URL. Check which features, such as tool calling, your chosen model supports.
Do these tools send my data to the cloud?
Running a downloaded model through these tools happens on your machine, which is the main privacy advantage of local LLMs. Downloading models and checking for updates does use the network.
Prefer plain text? Read this page as Markdown.