Local LLMs

Ollama vs LM Studio vs llama.cpp: which should you use?

Three names come up whenever someone asks "how do I run an LLM on my own machine?" They are not exact rivals: one is an engine, the other two are friendly layers around it. Here is how to choose.

How they relate#

Your app or editorcalls a local HTTP APIOllama | LM Studiomodel downloads, runner, CLI or GUI, local serverllama.cppthe inference engine that runs GGUF modelsYour hardwareCPU, NVIDIA / AMD GPU, Apple Silicon
Ollama and LM Studio make the engine easy to use. llama.cpp is the engine itself.

At a glance#

OllamaLM Studiollama.cpp
What it isCLI + local APIDesktop app + local serverInference engine and server
Best forDevelopers, scripts, appsExploring models, chat, non-CLI usersControl, edge devices, custom builds
SetupInstaller, then one commandInstaller, point and clickBuild or download binaries
Model discoveryollama pull <model>Built-in model browserDownload GGUF files yourself
APIREST on localhost:11434Local server, OpenAI-compatiblellama-server, OpenAI-compatible
Tuning knobsModelfile, a few optionsGUI slidersEverything, via flags
Learning curveLowLowestHighest

Ollama#

You install it, pull a model and run it:

bash
ollama pull <model>
ollama run <model>

It also runs a local server. Hit it from code:

ts
const res = await fetch("http://localhost:11434/api/chat", {
  method: "POST",
  body: JSON.stringify({
    model: "<model>",
    stream: false,
    messages: [{ role: "user", content: "Explain closures in one paragraph." }],
  }),
});
const data = await res.json();
console.log(data.message.content);

Pros: simplest developer workflow, scriptable, large library of models. Cons: fewer low-level knobs than raw llama.cpp.

LM Studio#

A desktop app. You search for a model, download it, and chat. It can also start a local server so your editor or app can use it. It is the easiest way to try models and compare them without touching a terminal. Cons: GUI-first, less natural for automation and servers.

llama.cpp#

The engine under the others. You choose the model file, the quantization, the number of layers on GPU, the context size, the threads. It runs on almost anything: NVIDIA, AMD, Apple Silicon, plain CPUs.

bash
llama-server -m ./model-Q4_K_M.gguf -c 8192

Pros: maximum control and portability, great for edge devices. Cons: you manage files and flags yourself.

Which one should you pick?#

  • "I want to build something with a local model." Ollama.
  • "I want to try many models and chat with them." LM Studio.
  • "I need to run on a Raspberry Pi, tune every flag, or embed it." llama.cpp.
  • "I need many users hitting one model at high throughput." Look at a serving engine such as vLLM, which is a different class of tool.

You can use more than one. They all read the same kind of GGUF file.

Before you start#

New to this? Follow how to run an LLM locally first.

  1. Estimate memory first: how much RAM do you need.
  2. Pick a quantization, usually 4-bit.
  3. Test the model on your own cases with a small eval. A local model that is fine for classification may not be fine for complex coding.

Frequently asked questions

How do I run an LLM locally?

Check that the model fits in your memory, install a runner such as Ollama, pull a model and run it. The step-by-step guide is in the post How to run an LLM locally.

What is the difference between Ollama, LM Studio and llama.cpp?

llama.cpp is the inference engine. Ollama wraps a model runner with a simple CLI and local API. LM Studio is a desktop app with a model browser, chat UI and a local server.

Which is best for developers?

Ollama is the usual pick for developers because it is scriptable, has a REST API and is easy to use in code. LM Studio is better if you prefer a GUI. llama.cpp is for when you need full control.

Can I use a local model with code written for a cloud API?

Often yes. Ollama, LM Studio and llama.cpp's server expose OpenAI-compatible endpoints, so many clients work by changing the base URL. Check which features, such as tool calling, your chosen model supports.

Do these tools send my data to the cloud?

Running a downloaded model through these tools happens on your machine, which is the main privacy advantage of local LLMs. Downloading models and checking for updates does use the network.