How to run an LLM locally: step by step with Ollama
Running a model on your own machine takes about ten minutes. You get privacy (nothing leaves your computer), no per-token bill and a model that works offline. Here is the shortest path from nothing to a working local LLM, plus how to call it from code.
1. Check your memory#
The model has to fit in memory. A quick estimate: parameters (in billions) × 0.6 gives roughly the gigabytes a 4-bit model needs. A 7B model is around 5 GB, a 14B around 9 GB. The full table and formula are in how much RAM do you need to run an LLM locally.
If you have 8 GB of RAM, choose a small model (about 3B). With 16 GB, 7B or 8B models are comfortable.
2. Install Ollama#
Download the installer for your OS from ollama.com, or use the package manager your platform supports. Then confirm it works:
ollama --versionOllama runs a small local server in the background and downloads models for you. If you prefer a desktop app with a model browser, use LM Studio instead; see Ollama vs LM Studio vs llama.cpp.
3. Pull a model#
Browse the model library on ollama.com, pick one whose size fits your memory, and pull it. Replace <model> with the name from the library page:
ollama pull <model>Model names change often, which is why this post does not hardcode one. Prefer the default 4-bit build; it is the usual sweet spot (what is quantization).
4. Chat in the terminal#
ollama run <model>Type a prompt and press Enter. /bye exits. Try something you can judge, like "explain this regex" or "write a TypeScript function that groups an array by key".
If it feels slow, the usual causes are a model too big for your memory (it spills to disk), a long context, or CPU-only inference. Try a smaller model or a lower-bit build.
5. Call it from code#
Ollama serves an HTTP API on localhost:11434.
const res = await fetch("http://localhost:11434/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: "<model>",
stream: false,
messages: [{ role: "user", content: "Give me three names for a CLI tool that cleans git branches." }],
}),
});
const data = await res.json();
console.log(data.message.content);No API key, no billing. Many tools also expose an OpenAI-compatible endpoint, so existing clients often work by changing the base URL. Check which features your chosen model supports, such as tool calling or structured output.
When a local model is the right choice#
| Good fit | Poor fit |
|---|---|
| Private or regulated data | The hardest reasoning and coding tasks |
| Offline or air-gapped work | Many users at once on one laptop |
| Classification, extraction, summaries | Very long documents on small memory |
| Cheap experiments and evals | When latency from a laptop is not acceptable |
A hybrid is common: a local model for the bulk, high-volume work and a hosted model for the hard cases. Whichever you pick, measure it on your own inputs with a small eval.
Troubleshooting#
- Out of memory or crawling: pick a smaller model or a more aggressive quantization.
- Command not found: reopen the terminal after installing.
- Port in use: another process is using 11434; stop it or change Ollama's address.
- Answers are poor: try a larger model, give clearer prompts, or check you are not on a heavily compressed build.
Frequently asked questions
How do I run an LLM locally?
Install a local runner such as Ollama, LM Studio or llama.cpp, download a model that fits your memory, and run it. With Ollama that is ollama pull
Do I need a GPU to run an LLM locally?
No. Quantized models run on the CPU using system RAM, just more slowly. A GPU or Apple Silicon makes responses much faster.
How much RAM do I need to run a local LLM?
Roughly the model's parameters times its bytes per weight, plus 10 to 20 percent. A 7B model at 4-bit needs about 5 GB. See the memory guide for a full table.
Is running an LLM locally private?
Prompts and responses stay on your machine when you run a downloaded model. Downloading models and updates still uses the network.
Prefer plain text? Read this page as Markdown.