How much RAM do you need to run an LLM locally?
"Will this model run on my machine?" has a quick answer you can calculate in your head. The two numbers you need: how many parameters, and how many bits each one is stored in.
The formula#
memory (GB) ≈ parameters (billions) × bits per weight ÷ 8 + 10-20% overheadA 7-billion-parameter model stored at 16 bits is 7 × 16 ÷ 8 = 14 GB before overhead. The same model quantized to roughly 4.8 bits per weight (a common 4-bit format) is about 4.2 GB. That drop is why people can run these models at home. See what is quantization.
Quick reference#
These are estimates for the weights. Add overhead for context and the runtime.
Context needs memory too#
The model also keeps a KV cache for every token in the conversation. It grows with context length, so a long prompt or a big document can add several gigabytes on top of the weights. If you plan to feed large documents, leave headroom, or set a smaller context window in your tool.
Which memory matters#
A practical sizing guide#
| You have | Comfortable choices |
|---|---|
| 8 GB RAM | Small models (about 3B at 4-bit), short context |
| 16 GB RAM | 7B to 8B at 4 or 8-bit |
| 32 GB RAM | 14B at 8-bit, or up to about 30B at 4-bit |
| 64 GB or more | 70B at 4-bit |
Leave memory for your OS and apps. If the model barely fits, it will swap and crawl.
Before you download#
- Check the parameter count and quantization in the model name or file name.
- Compute the estimate. Add 20 percent.
- Compare it to free RAM or VRAM, not total.
- If it does not fit, pick a smaller model or a lower-bit version rather than a bigger machine.
Not sure which tool to use? Read Ollama vs LM Studio vs llama.cpp.
Frequently asked questions
How much RAM do I need to run an LLM locally?
Multiply the parameter count by the bytes per weight, then add 10 to 20 percent for context and overhead. At 4-bit quantization that is about 0.5 to 0.6 bytes per parameter, so a 7B model needs roughly 5 GB and a 70B model roughly 42 GB.
Can I run an LLM without a GPU?
Yes. Quantized models run on the CPU using system RAM. It is slower than on a GPU, but small models are usable on a laptop.
Does context length use memory?
Yes. The KV cache grows with the number of tokens in context, so long prompts need extra memory on top of the weights.
What is the difference between RAM and VRAM?
VRAM is the memory on your graphics card, and speed is best when the whole model fits in it. System RAM is slower but larger. Apple Silicon uses unified memory shared by CPU and GPU.
Prefer plain text? Read this page as Markdown.