Local LLMs

How much RAM do you need to run an LLM locally?

"Will this model run on my machine?" has a quick answer you can calculate in your head. The two numbers you need: how many parameters, and how many bits each one is stored in.

The formula#

text
memory (GB) ≈ parameters (billions) × bits per weight ÷ 8   +   10-20% overhead

A 7-billion-parameter model stored at 16 bits is 7 × 16 ÷ 8 = 14 GB before overhead. The same model quantized to roughly 4.8 bits per weight (a common 4-bit format) is about 4.2 GB. That drop is why people can run these models at home. See what is quantization.

Quick reference#

These are estimates for the weights. Add overhead for context and the runtime.

7B, 16-bit14 GB7B, 8-bit7.5 GB7B, 4-bit4.2 GB14B, 4-bit8.4 GB32B, 4-bit19 GB70B, 4-bit42 GB
Approximate weight memory. Real files and overhead vary by format and runtime.

Context needs memory too#

The model also keeps a KV cache for every token in the conversation. It grows with context length, so a long prompt or a big document can add several gigabytes on top of the weights. If you plan to feed large documents, leave headroom, or set a smaller context window in your tool.

Which memory matters#

Discrete GPUFastest when the whole model fits inVRAMVRAM is the limit (8, 12, 24 GB cards)Offload some layers to RAM if it doesnot fit, with a speed hitCPU or Apple SiliconRuns from system RAM, no GPU neededSlower per token than a good GPUApple unified memory lets big modelsrun on one machinevs

A practical sizing guide#

You haveComfortable choices
8 GB RAMSmall models (about 3B at 4-bit), short context
16 GB RAM7B to 8B at 4 or 8-bit
32 GB RAM14B at 8-bit, or up to about 30B at 4-bit
64 GB or more70B at 4-bit

Leave memory for your OS and apps. If the model barely fits, it will swap and crawl.

Before you download#

  1. Check the parameter count and quantization in the model name or file name.
  2. Compute the estimate. Add 20 percent.
  3. Compare it to free RAM or VRAM, not total.
  4. If it does not fit, pick a smaller model or a lower-bit version rather than a bigger machine.

Not sure which tool to use? Read Ollama vs LM Studio vs llama.cpp.

Frequently asked questions

How much RAM do I need to run an LLM locally?

Multiply the parameter count by the bytes per weight, then add 10 to 20 percent for context and overhead. At 4-bit quantization that is about 0.5 to 0.6 bytes per parameter, so a 7B model needs roughly 5 GB and a 70B model roughly 42 GB.

Can I run an LLM without a GPU?

Yes. Quantized models run on the CPU using system RAM. It is slower than on a GPU, but small models are usable on a laptop.

Does context length use memory?

Yes. The KV cache grows with the number of tokens in context, so long prompts need extra memory on top of the weights.

What is the difference between RAM and VRAM?

VRAM is the memory on your graphics card, and speed is best when the whole model fits in it. System RAM is slower but larger. Apple Silicon uses unified memory shared by CPU and GPU.