# How much RAM do you need to run an LLM locally?

> Estimate the memory a local LLM needs: parameters times bits per weight, plus context and overhead. Includes a size table for 7B to 70B models and GPU vs CPU notes.

Source: https://devaiper.com/blog/how-much-ram-to-run-an-llm-locally
Published: 2026-10-08
Topics: Local LLMs, Hardware, Quantization

**Short answer:** Memory is roughly parameters times bytes per weight, plus 10 to 20 percent for context and overhead. A 7B model at 4-bit needs about 5 GB, a 14B about 9 GB and a 70B about 42 GB. Quantization is what makes these fit.

"Will this model run on my machine?" has a quick answer you can calculate in your head. The two numbers you need: how many parameters, and how many bits each one is stored in.

## The formula

```text
memory (GB) ≈ parameters (billions) × bits per weight ÷ 8   +   10-20% overhead
```

A 7-billion-parameter model stored at 16 bits is 7 × 16 ÷ 8 = **14 GB** before overhead. The same model quantized to roughly 4.8 bits per weight (a common 4-bit format) is about **4.2 GB**. That drop is why people can run these models at home. See [what is quantization](https://devaiper.com/blog/what-is-llm-quantization).

## Quick reference

These are estimates for the weights. Add overhead for context and the runtime.

> **Diagram:** Approximate weight memory in gigabytes. A 7B model needs 14 GB at 16 bit, about 7.5 GB at 8 bit and about 4.2 GB at 4 bit. A 14B model at 4 bit needs about 8.4 GB. A 32B model at 4 bit needs about 19 GB. A 70B model at 4 bit needs about 42 GB.
> 7B, 16-bit: 14; 7B, 8-bit: 7.5; 7B, 4-bit: 4.2; 14B, 4-bit: 8.4; 32B, 4-bit: 19; 70B, 4-bit: 42
> Approximate weight memory. Real files and overhead vary by format and runtime.

## Context needs memory too

The model also keeps a **KV cache** for every token in the conversation. It grows with context length, so a long prompt or a big document can add several gigabytes on top of the weights. If you plan to feed large documents, leave headroom, or set a smaller context window in your tool.

## Which memory matters

> **Diagram:** Where a local model can live. GPU VRAM is fastest but limited by the card. System RAM on the CPU is slower but usually larger. Apple Silicon unified memory is shared by CPU and GPU, so a large amount of RAM can hold big models.
> Discrete GPU: Fastest when the whole model fits in VRAM; VRAM is the limit (8, 12, 24 GB cards); Offload some layers to RAM if it does not fit, with a speed hit. CPU or Apple Silicon: Runs from system RAM, no GPU needed; Slower per token than a good GPU; Apple unified memory lets big models run on one machine.

## A practical sizing guide

| You have | Comfortable choices |
|---|---|
| 8 GB RAM | Small models (about 3B at 4-bit), short context |
| 16 GB RAM | 7B to 8B at 4 or 8-bit |
| 32 GB RAM | 14B at 8-bit, or up to about 30B at 4-bit |
| 64 GB or more | 70B at 4-bit |

Leave memory for your OS and apps. If the model barely fits, it will swap and crawl.

## Before you download

1. Check the **parameter count** and **quantization** in the model name or file name.
2. Compute the estimate. Add 20 percent.
3. Compare it to free RAM or VRAM, not total.
4. If it does not fit, pick a smaller model or a lower-bit version rather than a bigger machine.

Not sure which tool to use? Read [Ollama vs LM Studio vs llama.cpp](https://devaiper.com/blog/ollama-vs-lm-studio-vs-llama-cpp).

## FAQ

### How much RAM do I need to run an LLM locally?

Multiply the parameter count by the bytes per weight, then add 10 to 20 percent for context and overhead. At 4-bit quantization that is about 0.5 to 0.6 bytes per parameter, so a 7B model needs roughly 5 GB and a 70B model roughly 42 GB.

### Can I run an LLM without a GPU?

Yes. Quantized models run on the CPU using system RAM. It is slower than on a GPU, but small models are usable on a laptop.

### Does context length use memory?

Yes. The KV cache grows with the number of tokens in context, so long prompts need extra memory on top of the weights.

### What is the difference between RAM and VRAM?

VRAM is the memory on your graphics card, and speed is best when the whole model fits in it. System RAM is slower but larger. Apple Silicon uses unified memory shared by CPU and GPU.

