What is LLM quantization? Q4 vs Q8 explained
A model is a huge pile of numbers called weights. Stored as 16-bit values, a 7B model is about 14 GB. Quantization rounds those numbers to coarser values, stored in fewer bits. Fewer bits means a smaller file, less memory and usually faster generation.
The idea in one picture#
Imagine storing a photo with fewer colors. At 256 colors it looks nearly identical. At 16 colors you start to see banding. Weights behave the same way: a little rounding is invisible, a lot hurts.
What the labels mean#
| Label | About how many bits | Size vs 16-bit | Quality | Use when |
|---|---|---|---|---|
| F16 / BF16 | 16 | 100% | Reference | You have lots of memory |
| Q8 | 8 | about 50% | Near-identical | Memory is plentiful |
| Q5 | 5 | about 33% | Very close | Good balance |
| Q4 | 4 to 5 | about 28 to 30% | Slightly lower | Default for local use |
| Q3 / Q2 | 2 to 3 | 15 to 20% | Visible drops | Last resort |
Names like Q4_K_M add detail: the K variants group weights smartly, and S, M, L are small, medium and large mixes that keep the most sensitive layers at higher precision.
GGUF#
GGUF is the single-file format used by llama.cpp and the tools built on it (Ollama, LM Studio). When you see a model file ending in .gguf with Q4_K_M in the name, that is a 4-bit quantized model ready to run locally.
Does quality really drop?#
A little, and it depends on the task:
- Chat and summarizing: hard to notice at Q4.
- Code, math and precise instruction following: more sensitive; prefer Q5 to Q8 if you can afford it.
- Smaller models lose more from quantization than bigger ones, so a larger model at Q4 often beats a smaller one at Q8 in the same memory.
That last point is the useful rule: with a fixed memory budget, prefer a bigger model at 4-bit over a smaller model at 8-bit.
Test, do not guess#
Run your own cases on two quantizations and compare. A small eval tells you in minutes whether the lighter version is good enough for your task.
What it saves in practice#
| Model | 16-bit | 4-bit (approx.) |
|---|---|---|
| 7B | 14 GB | 4 to 5 GB |
| 14B | 28 GB | 8 to 9 GB |
| 70B | 140 GB | about 42 GB |
Memory math in full: how much RAM to run an LLM locally. Tools that load these files: Ollama vs LM Studio vs llama.cpp.
Frequently asked questions
What is quantization in LLMs?
Quantization reduces the precision of a model's weights, for example from 16-bit to 8-bit or 4-bit numbers, so the model uses less memory and often runs faster, at a small cost in accuracy.
What does Q4 or Q8 mean?
The number is roughly the bits stored per weight. Q8 is about 8 bits, Q4 about 4 bits. Lower numbers are smaller and faster but lose more quality.
Which quantization should I choose?
A 4-bit variant is the usual sweet spot for most local use. Choose 8-bit if you have memory to spare and want to stay closer to full quality, and avoid going below 4-bit for tasks that need accuracy.
What is GGUF?
GGUF is a file format used by llama.cpp and tools built on it, such as Ollama and LM Studio, that stores quantized model weights in a single file.
Prefer plain text? Read this page as Markdown.