Local LLMs

What is LLM quantization? Q4 vs Q8 explained

A model is a huge pile of numbers called weights. Stored as 16-bit values, a 7B model is about 14 GB. Quantization rounds those numbers to coarser values, stored in fewer bits. Fewer bits means a smaller file, less memory and usually faster generation.

The idea in one picture#

Imagine storing a photo with fewer colors. At 256 colors it looks nearly identical. At 16 colors you start to see banding. Weights behave the same way: a little rounding is invisible, a lot hurts.

16-bitweightsfull precision,largest8-bit (Q8)about half thesize,near-lossless4-bit (Q4)about aquarter, smallloss2 to 3-bittiny,noticeablequality drop
Each step down saves memory and costs some accuracy.

What the labels mean#

LabelAbout how many bitsSize vs 16-bitQualityUse when
F16 / BF1616100%ReferenceYou have lots of memory
Q88about 50%Near-identicalMemory is plentiful
Q55about 33%Very closeGood balance
Q44 to 5about 28 to 30%Slightly lowerDefault for local use
Q3 / Q22 to 315 to 20%Visible dropsLast resort

Names like Q4_K_M add detail: the K variants group weights smartly, and S, M, L are small, medium and large mixes that keep the most sensitive layers at higher precision.

GGUF#

GGUF is the single-file format used by llama.cpp and the tools built on it (Ollama, LM Studio). When you see a model file ending in .gguf with Q4_K_M in the name, that is a 4-bit quantized model ready to run locally.

Does quality really drop?#

A little, and it depends on the task:

  • Chat and summarizing: hard to notice at Q4.
  • Code, math and precise instruction following: more sensitive; prefer Q5 to Q8 if you can afford it.
  • Smaller models lose more from quantization than bigger ones, so a larger model at Q4 often beats a smaller one at Q8 in the same memory.

That last point is the useful rule: with a fixed memory budget, prefer a bigger model at 4-bit over a smaller model at 8-bit.

Test, do not guess#

Run your own cases on two quantizations and compare. A small eval tells you in minutes whether the lighter version is good enough for your task.

What it saves in practice#

Model16-bit4-bit (approx.)
7B14 GB4 to 5 GB
14B28 GB8 to 9 GB
70B140 GBabout 42 GB

Memory math in full: how much RAM to run an LLM locally. Tools that load these files: Ollama vs LM Studio vs llama.cpp.

Frequently asked questions

What is quantization in LLMs?

Quantization reduces the precision of a model's weights, for example from 16-bit to 8-bit or 4-bit numbers, so the model uses less memory and often runs faster, at a small cost in accuracy.

What does Q4 or Q8 mean?

The number is roughly the bits stored per weight. Q8 is about 8 bits, Q4 about 4 bits. Lower numbers are smaller and faster but lose more quality.

Which quantization should I choose?

A 4-bit variant is the usual sweet spot for most local use. Choose 8-bit if you have memory to spare and want to stay closer to full quality, and avoid going below 4-bit for tasks that need accuracy.

What is GGUF?

GGUF is a file format used by llama.cpp and tools built on it, such as Ollama and LM Studio, that stores quantized model weights in a single file.