# What is LLM quantization? Q4 vs Q8 explained

> Quantization stores a model's weights in fewer bits so it fits in less memory and runs faster. What Q4, Q8 and GGUF mean and how much quality you give up.

Source: https://devaiper.com/blog/what-is-llm-quantization
Published: 2026-10-08
Topics: Quantization, Local LLMs, GGUF

**Short answer:** Quantization shrinks a model by storing each weight in fewer bits, for example 4 instead of 16. A 4-bit model uses roughly a quarter of the memory with a small quality loss, which is why it is the default for local LLMs.

A model is a huge pile of numbers called **weights**. Stored as 16-bit values, a 7B model is about 14 GB. **Quantization** rounds those numbers to coarser values, stored in fewer bits. Fewer bits means a smaller file, less memory and usually faster generation.

## The idea in one picture

Imagine storing a photo with fewer colors. At 256 colors it looks nearly identical. At 16 colors you start to see banding. Weights behave the same way: a little rounding is invisible, a lot hurts.

> **Diagram:** Quantization shrinks a model. Original 16-bit weights are rounded to 8-bit, which halves the size with very little quality loss, or to 4-bit, which cuts it to about a quarter with a small quality loss, enabling large models on consumer hardware.
> 16-bit weights (full precision, largest) → 8-bit (Q8) (about half the size, near-lossless) → 4-bit (Q4) (about a quarter, small loss) → 2 to 3-bit (tiny, noticeable quality drop)
> Each step down saves memory and costs some accuracy.

## What the labels mean

| Label | About how many bits | Size vs 16-bit | Quality | Use when |
|---|---|---|---|---|
| F16 / BF16 | 16 | 100% | Reference | You have lots of memory |
| Q8 | 8 | about 50% | Near-identical | Memory is plentiful |
| Q5 | 5 | about 33% | Very close | Good balance |
| **Q4** | 4 to 5 | about 28 to 30% | Slightly lower | **Default for local use** |
| Q3 / Q2 | 2 to 3 | 15 to 20% | Visible drops | Last resort |

Names like `Q4_K_M` add detail: the `K` variants group weights smartly, and `S`, `M`, `L` are small, medium and large mixes that keep the most sensitive layers at higher precision.

## GGUF

**GGUF** is the single-file format used by llama.cpp and the tools built on it (Ollama, LM Studio). When you see a model file ending in `.gguf` with `Q4_K_M` in the name, that is a 4-bit quantized model ready to run locally.

## Does quality really drop?

A little, and it depends on the task:

- **Chat and summarizing:** hard to notice at Q4.
- **Code, math and precise instruction following:** more sensitive; prefer Q5 to Q8 if you can afford it.
- **Smaller models** lose more from quantization than bigger ones, so a larger model at Q4 often beats a smaller one at Q8 in the same memory.

That last point is the useful rule: **with a fixed memory budget, prefer a bigger model at 4-bit over a smaller model at 8-bit.**

## Test, do not guess

Run your own cases on two quantizations and compare. A small [eval](https://devaiper.com/blog/llm-evals-explained) tells you in minutes whether the lighter version is good enough for *your* task.

## What it saves in practice

| Model | 16-bit | 4-bit (approx.) |
|---|---|---|
| 7B | 14 GB | 4 to 5 GB |
| 14B | 28 GB | 8 to 9 GB |
| 70B | 140 GB | about 42 GB |

Memory math in full: [how much RAM to run an LLM locally](https://devaiper.com/blog/how-much-ram-to-run-an-llm-locally). Tools that load these files: [Ollama vs LM Studio vs llama.cpp](https://devaiper.com/blog/ollama-vs-lm-studio-vs-llama-cpp).

## FAQ

### What is quantization in LLMs?

Quantization reduces the precision of a model's weights, for example from 16-bit to 8-bit or 4-bit numbers, so the model uses less memory and often runs faster, at a small cost in accuracy.

### What does Q4 or Q8 mean?

The number is roughly the bits stored per weight. Q8 is about 8 bits, Q4 about 4 bits. Lower numbers are smaller and faster but lose more quality.

### Which quantization should I choose?

A 4-bit variant is the usual sweet spot for most local use. Choose 8-bit if you have memory to spare and want to stay closer to full quality, and avoid going below 4-bit for tasks that need accuracy.

### What is GGUF?

GGUF is a file format used by llama.cpp and tools built on it, such as Ollama and LM Studio, that stores quantized model weights in a single file.

