What is RAG? Retrieval-augmented generation explained
An LLM knows what it learned in training. It does not know your handbook, your tickets or last week's incident report. RAG fixes that without retraining: at question time, fetch the relevant pieces of your documents and put them in the prompt.
The pipeline in two halves#
Step by step#
1. Chunk. Split documents into passages of a few hundred tokens, usually with a little overlap so ideas are not cut in half. Respect structure: split on headings and paragraphs before you split on length.
2. Embed. An embedding model turns each chunk into a vector, a list of numbers. Chunks with similar meaning end up near each other.
3. Store. Save the vectors with the original text and metadata (source, page, date) in an index.
4. Retrieve. Embed the user's question and find the nearest chunks. Take the top few, often 3 to 8.
5. Augment and generate. Put the chunks in the prompt, mark them as source material, and ask the model to answer using only them.
Answer the question using only the sources below. If the sources do not
contain the answer, say you do not know. Cite the source id for each claim.
<sources>
<source id="handbook-12">Refunds are available within 30 days of purchase...</source>
<source id="handbook-13">Annual plans are refunded pro rata...</source>
</sources>
<question>Can I get a refund on an annual plan?</question>A tiny retrieval sketch in TypeScript#
type Chunk = { id: string; text: string; vector: number[] };
function cosine(a: number[], b: number[]): number {
let dot = 0, na = 0, nb = 0;
for (let i = 0; i < a.length; i++) {
dot += a[i] * b[i];
na += a[i] ** 2;
nb += b[i] ** 2;
}
return dot / (Math.sqrt(na) * Math.sqrt(nb));
}
function topK(queryVector: number[], index: Chunk[], k = 5): Chunk[] {
return [...index]
.sort((x, y) => cosine(queryVector, y.vector) - cosine(queryVector, x.vector))
.slice(0, k);
}Your embedding model produces the vectors; everything else is arithmetic. A vector database does the same thing faster at scale.
Where RAG fails (and the fix)#
| Symptom | Likely cause | Fix |
|---|---|---|
| Right doc exists, answer is wrong | Retrieval missed it | Smaller chunks, hybrid search, rerank |
| Exact terms (error codes, SKUs) not found | Embeddings blur exact strings | Add keyword (BM25) search and combine |
| Answer mixes two documents | Chunks lack context | Add title or section to each chunk |
| Made-up details | Model not told to stay in sources | "Only use the sources, say if missing" |
| Stale answers | Index not refreshed | Re-index changed docs; store dates |
| Slow and expensive | Too many chunks | Lower k, rerank, use prompt caching |
Measure retrieval separately from generation#
Two questions, two metrics:
- Did retrieval return the chunk that contains the answer? (recall at k)
- Given the right chunks, did the model answer correctly and stay grounded?
If you only judge final answers you will tune the wrong half. Build a small set of real questions with known source passages. See LLM evals explained.
RAG is context engineering#
RAG decides what enters the context window. The same thinking applies: retrieve less, label it clearly and keep instructions separate from data. When you wonder whether to use RAG, fine-tuning or tools, read RAG vs fine-tuning vs MCP.
Frequently asked questions
What does RAG stand for?
Retrieval-augmented generation: retrieve relevant information first, then have the model generate an answer using it.
Why use RAG instead of pasting documents into the prompt?
Your knowledge base is usually bigger than the context window, and sending everything every time is slow, expensive and less accurate. RAG sends only the few relevant chunks.
What is an embedding?
A list of numbers that represents the meaning of a piece of text, so texts with similar meaning have similar vectors. It lets you search by meaning instead of exact keywords.
Why does my RAG system give wrong answers?
Most often retrieval returned the wrong chunks: bad chunk size, no keyword matching, no reranking, or the question did not match the document wording. Fix retrieval before changing the prompt.
Do I need a vector database for RAG?
Not at the start. For a few thousand chunks, an in-memory index or a Postgres extension is enough. Pick a dedicated vector store when scale or filtering requires it.
Prefer plain text? Read this page as Markdown.