# What is RAG? Retrieval-augmented generation explained

> RAG gives an LLM your own documents at answer time: split, embed, search, then prompt. How the pipeline works, where it fails and how to make it accurate.

Source: https://devaiper.com/blog/what-is-rag
Published: 2026-10-08
Topics: RAG, Embeddings, Context engineering

**Short answer:** RAG (retrieval-augmented generation) finds the most relevant pieces of your documents for each question and puts them in the prompt, so the model answers from your data instead of from memory. Most RAG quality problems are retrieval problems.

An LLM knows what it learned in training. It does not know your handbook, your tickets or last week's incident report. **RAG** fixes that without retraining: at question time, fetch the relevant pieces of your documents and put them in the prompt.

## The pipeline in two halves

> **Diagram:** The RAG pipeline in two stages. Indexing, done once: load documents, split into chunks, embed each chunk, store in an index. Querying, done per question: embed the question, search the index for similar chunks, build a prompt with those chunks, generate the answer.
> Load documents → Split into chunks (indexing, done once) → Embed and store (vector index) → Embed the question (querying, per request) → Search for similar chunks → Prompt + answer (model cites the chunks)
> The first three steps run when documents change. The last three run on every question.

## Step by step

**1. Chunk.** Split documents into passages of a few hundred tokens, usually with a little overlap so ideas are not cut in half. Respect structure: split on headings and paragraphs before you split on length.

**2. Embed.** An embedding model turns each chunk into a vector, a list of numbers. Chunks with similar meaning end up near each other.

**3. Store.** Save the vectors with the original text and metadata (source, page, date) in an index.

**4. Retrieve.** Embed the user's question and find the nearest chunks. Take the top few, often 3 to 8.

**5. Augment and generate.** Put the chunks in the prompt, mark them as source material, and ask the model to answer using only them.

```text
Answer the question using only the sources below. If the sources do not
contain the answer, say you do not know. Cite the source id for each claim.

<sources>
<source id="handbook-12">Refunds are available within 30 days of purchase...</source>
<source id="handbook-13">Annual plans are refunded pro rata...</source>
</sources>

<question>Can I get a refund on an annual plan?</question>
```

## A tiny retrieval sketch in TypeScript

```ts
type Chunk = { id: string; text: string; vector: number[] };

function cosine(a: number[], b: number[]): number {
  let dot = 0, na = 0, nb = 0;
  for (let i = 0; i < a.length; i++) {
    dot += a[i] * b[i];
    na += a[i] ** 2;
    nb += b[i] ** 2;
  }
  return dot / (Math.sqrt(na) * Math.sqrt(nb));
}

function topK(queryVector: number[], index: Chunk[], k = 5): Chunk[] {
  return [...index]
    .sort((x, y) => cosine(queryVector, y.vector) - cosine(queryVector, x.vector))
    .slice(0, k);
}
```

Your embedding model produces the vectors; everything else is arithmetic. A vector database does the same thing faster at scale.

## Where RAG fails (and the fix)

| Symptom | Likely cause | Fix |
|---|---|---|
| Right doc exists, answer is wrong | Retrieval missed it | Smaller chunks, hybrid search, rerank |
| Exact terms (error codes, SKUs) not found | Embeddings blur exact strings | Add keyword (BM25) search and combine |
| Answer mixes two documents | Chunks lack context | Add title or section to each chunk |
| Made-up details | Model not told to stay in sources | "Only use the sources, say if missing" |
| Stale answers | Index not refreshed | Re-index changed docs; store dates |
| Slow and expensive | Too many chunks | Lower k, rerank, use [prompt caching](https://devaiper.com/blog/prompt-caching-explained) |

## Measure retrieval separately from generation

Two questions, two metrics:

1. **Did retrieval return the chunk that contains the answer?** (recall at k)
2. **Given the right chunks, did the model answer correctly and stay grounded?**

If you only judge final answers you will tune the wrong half. Build a small set of real questions with known source passages. See [LLM evals explained](https://devaiper.com/blog/llm-evals-explained).

## RAG is context engineering

RAG decides what enters the [context window](https://devaiper.com/blog/what-is-context-engineering). The same thinking applies: retrieve less, label it clearly and keep instructions separate from data. When you wonder whether to use RAG, fine-tuning or tools, read [RAG vs fine-tuning vs MCP](https://devaiper.com/blog/rag-vs-fine-tuning-vs-mcp).

## FAQ

### What does RAG stand for?

Retrieval-augmented generation: retrieve relevant information first, then have the model generate an answer using it.

### Why use RAG instead of pasting documents into the prompt?

Your knowledge base is usually bigger than the context window, and sending everything every time is slow, expensive and less accurate. RAG sends only the few relevant chunks.

### What is an embedding?

A list of numbers that represents the meaning of a piece of text, so texts with similar meaning have similar vectors. It lets you search by meaning instead of exact keywords.

### Why does my RAG system give wrong answers?

Most often retrieval returned the wrong chunks: bad chunk size, no keyword matching, no reranking, or the question did not match the document wording. Fix retrieval before changing the prompt.

### Do I need a vector database for RAG?

Not at the start. For a few thousand chunks, an in-memory index or a Postgres extension is enough. Pick a dedicated vector store when scale or filtering requires it.

