# LLM evals explained: how to test AI features properly

> LLM evals are repeatable tests for AI features: a dataset, a grader and a score. How to build your first eval with code graders and an LLM judge, in TypeScript.

Source: https://devaiper.com/blog/llm-evals-explained
Published: 2026-10-08
Topics: Evals, Testing, Prompt engineering

**Short answer:** An eval is a set of test cases, a way to grade each output, and a score you track. Start with 20 real cases and code-based graders, add an LLM judge for fuzzy qualities, and run it every time you change a prompt or a model.

You changed the prompt. It looks better on the three examples you tried. Did you make it better, or did you break two cases you did not try? Without an eval you do not know. **An eval turns "feels better" into a number.**

## The three parts

> **Diagram:** The three parts of an LLM eval. A dataset of test cases feeds your prompt or app, which produces outputs. A grader scores each output. The scores are combined into a pass rate you track over time.
> Dataset (inputs + expected outcomes) → Your prompt or app (produces an output) → Grader (code check or LLM judge) → Score (pass rate, tracked over time)
> Run it on every change. Compare the score before and after.

## 1. Build a dataset from real inputs

Start with 20 to 50 cases. Pull them from real tickets, logs or user messages, not from your imagination. Include the easy cases, the awkward ones, and every failure you already know about.

```ts
type Case = { id: string; input: string; expected: "billing" | "bug" | "other" };

const cases: Case[] = [
  { id: "t1", input: "I was charged twice this month.", expected: "billing" },
  { id: "t2", input: "The export button does nothing.", expected: "bug" },
  { id: "t3", input: "Do you have a student discount?", expected: "other" },
  // 17 more, including the weird ones
];
```

## 2. Prefer code graders when you can

If the answer can be checked by code, use code. It is fast, free and deterministic: exact match, JSON schema validity, regex, "contains this field", "code compiles", "tests pass".

```ts
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

async function classify(input: string): Promise<string> {
  const res = await client.messages.create({
    model: "claude-opus-5-5",
    max_tokens: 20,
    system: "Classify the support ticket as billing, bug or other. Reply with one word.",
    messages: [{ role: "user", content: input }],
  });
  const block = res.content.find((b) => b.type === "text");
  return block && block.type === "text" ? block.text.trim().toLowerCase() : "";
}

let passed = 0;
for (const c of cases) {
  const got = await classify(c.input);
  const ok = got === c.expected;
  if (ok) passed++;
  else console.log(`FAIL ${c.id}: expected ${c.expected}, got "${got}"`);
}
console.log(`score: ${passed}/${cases.length}`);
```

That is a complete eval. It is small enough to run before every commit.

## 3. Use an LLM judge for fuzzy qualities

Some things code cannot grade: "is this summary faithful to the source?", "is the tone polite?". A second model call can score against a rubric.

```text
You are grading a summary against its source.
Score 1 if every claim in the summary is supported by the source, else 0.
Reply with JSON: {"score": 0 or 1, "reason": "one sentence"}.

<source>...</source>
<summary>...</summary>
```

Rules for judges: write a **specific rubric**, ask for a **reason** so you can audit it, use **binary or small scales**, and **check a sample by hand** to confirm the judge agrees with you.

## 4. Grow the dataset from failures

Every bug a user reports becomes a new case. That is how an eval gets better than your first guess.

> **Diagram:** The eval improvement loop. Ship the feature, collect real failures, add them as test cases, improve the prompt against the whole set, and ship again.
> Ship → ... → Collect failures (tickets, logs, thumbs-down) → ... → Add as cases → ... → Improve and re-run

## Mistakes to avoid

| Mistake | Better |
|---|---|
| Invented, too-easy test cases | Real inputs and real failures |
| Judging only the final answer in RAG | Score retrieval and generation separately ([RAG](https://devaiper.com/blog/what-is-rag)) |
| One run, one number | Run several times; outputs vary |
| Trusting an LLM judge blindly | Spot-check with humans |
| Evals that never run | Put them in CI so regressions fail the build |
| Changing prompt and model at once | Change one thing per run |

## Where evals fit

Evals are the safety net under everything else: [prompt engineering](https://devaiper.com/blog/prompt-engineering-techniques), [structured outputs](https://devaiper.com/blog/how-to-get-json-from-an-llm), [agents](https://devaiper.com/blog/ai-agents-vs-workflows). Discernment, the third D of the [4D framework](https://devaiper.com/blog/4d-framework-ai-fluency), done at scale.

## FAQ

### What is an LLM eval?

A repeatable test for an LLM feature: a dataset of inputs with expected outcomes, a grader that scores each output, and an aggregate score you can compare before and after a change.

### How many test cases do I need for an eval?

Start with 20 to 50 real examples that cover your common cases and known failures. A small honest set beats a large invented one, and you can grow it from production failures.

### What is LLM-as-a-judge?

Using a model to grade another model's output against a rubric. It handles fuzzy qualities like helpfulness or tone, but it needs clear criteria and should be spot-checked against human judgment.

### How is an eval different from a unit test?

A unit test expects an exact result. LLM output varies, so evals score across many cases and track a pass rate or average rather than demanding identical text.

