LLM evals explained: how to test AI features properly
You changed the prompt. It looks better on the three examples you tried. Did you make it better, or did you break two cases you did not try? Without an eval you do not know. An eval turns "feels better" into a number.
The three parts#
1. Build a dataset from real inputs#
Start with 20 to 50 cases. Pull them from real tickets, logs or user messages, not from your imagination. Include the easy cases, the awkward ones, and every failure you already know about.
type Case = { id: string; input: string; expected: "billing" | "bug" | "other" };
const cases: Case[] = [
{ id: "t1", input: "I was charged twice this month.", expected: "billing" },
{ id: "t2", input: "The export button does nothing.", expected: "bug" },
{ id: "t3", input: "Do you have a student discount?", expected: "other" },
// 17 more, including the weird ones
];2. Prefer code graders when you can#
If the answer can be checked by code, use code. It is fast, free and deterministic: exact match, JSON schema validity, regex, "contains this field", "code compiles", "tests pass".
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
async function classify(input: string): Promise<string> {
const res = await client.messages.create({
model: "claude-opus-5-5",
max_tokens: 20,
system: "Classify the support ticket as billing, bug or other. Reply with one word.",
messages: [{ role: "user", content: input }],
});
const block = res.content.find((b) => b.type === "text");
return block && block.type === "text" ? block.text.trim().toLowerCase() : "";
}
let passed = 0;
for (const c of cases) {
const got = await classify(c.input);
const ok = got === c.expected;
if (ok) passed++;
else console.log(`FAIL ${c.id}: expected ${c.expected}, got "${got}"`);
}
console.log(`score: ${passed}/${cases.length}`);That is a complete eval. It is small enough to run before every commit.
3. Use an LLM judge for fuzzy qualities#
Some things code cannot grade: "is this summary faithful to the source?", "is the tone polite?". A second model call can score against a rubric.
You are grading a summary against its source.
Score 1 if every claim in the summary is supported by the source, else 0.
Reply with JSON: {"score": 0 or 1, "reason": "one sentence"}.
<source>...</source>
<summary>...</summary>Rules for judges: write a specific rubric, ask for a reason so you can audit it, use binary or small scales, and check a sample by hand to confirm the judge agrees with you.
4. Grow the dataset from failures#
Every bug a user reports becomes a new case. That is how an eval gets better than your first guess.
Mistakes to avoid#
| Mistake | Better |
|---|---|
| Invented, too-easy test cases | Real inputs and real failures |
| Judging only the final answer in RAG | Score retrieval and generation separately (RAG) |
| One run, one number | Run several times; outputs vary |
| Trusting an LLM judge blindly | Spot-check with humans |
| Evals that never run | Put them in CI so regressions fail the build |
| Changing prompt and model at once | Change one thing per run |
Where evals fit#
Evals are the safety net under everything else: prompt engineering, structured outputs, agents. Discernment, the third D of the 4D framework, done at scale.
Frequently asked questions
What is an LLM eval?
A repeatable test for an LLM feature: a dataset of inputs with expected outcomes, a grader that scores each output, and an aggregate score you can compare before and after a change.
How many test cases do I need for an eval?
Start with 20 to 50 real examples that cover your common cases and known failures. A small honest set beats a large invented one, and you can grow it from production failures.
What is LLM-as-a-judge?
Using a model to grade another model's output against a rubric. It handles fuzzy qualities like helpfulness or tone, but it needs clear criteria and should be spot-checked against human judgment.
How is an eval different from a unit test?
A unit test expects an exact result. LLM output varies, so evals score across many cases and track a pass rate or average rather than demanding identical text.
Prefer plain text? Read this page as Markdown.