# LLM evals and observability: test AI features properly

> Learn to build evals for LLM features: datasets, code and model graders, RAG and agent evals, CI, tracing and production monitoring, using a support-ticket triage bot.

Source: https://devaiper.com/courses/llm-evals
Published: 2026-10-08

Stop shipping prompts on vibes, in eight parts. Every part has a short concept video, then a hands-on video where we build it live. The written version of each part appears on this site as its video goes live, with code you can copy.

## Course outline

1. Why evals, and your first one (coming soon): The cheapest test that catches the most.
2. Building a dataset (coming soon): Cases that reflect real traffic.
3. Code-based graders (coming soon): Exact match, schema checks and rules.
4. Model-graded evals (coming soon): Using an LLM as the judge, carefully.
5. Evals for RAG and agents (coming soon): Retrieval quality and multi-step behavior.
6. Evals in CI (coming soon): Fail the pull request when quality drops.
7. Tracing and logging (coming soon): See what the model actually did.
8. Production monitoring and the improvement loop (coming soon): Turn real failures into new test cases.

## FAQ

### What is an LLM eval?

A repeatable test for an LLM feature: a set of inputs, a way to grade each output, and a score you can track as you change the prompt or model.

### Why not just test prompts by hand?

Hand testing misses regressions and does not scale. An eval gives you a number you can compare before and after every change.

