Free course · 8 parts · Intermediate
LLM evals and observability: test AI features properly
Learn to build evals for LLM features: datasets, code and model graders, RAG and agent evals, CI, tracing and production monitoring, using a support-ticket triage bot.
Course outline
- 1Why evals, and your first oneComing soon
The cheapest test that catches the most.
- 2Building a datasetComing soon
Cases that reflect real traffic.
- 3Code-based gradersComing soon
Exact match, schema checks and rules.
- 4Model-graded evalsComing soon
Using an LLM as the judge, carefully.
- 5Evals for RAG and agentsComing soon
Retrieval quality and multi-step behavior.
- 6Evals in CIComing soon
Fail the pull request when quality drops.
- 7Tracing and loggingComing soon
See what the model actually did.
- 8Production monitoring and the improvement loopComing soon
Turn real failures into new test cases.
Stop shipping prompts on vibes, in eight parts. Every part has a short concept video, then a hands-on video where we build it live. The written version of each part appears on this site as its video goes live, with code you can copy.
Frequently asked questions
What is an LLM eval?
A repeatable test for an LLM feature: a set of inputs, a way to grade each output, and a score you can track as you change the prompt or model.
Why not just test prompts by hand?
Hand testing misses regressions and does not scale. An eval gives you a number you can compare before and after every change.