Free course · 8 parts · Intermediate

LLM evals and observability: test AI features properly

Learn to build evals for LLM features: datasets, code and model graders, RAG and agent evals, CI, tracing and production monitoring, using a support-ticket triage bot.

Follow on YouTube

Course outline

  1. 1
    Why evals, and your first one

    The cheapest test that catches the most.

    Coming soon
  2. 2
    Building a dataset

    Cases that reflect real traffic.

    Coming soon
  3. 3
    Code-based graders

    Exact match, schema checks and rules.

    Coming soon
  4. 4
    Model-graded evals

    Using an LLM as the judge, carefully.

    Coming soon
  5. 5
    Evals for RAG and agents

    Retrieval quality and multi-step behavior.

    Coming soon
  6. 6
    Evals in CI

    Fail the pull request when quality drops.

    Coming soon
  7. 7
    Tracing and logging

    See what the model actually did.

    Coming soon
  8. 8
    Production monitoring and the improvement loop

    Turn real failures into new test cases.

    Coming soon

Stop shipping prompts on vibes, in eight parts. Every part has a short concept video, then a hands-on video where we build it live. The written version of each part appears on this site as its video goes live, with code you can copy.

Frequently asked questions

What is an LLM eval?

A repeatable test for an LLM feature: a set of inputs, a way to grade each output, and a score you can track as you change the prompt or model.

Why not just test prompts by hand?

Hand testing misses regressions and does not scale. An eval gives you a number you can compare before and after every change.