Evaluating AI Agents
Building a test set, deterministic checks before model judges, rubrics that discriminate, measuring the trajectory, and cost and latency as results.
10 lessons · about 92 minutes
What you'll go through
- 01You tried it five times and it worked9 min
- 02«The output looks good» is not a measurement9 min
- 03Where a test set comes from when you have no data9 min
- 04Most of what you want to verify is a string comparison9 min
- 05Grading with the thing you are grading9 min
- 06Every answer scores four out of five10 min
- 07Measuring the trajectory, not just the result9 min
- 08It passes, and it costs eleven dollars9 min
- 09The suite passes and production does not9 min
- 10Small, fast, and in CI10 min
Included in this pack
AI Agents — Tool design, context and cost, and how to evaluate an agent without fooling yourself. Three courses, 30 lessons.
Also in this pack

Agent Tools
The agent loop, tool descriptions and schemas, error handling, parallel calls, and how to design a tool surface you can still say no to.
10 lessons · ~86 min$19.90

Context Engineering
Token budgets, prompt caching and what silently breaks it, pruning versus summarising, memory between sessions, and keeping tool output out of the window.
10 lessons · ~91 min$19.90