Skip to content

Evaluating AI Agents

Building a test set, deterministic checks before model judges, rubrics that discriminate, measuring the trajectory, and cost and latency as results.

10 lessons · about 92 minutes

What you'll go through

  • 01You tried it five times and it worked9 min
  • 02«The output looks good» is not a measurement9 min
  • 03Where a test set comes from when you have no data9 min
  • 04Most of what you want to verify is a string comparison9 min
  • 05Grading with the thing you are grading9 min
  • 06Every answer scores four out of five10 min
  • 07Measuring the trajectory, not just the result9 min
  • 08It passes, and it costs eleven dollars9 min
  • 09The suite passes and production does not9 min
  • 10Small, fast, and in CI10 min