Incident Response
The first five minutes, blast radius, cascades, graphs that mislead, partial failures, mitigation, communication and the post-mortem.
10 lessons · about 90 minutes
What you'll go through
- 01The page arrives at 04:12. What do you do first?9 min
- 02Everything is red and none of it is yours9 min
- 03Nothing changed, and yet9 min
- 04Slow, or full? They look identical and are not9 min
- 05One slow service and the whole platform is down9 min
- 06The metric recovered and the users did not9 min
- 07Only 3% of requests fail, and it is always the same 3%9 min
- 08Rolling back before you understand it is the right call9 min
- 09The status update nobody had to ask for8 min
- 10Blameless is a technique, not a courtesy10 min
Included in this pack
Observability — Metrics, logs and traces; SLOs and alerts; and what to do during an incident. Three courses, 30 lessons.
Also in this pack

Metrics, Logs and Traces
What each signal answers, why percentiles beat averages, how histograms work, and where cardinality and sampling send the bill.
10 lessons · ~85 min$19.90

SLOs and Alerts
Indicators, targets, error budgets, burn-rate alerting, what an alert must contain, and who gets woken up.
10 lessons · ~85 min$19.90