A blameless write-up of what happened. Structure:
- Summary — one paragraph, what happened
- Impact — duration, services affected, customers affected, financial estimate if applicable
- Timeline — minute-by-minute log of detection, response, key decisions, resolution
- Root cause — the actual underlying cause, not just the proximate trigger
- Five whys — drill into why the root cause was possible
- What went well — what helped resolve quickly
- What went poorly — gaps, slow detection, missing runbooks, etc.
- Action items — concrete tickets with owners and due dates
Blameless means: no naming individuals as fault. Focus on systemic causes: missing alerts, insufficient testing, unclear runbooks, complex deployment process. The assumption is people are doing their best with the information they have.