"Who ran that command?" is the question that ends a postmortem before it's actually started anything useful. The engineer who ran it stops talking, or starts defending, and everyone in the room learns the same lesson: if something goes wrong on your watch, the conversation afterward is about you, not about the system. That lesson gets absorbed fast, and it produces exactly the wrong incentive — hide mistakes, under-report near-misses, and definitely don't volunteer the detail that would actually explain what happened, because that detail is now evidence.
A blameless postmortem asks a different question: not "who did it," but "why did the system allow it to happen, and why did it feel like the reasonable thing to do at the time." That reframe isn't softer — it's more demanding, because it doesn't let the investigation stop at a human decision. It has to go one level deeper, to whatever set of conditions made that decision look correct in the moment.
Note
The single sentence that captures the whole practice: fix systems, not people. A person can be replaced and the exact same incident will recur, because the thing that actually caused it — a missing alert, an unclear runbook, an easy-to-misread dashboard — is still sitting there waiting for the next person to hit it.
Blameful vs. blameless, side by side
| Traditional (blameful) | Blameless |
|---|---|
| "Who deployed the broken change?" | "What in our review or testing process let it ship?" |
| Focus on the individual | Focus on the system and process |
| Defensiveness, hidden details | Open, honest reconstruction of events |
| Same incident recurs with a different name attached | Root cause addressed, recurrence actually prevented |
| Fear of the next postmortem | Curiosity about what the next one might reveal |
The recurrence row is the one that should end any argument about which approach is more rigorous. A blameful postmortem that identifies "Sarah forgot to check the dashboard" produces exactly one fix: don't be Sarah. The next person in that role, doing the exact same reasonable-seeming thing under the exact same pressure, hits the exact same failure. A blameless postmortem that identifies "the dashboard didn't surface the one metric that mattered, and nothing forced a check before deploy" produces a fix that actually holds regardless of who's on call next.
What belongs in one
1. Summary — what happened, in plain language, one paragraph.
2. Timeline — every relevant event, with timestamps, no editorializing.
3. Impact — who and what was affected, and for how long.
4. Root cause(s) — the actual chain of conditions, not a single scapegoat.
5. What went well — the detection, the response, whatever worked.
6. Action items — specific, owned, dated fixes. Not "be more careful."
7. Follow-up — did the action items actually ship, and did they help?Step 7 is the one that separates postmortems that matter from postmortems that are theater. An action item that never gets tracked to completion is just a paragraph that made the meeting feel productive. The postmortem process itself needs the same rigor as the system it's investigating — if "add monitoring for X" sits open for six months, that's a process failure worth its own review.
Ground rules that keep it blameless in practice
Ask "why" and "what," never "who" or "how could you." Assume everyone made the most reasonable decision available to them with the information and time pressure they had — because they almost always did. Keep it constructive: the shared goal in the room is a stronger system, not a settled score.
Why blameless is the harder discipline, not the easier one
There's a version of "blameless" that's actually just "consequence-free," and that version doesn't work — it lets genuine negligence or repeated carelessness slide under a process designed for something else. Real blameless culture isn't the absence of accountability; it's accountability aimed at the system and the process that let a mistake become an incident, which is almost always the more useful target. An engineer who fat-fingered a command isn't the problem if the deploy process allowed that single command to take down production with no review, no staging step, and no automatic rollback. Fix that, and the next fat-fingered command — because there will be one — becomes a non-event instead of an outage.
Tomorrow's article closes out the SRE Foundations phase with incident management itself: the real-time discipline of detecting, triaging, and resolving an incident well while it's still in progress — which is exactly the process a good blameless postmortem exists to keep improving.





