A bad postmortem starts with a culprit and works backward. A useless postmortem starts with “the system failed” and removes every decision until nobody owns anything. A useful postmortem does harder work: it reconstructs what people saw, what the system allowed, why actions were locally reasonable and which changes will reduce the chance or cost of recurrence.
Blamelessness is a method for obtaining better evidence. People hide less when the review is not a trial. That does not erase standards, negligence, deliberate misconduct or accountability. Those questions may require a separate personnel, legal, safety or regulatory process. Mixing them into the learning review can corrupt both.

Decide when a postmortem is required before the incident
Trigger criteria reduce politics. Google SRE describes thresholds such as meaningful user impact, data loss, manual intervention, long resolution time or a monitoring failure. A household, project or small team can use proportional versions: irreversible data loss, missed critical deadline, safety near-miss, financial harm, privacy exposure, repeated outage or any event whose workaround became the new normal.
Do not require a full review for every small mistake. Define a light version for contained incidents and a formal version when consequence, recurrence or uncertainty crosses the threshold. Any affected person should be able to request one. Record why a review was or was not opened.
Stabilize first, learn second
During an active incident, prioritize safety, containment, continuity and clear authority. Preserve logs and observations as you go, but do not force responders to write the polished history while they are still restoring service. Mark uncertain timestamps. Record consequential commands, changes and communications. Use the digital change log to capture evidence without turning response into paperwork theater.
Choose a review time soon enough that memory and records remain available, but not while participants are exhausted or still managing harm. For high-stakes medical, legal, security, employment or public-safety incidents, follow applicable professional and reporting requirements; this general protocol is not a substitute.
Give the review four roles
- Facilitator: protects the method, separates evidence from judgment and prevents one voice from owning the narrative.
- Scribe: maintains the shared timeline, unresolved questions and source links.
- Participants: contribute what they observed and decided, including uncertainty and workarounds.
- Action owners: accept specific follow-up work and verification. They need not be the people who acted during the incident.
If power differences make candor unsafe, gather private written input or use an independent facilitator. “Speak freely” is not a control. The review design must account for incentives.
Pass one: reconstruct the timeline
Build a chronology from logs, messages, tickets, sensor data, calendars, screenshots and direct recollection. Label each entry by source. Use ranges when time is uncertain. Include detection, first interpretation, escalation, mitigation, recovery and confirmation—not only the dramatic failure.
| Time | Observed signal | Available interpretation | Action | Result | Source |
|---|---|---|---|---|---|
| 09:12 | Service latency alert | Expected traffic spike | Raised capacity | Latency persisted | Monitoring log |
| 09:19 | Error rate widened | Dependency failure possible | Escalated owner | Diagnosis began | Incident channel |
| 09:31 | Recent config identified | Change may contribute | Rolled back | Recovery confirmed | Change record |
The example is structural, not a claim that incidents are linear. Parallel action and delayed effects are common. Preserve disagreement rather than forcing one smooth story.
Pass two: recover the local logic
For each consequential action, ask: What did the person believe? What information was visible? Which goal or pressure were they serving? What alternatives were practical? What signal would have changed the choice? This is not an excuse-making exercise. It is how the review finds the conditions that can recur.
Replace “operator ignored the warning” with questions about alert meaning, frequency, false positives, interface placement, training, workload and authority. Replace “they should have known” with the document, signal or rehearsal that would have made knowledge available at that moment. A person can violate a clear standard; the review should still ask why detection, prevention or recovery depended on perfection.
Pass three: map contributing factors
Do not stop at the first cause that accepts a verb. Use at least five lenses:
- Trigger: what changed or initiated the event?
- Latent conditions: what configuration, debt or exception made the trigger consequential?
- Detection: what revealed the problem, and what should have revealed it sooner?
- Response: what reduced harm, and what slowed or confused response?
- Governance: what incentives, ownership gaps, standards or resource choices shaped the conditions?
A causal diagram is optional. Multiple contributing factors are not a refusal to decide. They are often the accurate shape of a complex event. The causal-claims guide helps distinguish sequence, association and interventions that would plausibly change the outcome.
Name what worked
Record successful detection, safe rollback, clear escalation, good judgment and protective redundancy. Otherwise the review may remove a control that limited damage or teach responders that only failure is visible. “What went well” is not morale padding. It identifies resilience worth preserving.
Turn findings into changes, not wishes
| Weak action | Stronger action | Verification |
|---|---|---|
| Be more careful | Add a preflight check for the destructive parameter | Run against a safe test case |
| Improve monitoring | Alert on a defined failure condition with owner and route | Inject the condition and observe delivery |
| Update docs | Revise the exact runbook step and remove the stale path | Another person completes the scenario |
| Train the team | Run a bounded drill on the missed decision | Record time, errors and questions |
| Never happen again | Reduce likelihood or impact by a stated mechanism | Review the leading indicator on a date |
Each action needs an owner, due date, priority, expected mechanism and proof of completion. Track it in the system where ordinary work is tracked. Google SRE warns that action items without ownership and follow-through turn the postmortem into a document archive.
Keep responsibility without scapegoating
Blameless analysis assumes good intent for the learning process. It does not promise that every action was acceptable. Responsibility can include acknowledging harm, correcting a record, notifying affected people, restoring access, paying a cost, changing authority, meeting professional obligations or addressing repeated disregard of a known standard.
Separate three questions:
- Learning: which conditions produced the outcome and what will change?
- Accountability: who owns repair, follow-up and the integrity of the record?
- Conduct: was there deliberate harm, concealment, impairment, discrimination or reckless disregard requiring another process?
One meeting cannot safely answer every question. A learning review should not become a back door for discipline, and blameless language should not become a shield against legitimate conduct review.
The one-page postmortem record
- Incident name, date, scope and review trigger.
- Impact stated in observable terms.
- Detection and recovery summary.
- Evidence-linked timeline with uncertainty labels.
- Contributing factors across trigger, conditions, detection, response and governance.
- What worked and should be preserved.
- Unresolved questions and missing evidence.
- Actions with owner, deadline, mechanism and verification test.
- Separate accountability or conduct path, if required.
- Thirty-day review date and status.
Store the record where the people maintaining the system can find it. Protect sensitive personal, security and legal details. Publish a narrower version when sharing the lesson is valuable but raw evidence would create new harm.
Close the loop after thirty days
Reopen the action list. Confirm whether changes were implemented, tested and adopted. Ask whether the fix created a new failure mode or merely moved work to another person. Link recurring conditions to the exception log and update the runbook. A postmortem is complete when learning changed the operating system—not when the document received a final paragraph.
Claims and boundaries
Sourced fact: Google SRE and NIST incident-response guidance treat post-incident review, lessons learned and tracked improvement as parts of reliable operations. Inference: blameless methods improve the evidence available to a review when participants can describe local conditions without anticipating humiliation. Judgment: learning, accountability and conduct should be explicit but may require separate processes. Not claimed: intent erases harm, every failure is systemic, personnel consequences are never appropriate or this protocol replaces regulated investigation.
Official and primary practice sources
- Google SRE: Postmortem Culture—Learning from Failure, for triggers, blameless analysis, documentation and follow-up.
- Google SRE Workbook: Postmortem Culture, for facilitation, templates and action-item practice.
- NIST Incident Response project, for integrating lessons across response, recovery and improvement.
- NIST SP 800-184: Guide for Cybersecurity Event Recovery, for recovery planning and learning from prior events.
- National Transportation Safety Board investigative process, for the distinction between evidence, probable cause and safety recommendations in a regulated domain.
END OF FIELD GUIDE 084
Keep the question. Test the model.
Choose the narrowest claim the evidence can carry, then leave room for revision.