Most postmortems fail after the meeting, not in it. The document is written and the actions are agreed, then three months later the same alert fires because the fix is still sitting in a backlog with nobody’s name on it.
A postmortem turns one bad day into changes that make the next one less likely. This free incident postmortem template is for SRE teams, platform and infrastructure engineers, IT operations managers and MSPs who run reviews after significant incidents and want every one to end with finished work rather than a document. It starts with a decision on whether the incident needs a review at all, then walks the owner through freezing the evidence, rebuilding the timeline, drafting the write-up, running a facilitated blameless meeting and publishing the result. The last two phases are the ones most teams skip: tracking every action to a verified close, with a hand-off to change management or problem management when the fix needs one.
Google’s Site Reliability Engineering book describes a postmortem as a written record of an incident, its impact, the actions taken to mitigate or resolve it, its causes and the follow-up actions to stop it happening again. A blameless postmortem assumes that everyone involved acted in good faith and made reasonable decisions with the information they had at the time. The review looks for the conditions that made a mistake possible, because a system can be changed and a person’s fallibility cannot. It is not a way of avoiding accountability. People who expect to be blamed leave details out, and the review loses the facts it needs most.
The same book recommends agreeing postmortem criteria before an incident happens, so nobody has to argue about whether one is needed. Its example triggers are user-visible downtime or degradation beyond a threshold, any data loss, on-call intervention such as a rollback or rerouting traffic, a resolution time above a threshold, and a monitoring failure that meant the incident was found by hand. Any stakeholder can also ask for one.
This template picks up where live response ends. Phase 6 of the Incident Management Checklist is a short post-incident review held while the team is still close to the event. This is the full version: the written review, the facilitated meeting and the weeks of action tracking that follow. It covers one incident. When the review finds an underlying cause that also sits behind other incidents, it hands that cause to problem management, which traces it across every incident it produces.
Blame-oriented review
Asks who got it wrong
Typical line: “The engineer ran the migration against production without checking.”
Ends with: a name, a warning and a reminder to be more careful.
Next incident: responders share less, and the same gap is still there.
Blameless postmortem
Asks what allowed it to happen
Typical line: “The migration tool defaulted to the production connection, and nothing prompted for confirmation.”
Ends with: owned actions such as a confirmation prompt, a safer default and a dry-run mode.
Next incident: the gap is closed for everyone, not just the person involved.
What the Incident Postmortem Template Covers
Seven phases take one incident from the review decision to closed actions. Answers on the checklist decide whether the review runs at all, whether customers get a report and whether change or problem records are raised.
Phase 1
Phase 1: Decide & Assign
If no trigger applies, task 5 is the only task shown and the rest of the checklist is hidden.
Check the incident against your postmortem triggers — severity, data loss, customer impact, a rollback or failover, a monitoring miss, or a stakeholder request
Assign one postmortem owner — often the incident commander or the engineer closest to the fix; the owner writes the draft and chases the actions
Set dates for the draft and the review meeting — for example a draft within five working days and the meeting within ten
Freeze the evidence before retention deletes it — export the incident channel, bridge notes, alert history, dashboards and deployment logs
Record why no postmortem is needed — which triggers were checked and who decided, so the choice can be questioned later
Phase 2
Phase 2: Rebuild the Timeline
Rebuild the timeline from timestamps, not memory — alerts, pages, chat messages, commands and deploys, all in UTC
Mark the key moments — start of impact, detection, first response, mitigation and full recovery; the gaps between them are your time to detect and time to mitigate
Talk to each responder one to one — ask what they saw and what they expected at each decision point, before the group meeting fixes a shared story
Note what responders knew at each decision — a blameless review judges a decision against the information available then, not hindsight
Quantify the impact — users or customers affected, failed requests, data lost or delayed, error budget consumed and any SLA breach
Phase 3
Phase 3: Draft the Postmortem
Write a summary a non-engineer can follow — what broke, who noticed, how long it lasted and what fixed it, in four sentences or fewer
List the contributing factors, not a single culprit — the trigger, the latent conditions it met, and gaps in monitoring, runbooks, tooling or process
Record what went well and where you got lucky — luck that shortened the outage is a risk that will not repeat itself
Draft actions against each contributing factor — specific and testable, such as “alert when queue depth exceeds 10,000 for five minutes”, not “improve monitoring”
Check the draft for blame — replace names with roles, and “should have” with what the system allowed or failed to show
Phase 4
Phase 4: Facilitated Review Meeting
The facilitator should be someone who was not in the incident and does not manage the people who were.
Circulate the draft a day before the meeting — attendees comment in advance, so the meeting spends its time on disagreements
Walk the timeline and challenge the gaps — is the impact complete, and does the analysis go deep enough to explain why the defences failed?
Agree an owner, priority and due date for every action — no action leaves the room without a named person
Decide whether an underlying cause needs a problem record — if the cause is unconfirmed or also sits behind other incidents, the answer is yes
Decide whether any action changes production — those actions go through change management rather than straight into a backlog
Phase 5
Phase 5: Publish & Share
Task 4 appears only when the checklist records that customers were affected.
Get sign-off on the final document — the service owner or engineering lead confirms it is complete and accurate
Publish it to your postmortem library — tagged by service, date and contributing factor, so later reviews can search for patterns
Share a summary with the wider engineering team — the people who run similar systems learn most from it
Send the customer incident report — plain language, no internal names; check contracts for notice deadlines and content requirements
Phase 6
Phase 6: Track Actions & Hand Off
Task 2 appears when an action changes production. Task 3 appears when the review asked for a problem record.
Log every action in the team’s work tracker and link it here — the postmortem stays open until each one is done or formally declined
Raise a change record for each action that alters production — the postmortem is the justification, and the change follows your normal approval route
Raise a problem record for the underlying cause — link the incident and this postmortem so the investigation starts with the evidence
Review open actions with their owners every week — anything past its due date is escalated to the service owner
Record a reason for every declined or deferred action — who decided, why, and when it will be looked at again
Phase 7
Phase 7: Verify & Close
This phase is due when the last action’s due date passes, not when the meeting ends.
Confirm each action is complete, with evidence — a merged change, a new alert firing in test, an updated runbook
Test that the fixes work — a failure-injection test, a game day, or a tabletop replay of this incident with a new on-call engineer
Update the runbooks and alert descriptions this incident exposed — the next responder should not need to read the postmortem to act
Close the postmortem with owner sign-off — timeline, document, actions and evidence stay attached as the record
The checklist runs the process. The document is what people read later, often months later, when a similar alert fires. These are the sections Phase 3 asks the owner to write, with the mistake that most often weakens each one.
Section
What goes in it
Common mistake
Summary
What broke, for how long, who noticed and what fixed it
Written for the team that was there, so nobody else can follow it
Impact
Users, customers, requests, data and SLAs affected, with numbers
“Some users were affected” with no figures behind it
Timeline
Timestamped events from the first signal to full recovery
Rebuilt from memory, so the times drift and gaps disappear
Contributing factors
The trigger and the conditions that let it cause an outage
Stopping at the trigger, such as “bad deploy”, and ignoring why it reached production
What went well
Detection, tools and decisions that shortened the incident
Left out, so good practice is never reinforced
Where we got lucky
Things that helped by chance, such as low traffic or the right person being online
Counted as resilience when it was chance
Action items
Specific changes, each with an owner, priority, due date and tracker link
Vague actions with no owner, never checked again
Why Run Postmortems in CheckFlow?
1
The review starts when the incident closes
A webhook or the REST API can start the postmortem checklist when a qualifying incident is resolved in your incident tool, or you can connect it through Zapier. The owner is assigned straight away and the evidence-freeze task is due that day.
2
Actions with dates that cannot drift
Dynamic due dates are set from the review date, and each action is assigned to a named person. Actions are captured in a table inside the task, with owner, priority and tracker link, and the weekly review task keeps them visible until every row is closed.
3
Only the steps this incident needs
Conditional logic hides the customer report when no customer was affected and adds change and problem hand-offs only when the review calls for them. Chat exports and timelines are attached to the task they support, and the activity trail records who signed off and when.
Many postmortem actions end in a runbook. Our IT runbook template guide covers how to write one that responders actually follow, and Phase 7 above makes sure the runbooks an incident exposed are updated before the postmortem closes.
It is a written review of an incident that looks for the systemic causes of what went wrong rather than a person to hold responsible. It assumes everyone acted in good faith with the information they had, and asks why the system let a reasonable action cause an outage. The practice is closely associated with Google’s Site Reliability Engineering teams. The point is practical: people who fear blame hold back the details a review needs to find the real gaps.
When should you write an incident postmortem?
+
Whenever an incident meets triggers you agreed in advance. Typical triggers are user-visible downtime or degradation beyond a threshold, any data loss, on-call intervention such as a rollback, a resolution time above a threshold, and a monitoring failure. You can also simply require one for every SEV-1 and SEV-2. Anyone should be able to request a postmortem for an incident below the line, and a review of a near miss is often cheaper than a review of the outage it predicted.
What should an incident postmortem include?
+
A summary, the impact in numbers, a timestamped timeline, the contributing factors, what went well, where the team got lucky, and action items with an owner, priority and due date each. The table above describes each section and the mistake that most often weakens it. Keep the document factual and written in terms of roles and systems rather than named individuals.
How soon after an incident should the postmortem happen?
+
Freeze the evidence the same day, because chat history and logs expire and memories fade fast. The draft and meeting can follow within a week or two, once the team has recovered and the timeline has been rebuilt properly. Set the targets in your policy and track them, so that a busy week does not quietly turn into a review that never happens.
What is the difference between a postmortem and root cause analysis?
+
A postmortem reviews one incident end to end: what happened, how the response went and what should change. Root cause analysis is one technique used inside it, and complex incidents rarely have a single root cause anyway. When the underlying cause is unconfirmed or links several incidents, the investigation moves into a problem record, where it can run for weeks without holding up the postmortem.
How do you make sure postmortem action items get done?
+
Give every action one named owner, a priority and a due date before the meeting ends, and keep the postmortem open until every action is complete or formally declined with a reason. Review open actions weekly, escalate overdue ones to the service owner, and check the fix works before closing. A postmortem with open actions is still an open risk.
Is CheckFlow free for this template?
+
14-day free trial, no card required. The Business plan is $10 per user per month after the trial. Full details at checkflow.io/pricing.
Every Postmortem Action Owned, Tracked and Closed
Free trial — no credit card required.
Do you like cookies? 🍪 We use cookies to ensure you get the best experience on our website. Learn more