Postmortem triggers worth agreeing on in advance
Decide before the next incident which events require a written postmortem, so nobody has to argue about it afterwards. Google's Site Reliability Engineering book lists a set of triggers that most teams can adopt with their own thresholds.
- User-visible downtime or degradation beyond an agreed threshold.
- Data loss of any kind.
- On-call intervention to restore service, such as a rollback or rerouting traffic.
- Resolution time above an agreed threshold.
- A monitoring failure, meaning the incident was found by a person or a customer rather than by an alert.
Anyone involved should also be able to request a postmortem for an incident that did not meet a trigger. Tie the requirement to your incident severity levels: a common rule is that every SEV1 and SEV2 gets a full postmortem, and SEV3 gets one when it reveals something new.
Incident postmortem template, section by section
Copy the outline below into your document tool. The right column says what a good entry looks like.
| Section | What to write |
|---|---|
| Title and incident ID | Short and specific: service, symptom, date. INC-2481 checkout-api latency, for instance |
| Status | Draft, in review, or final, plus the owner of the document |
| Summary | Three or four sentences a person outside the team understands: what broke, for whom, for how long, how it was fixed |
| Impact | Who was affected and how: requests failed, customers unable to pay, internal tools unavailable. Use measured numbers from your own dashboards, or say the number is unknown |
| Detection | How the incident was noticed: which alert, at what time, or who reported it if no alert fired |
| Timeline | Timestamped events from first signal to resolution, in one time zone |
| Trigger | The change or event that started the incident |
| Root causes and contributing factors | Why the trigger turned into an outage. Usually several conditions, not one |
| Resolution | What restored service, and whether it is a mitigation or a permanent fix |
| What went well | Tools, runbooks and decisions that shortened the incident |
| What went wrong | Gaps in detection, response, tooling or process |
| Where we got lucky | Things that limited the damage by chance. These are future incidents waiting to happen |
| Action items | Each with an owner, a priority and a ticket link |
Why trigger and root cause are separate
A deploy, a certificate expiry or a traffic spike is usually the trigger, not the cause. The cause is why the system could not absorb it: no canary, no alert on certificate age, no load shedding. Keeping the two apart stops the document from ending at "someone shipped a bad change", which is where learning stops too.
How to keep a blameless postmortem honest
Blameless does not mean nobody did anything. It means the document assumes people acted reasonably with the information and tools they had at the time, and asks why the system let a reasonable action cause harm. The practice comes from safety-critical industries and is described at length in Google's SRE material on postmortem culture.
- Describe actions, not character. "The deploy was run without a canary stage" rather than "the engineer skipped the canary."
- Name roles, not people, when the name adds nothing. The incident commander, the on-call engineer, the database owner.
- Record what the person knew at the time, including what the dashboards showed. Hindsight is not evidence.
- If a human error appears, the next question is always what made that error easy to make.
Teams that punish mistakes get shorter, vaguer postmortems, because people protect themselves. Teams that do not, get the detail that actually prevents the next incident.
Build the timeline from records, not memory
The timeline is the most useful section and the hardest to rebuild after the fact. Pull it from the systems that recorded the incident while it happened: alert timestamps, paging and acknowledgement times, chat messages in the incident channel, deploys and rollbacks, status page updates.
| Time (UTC) | Event |
|---|---|
| 14:02 | Latency alert fires on checkout-api |
| 14:04 | Primary on-call acknowledges |
| 14:09 | Incident declared at SEV2, incident channel opened, commander assigned |
| 14:15 | First status page update published |
| 14:31 | Rollback of the most recent deploy started |
| 14:38 | Latency back within target, monitoring continues |
| 15:10 | Incident resolved |
From these timestamps you can read time to acknowledge and time to resolve directly, which also feeds your MTTR reporting without extra work.
The review meeting and the follow-through
A postmortem that nobody reviews is a diary entry. Hold a short review within a week while memories are fresh, with the people involved and the owners of affected services.
- Circulate the draft a day before the meeting and use the meeting for questions, not for reading.
- Limit action items to the few that address causes found in the document. Twenty low-priority tickets means none get done.
- Every action item has one owner and lives in the tracker the team already uses, not only in the document.
- Review open postmortem action items in a regular engineering meeting until they are closed or explicitly dropped.
- Share finished postmortems widely. Other teams often have the same weakness.
From incident to postmortem draft without rebuilding the timeline
IncidentBot captures the timeline while the incident runs: alerts, pages and acknowledgements, severity changes, role assignments, key messages from the Slack incident channel and status page updates. When the incident is resolved, it drafts a blameless postmortem with this structure, the timeline already filled in and time to acknowledge and time to resolve calculated. Action items sync to Jira and Linear so they are tracked where your team works. See the details on the incident postmortem page, compare plans on pricing, or run a sample incident on the demo and read the postmortem skeleton it produces.