Skip to content
IncidentBot

Postmortem template in a blameless format your team will actually fill in

Most postmortem templates fail for the same reason: they are long, they ask for things nobody remembers a week later, and the document ends up half written in a folder nobody opens. This template keeps the sections that produce learning, explains what goes in each one, and adds the habits that turn a document into fixed systems.

Postmortem triggers worth agreeing on in advance

Decide before the next incident which events require a written postmortem, so nobody has to argue about it afterwards. Google's Site Reliability Engineering book lists a set of triggers that most teams can adopt with their own thresholds.

  • User-visible downtime or degradation beyond an agreed threshold.
  • Data loss of any kind.
  • On-call intervention to restore service, such as a rollback or rerouting traffic.
  • Resolution time above an agreed threshold.
  • A monitoring failure, meaning the incident was found by a person or a customer rather than by an alert.

Anyone involved should also be able to request a postmortem for an incident that did not meet a trigger. Tie the requirement to your incident severity levels: a common rule is that every SEV1 and SEV2 gets a full postmortem, and SEV3 gets one when it reveals something new.

Incident postmortem template, section by section

Copy the outline below into your document tool. The right column says what a good entry looks like.

SectionWhat to write
Title and incident IDShort and specific: service, symptom, date. INC-2481 checkout-api latency, for instance
StatusDraft, in review, or final, plus the owner of the document
SummaryThree or four sentences a person outside the team understands: what broke, for whom, for how long, how it was fixed
ImpactWho was affected and how: requests failed, customers unable to pay, internal tools unavailable. Use measured numbers from your own dashboards, or say the number is unknown
DetectionHow the incident was noticed: which alert, at what time, or who reported it if no alert fired
TimelineTimestamped events from first signal to resolution, in one time zone
TriggerThe change or event that started the incident
Root causes and contributing factorsWhy the trigger turned into an outage. Usually several conditions, not one
ResolutionWhat restored service, and whether it is a mitigation or a permanent fix
What went wellTools, runbooks and decisions that shortened the incident
What went wrongGaps in detection, response, tooling or process
Where we got luckyThings that limited the damage by chance. These are future incidents waiting to happen
Action itemsEach with an owner, a priority and a ticket link

Why trigger and root cause are separate

A deploy, a certificate expiry or a traffic spike is usually the trigger, not the cause. The cause is why the system could not absorb it: no canary, no alert on certificate age, no load shedding. Keeping the two apart stops the document from ending at "someone shipped a bad change", which is where learning stops too.

How to keep a blameless postmortem honest

Blameless does not mean nobody did anything. It means the document assumes people acted reasonably with the information and tools they had at the time, and asks why the system let a reasonable action cause harm. The practice comes from safety-critical industries and is described at length in Google's SRE material on postmortem culture.

  • Describe actions, not character. "The deploy was run without a canary stage" rather than "the engineer skipped the canary."
  • Name roles, not people, when the name adds nothing. The incident commander, the on-call engineer, the database owner.
  • Record what the person knew at the time, including what the dashboards showed. Hindsight is not evidence.
  • If a human error appears, the next question is always what made that error easy to make.

Teams that punish mistakes get shorter, vaguer postmortems, because people protect themselves. Teams that do not, get the detail that actually prevents the next incident.

Build the timeline from records, not memory

The timeline is the most useful section and the hardest to rebuild after the fact. Pull it from the systems that recorded the incident while it happened: alert timestamps, paging and acknowledgement times, chat messages in the incident channel, deploys and rollbacks, status page updates.

Time (UTC)Event
14:02Latency alert fires on checkout-api
14:04Primary on-call acknowledges
14:09Incident declared at SEV2, incident channel opened, commander assigned
14:15First status page update published
14:31Rollback of the most recent deploy started
14:38Latency back within target, monitoring continues
15:10Incident resolved

From these timestamps you can read time to acknowledge and time to resolve directly, which also feeds your MTTR reporting without extra work.

The review meeting and the follow-through

A postmortem that nobody reviews is a diary entry. Hold a short review within a week while memories are fresh, with the people involved and the owners of affected services.

  • Circulate the draft a day before the meeting and use the meeting for questions, not for reading.
  • Limit action items to the few that address causes found in the document. Twenty low-priority tickets means none get done.
  • Every action item has one owner and lives in the tracker the team already uses, not only in the document.
  • Review open postmortem action items in a regular engineering meeting until they are closed or explicitly dropped.
  • Share finished postmortems widely. Other teams often have the same weakness.

From incident to postmortem draft without rebuilding the timeline

IncidentBot captures the timeline while the incident runs: alerts, pages and acknowledgements, severity changes, role assignments, key messages from the Slack incident channel and status page updates. When the incident is resolved, it drafts a blameless postmortem with this structure, the timeline already filled in and time to acknowledge and time to resolve calculated. Action items sync to Jira and Linear so they are tracked where your team works. See the details on the incident postmortem page, compare plans on pricing, or run a sample incident on the demo and read the postmortem skeleton it produces.

More guides

Escalation policy design with tiers, timeouts and fallbacks

how to choose tiers and timeouts so every alert reaches someone and nobody is woken without need.

Incident severity levels from SEV1 to SEV4 with definitions and examples

clear definitions with examples, and what each level triggers in paging and communication.

Major incident management checklist for the first 30 minutes

a minute-by-minute checklist for declaring, staffing and communicating a major incident.