Timeline
Every step of the incident with computed timestamps: alert received, duplicates grouped, incident opened at the chosen severity, each person paged and on which channel (push first, then SMS, then voice depending on severity and time of day), each escalation, the acknowledgement, the Slack channel created and the roles assigned.
- Alert received
- Deduplicated
- Incident opened
- Paged
- Escalated
- Acknowledged
- Slack channel created
- Roles assigned
Status page
A draft update your customers would see: the affected component, a status derived from severity and a short customer-facing message.
- Affected component
- Degraded performance
- Partial outage
- Major outage
- Investigating
Publish update
Postmortem
The skeleton of the review: a timeline table, the impact window, time to acknowledge and time to resolve from the run, and the sections your team completes.
- Summary
- Impact
- Timeline
- Contributing factors
- What went well
- What could be improved
- Action items
Metrics
- Time to acknowledge
- People woken up
- Escalations used
- Policy check
When the policy is weak, the metrics strip says why. Typical hints: a single tier with no fallback at night, a timeout longer than the severity allows, or a policy that wakes the whole team before a lead has seen the alert.