What must already exist
A checklist cannot create structure that is missing. Confirm these are in place before you rely on it.
- A severity scale that everyone uses the same way. Our guide to incident severity levels gives a SEV1 to SEV4 starting point.
- An on-call schedule and an escalation policy for every customer-facing service, with paging that does not depend on chat.
- A rotation or named list of people who can act as incident commander.
- A status page with components that match what customers recognise, and pre-approved wording for the first update.
- Runbooks linked to services, so the first responder has a starting point for the most common failures.
Acknowledge, assess, declare
The goal of the first five minutes is to confirm the problem is real and start the incident process. Resist the urge to fix things before anyone knows an incident is running.
- Acknowledge the page so escalation stops and others know someone is on it.
- Check impact from the customer's side: error rates, synthetic checks, support queue. Is this affecting customers now?
- Declare an incident and set a severity. If in doubt, choose the higher level; downgrading later is cheap.
- Open a dedicated incident channel with a predictable name, such as #inc-2481-checkout-api-latency, and post a one-line summary.
- Page the incident commander if the responder should stay hands-on.
Google's SRE guidance recommends declaring early for exactly this reason: the cost of an unnecessary incident is a few minutes of process, the cost of a late one is an hour of uncoordinated work.
Assign roles and set the rhythm
Once the incident is declared, structure beats speed. The roles below follow the incident command approach described in Google's Site Reliability Engineering book, adapted to typical software teams.
| Role | Owns | Does not |
|---|---|---|
| Incident commander | Decisions, priorities, role assignments, overall state | Debug or type commands |
| Operations lead | Hands-on investigation and changes to systems | Talk to customers or leadership |
| Communications lead | Status page, internal updates, support and leadership briefings | Make technical decisions |
| Scribe | Timeline, decisions, hypotheses tested, links | Change systems |
- Incident commander announces in the channel that they have command.
- Assign operations, communications and scribe. In a small team, one person can hold comms and scribe.
- Set an update interval, for example every 15 or 30 minutes for a SEV1, and keep to it even when there is nothing new.
- Move side discussions out of the channel. One channel, one thread of decisions.
Tell customers, then stabilise
Customers should hear about a major incident from you, not discover it on their own. The first update does not need a cause; it needs to show that you know and are working on it.
- Communications lead publishes the first status page update: affected component, status (degraded performance, partial outage or major outage) and a short plain-language description.
- Notify internal stakeholders: support, account managers, leadership. Point them to the incident channel or status page rather than answering one by one.
- Operations lead checks recent changes first: deploys, configuration, feature flags, infrastructure changes, certificate or credential expiries.
- Prefer mitigation over diagnosis. Roll back, fail over, shed load or disable a feature to restore service; find the root cause afterwards.
- Commander decides whether more help is needed and pages other teams through their own escalation policies rather than direct messages.
Confirm, update, plan the next hour
By now the incident is either mitigated or clearly going to take longer. Either way, the commander resets the plan.
- Confirm the effect of any mitigation from the customer's side, not only from internal dashboards.
- Post the second status page update on schedule, with what changed and when the next update will come.
- Reassess severity. Upgrade if impact grew, downgrade if mitigation holds.
- If the incident will run long, plan handoffs: who takes over as commander and operations lead, and when. Tired responders make mistakes.
- Scribe checks that the timeline is complete up to now: alert time, acknowledgement, declaration, each change made and its result.
Close properly
- Post a final status page update stating that service is restored, and keep monitoring for a period agreed in advance.
- Mark the incident resolved and record the resolution time.
- Schedule the postmortem review within a week and assign an owner for the draft. The postmortem template lists the sections that matter.
- Thank the people involved in the channel. It costs nothing and it matters to the next on-call shift.
Status pages and incident roles in one flow
IncidentBot runs most of this checklist as part of the incident itself. /incident in Slack opens the dedicated channel, sets severity and assigns commander, comms and scribe roles; runbooks attached to the service appear in the channel; the timeline records every step. The communications lead gets a status page update drafted from the incident and publishes it in one click to a public or private page, with subscriber notifications and a custom domain. See how customer updates work on the status page app page, compare plans on pricing, or run a sample incident on the demo to watch the first 30 minutes play out.