Four principles behind every step
- Restore service first. Root cause analysis happens after customers are safe. The Google SRE book describes the priority during an incident as stopping the bleeding, restoring service and preserving evidence for later.
- One person coordinates. A single incident commander keeps the response coherent. Everyone else knows who decides.
- Everything is written down. Decisions and actions go into one channel and one timeline, so late joiners can catch up without interrupting.
- Blameless by default. People act on the information they had at the time. The review improves systems, not reputations.
The incident response roles
The role model below follows the Incident Command System as adapted by Google SRE. Small teams combine roles; the important part is that each responsibility has exactly one owner at any moment.
| Role | Owns | Does not |
|---|---|---|
| Incident commander (IC) | Overall coordination, severity, decisions, delegation, declaring resolution | Debug or type commands on production |
| Operations lead | Investigation and mitigation, coordinating engineers working on the fix | Post external updates |
| Communications lead | Status page, stakeholder updates, support briefings | Guess at root cause in public |
| Scribe | Timeline of events, decisions and timestamps | Make decisions |
| Subject matter experts | Specific systems, pulled in by the IC or ops lead | Start parallel changes without telling the ops lead |
Step 1. Acknowledge and assess (first 5 minutes)
- Acknowledge the page so escalation stops and others know someone is on it.
- Check the alert against user-facing signals: dashboards, error rates, support tickets.
- Decide: is this an incident? If you are unsure, declare it. A false alarm costs minutes.
Step 2. Declare and assemble (minutes 5 to 10)
- Open the incident with a service, a severity and one line of description.
- Create or join the dedicated incident channel, for example #inc-2481-checkout-api-latency.
- Name the incident commander. Until someone else takes it, the first responder is the IC.
- Page additional people through the escalation policy or directly, based on the service owner.
Step 3. Stabilise and mitigate
- The ops lead states a working hypothesis in the channel before acting on it.
- Check recent changes first: deploys, configuration changes, feature flag flips, infrastructure changes.
- Prefer reversible mitigations: rollback, disabling a feature flag, shifting traffic, scaling out.
- Announce every change to production in the channel before running it, and record the result.
Step 4. Communicate
- First internal update within 15 minutes of declaring for SEV1 and SEV2: what is affected, what is known, next update time.
- First status page update as soon as customers can notice the impact. It does not need a cause, only the symptoms and that the team is working on it.
- Keep a fixed update rhythm, such as every 30 minutes for SEV1, even when there is nothing new. Silence reads as neglect.
Step 5. Resolve
- Confirm recovery with the same signals that showed the impact, not only with the fact that a fix was deployed.
- Post a final status page update and an internal summary with the time impact ended.
- The IC declares the incident resolved and hands open questions to the postmortem owner.
Step 6. Learn
- Assign a postmortem owner and a date. For SEV1 and SEV2, schedule the review within a week.
- Draft the postmortem from the timeline: impact, detection, response, contributing factors, what went well, what went poorly.
- Create action items with owners and due dates in the issue tracker, and track them to completion.
Incident commander checklist
- Confirm severity and update it as impact becomes clearer.
- Assign ops lead, communications lead and scribe, even if one person holds two roles.
- Every 15 to 30 minutes, summarise status in the channel: impact, current hypothesis, actions in flight, next update.
- Watch for fatigue. On long incidents, hand over roles explicitly and record the handover.
- Declare resolution and name the postmortem owner.
Operations lead checklist
- Share access to the relevant dashboards and logs in the channel.
- Keep a list of hypotheses and their status: open, ruled out, confirmed.
- Coordinate changes so two engineers do not change the same system at once.
- Tell the IC when mitigation is in place and what to watch.
Communications lead checklist
- Draft the first status page update from the incident facts: component, symptom, next update time.
- Brief support with a short answer they can give to customers.
- Update internal stakeholders on the agreed schedule.
- Post the resolution notice and, for major incidents, a follow-up once the postmortem is done.
Scribe checklist
- Record timestamps for detection, declaration, key decisions, mitigations and resolution.
- Note who did what and why, in neutral language.
- Capture links to graphs, deploys and commands used.
What breaks a playbook in practice
- The hero debugger as commander. The most senior engineer fixes the problem while nobody coordinates, updates stop, and duplicate work starts.
- Side conversations in direct messages. Decisions made off the channel never reach the timeline or the postmortem.
- Waiting for the root cause before updating customers. Customers need to know that you know. The cause can wait.
- Closing too early. Resolving when the fix ships, not when recovery is confirmed, leads to reopened incidents.
For the policy behind this playbook, see our incident response plan template, and for choosing levels, the guide on incident priority levels.
Roles, channel and timeline from one command
IncidentBot runs this playbook where your team already talks. Type /incident in Slack and it opens a dedicated channel, assigns commander, comms and scribe roles, sets the SEV level, attaches the service runbook and starts a live timeline. Status page updates are drafted from the incident and published in one click, and the postmortem draft comes from the timeline. Paging does not depend on Slack, so escalation still reaches people by push, SMS and voice. See the details on the Slack incident management page, compare plans on the pricing page, or run the playbook on a scenario in the incident simulator.