Plan, playbook and runbook are three different documents
- Incident response plan: the policy. Who is on call, how incidents are declared, which roles exist, how the team communicates and what happens afterwards. One per organisation or engineering department.
- Incident response playbook: the procedure for running an incident, step by step, with checklists per role. See our incident response playbook guide.
- Runbook: technical steps for one service or one failure mode, such as failing over a database or clearing a stuck queue. Many per service.
Keep the plan at the policy level and link out to playbooks and runbooks. A plan that tries to contain every procedure goes stale within a quarter.
Where the structure comes from
The template follows the lifecycle most frameworks share. NIST Special Publication 800-61, the widely used guide on computer security incident handling, describes preparation, detection and analysis, containment and recovery, and post-incident activity. The Google SRE book adapts the Incident Command System for production incidents, with clear roles and a single coordinator. The template below combines both for operational incidents such as outages, degradation and data problems.
Section 1. Purpose and scope
Write two or three sentences on what counts as an incident here. For example: "An incident is any unexpected event that degrades a customer-facing service, puts data at risk, or blocks the team from deploying fixes. Security incidents follow this plan plus the security addendum." List the services and teams in scope, and link to the service catalog if you have one.
Section 2. Severity levels
Define three to five levels with checkable criteria and the response each triggers. Our guides on incident severity levels and incident priority levels give ready definitions.
| Level | Criteria | Paging | Status page | Postmortem |
|---|---|---|---|---|
| SEV1 | Core journey down or data at risk | Immediately, day or night | Required | Required |
| SEV2 | Significant degradation, limited scope | Immediately, day or night | If customers notice | Required |
| SEV3 | Minor impact, workaround exists | Business hours | No | Optional |
| SEV4 | Cosmetic or internal only | None, ticket | No | No |
Section 3. Detection and declaring an incident
- Which alert sources page a human, and which only open tickets.
- Who can declare an incident. The best answer is anyone, including support.
- How to declare: the command, channel or button, and the information required (service, severity, short description).
- The rule of thumb: if you are wondering whether to declare, declare. Closing a false alarm costs little.
Section 4. On-call and escalation
- Rotations per team, their length and handover time.
- Escalation tiers and timeouts, such as primary for 5 minutes, then secondary for 10 minutes, then the engineering manager.
- Paging channels by severity, such as push first, then SMS and voice for SEV1 and SEV2.
- How overrides and swaps are recorded, so the schedule always matches reality.
Section 5. Roles and responsibilities
| Role | Responsibility | Who fills it |
|---|---|---|
| Incident commander | Coordinates the response, makes decisions, delegates. Does not debug. | On-call engineer or a trained commander |
| Operations lead | Investigates and applies mitigations | Engineers from the owning team |
| Communications lead | Status page updates, stakeholder and support updates | Engineer, support lead or product manager |
| Scribe | Keeps the timeline: decisions, actions, timestamps | Any available team member |
Write down that the commander role can be handed over, and how: an explicit statement in the incident channel, recorded on the timeline.
Section 6. Communication
- Internal: one incident channel per incident, named consistently, with a pinned summary.
- Stakeholders: who gets updates, how often per severity, and in which format.
- Customers: when a status page update is required, who approves the wording, and the target interval between updates.
- Language: facts and next update time, no speculation about cause during the incident.
Section 7. Mitigation principles
State the principles, not the procedures: restore service first and find root cause later, prefer rollback over forward fixes, and ask before any action that cannot be undone. Link to runbooks per service.
Section 8. Resolution and closing
Define when an incident is resolved (customer impact over and monitoring confirms it) and who closes it. Require a final status page update and a note in the channel with the resolution time.
Section 9. Post-incident review
- Which severities need a postmortem, and the deadline, such as within five business days.
- The blameless principle: the review looks at systems and decisions, not at who made a mistake.
- Action items have an owner and a due date, and live in the issue tracker, not in the document.
Section 10. Maintenance of this plan
Name an owner, set a review interval (quarterly works for most teams), and test the plan with a game day or a tabletop exercise. Record the date of the last review at the top.
Common reasons plans are ignored
- Too long. If it takes an hour to read, nobody reads it during onboarding.
- Not wired into the tools. A severity table in a document does nothing unless the paging and escalation rules match it.
- Never rehearsed. The first time a new commander runs an incident should not be a real SEV1.
- Stale contacts. Names and phone numbers in the plan go out of date. Point to schedules instead of listing people.
From document to working process
IncidentBot turns the plan into behaviour: escalation policies with tiers and timeouts, /incident in Slack opening a dedicated channel with commander, comms and scribe roles, SEV1 to SEV4 levels, runbooks attached to services, status page updates drafted from the incident and a postmortem draft built from the timeline. See how the pieces fit on the incident response tools page, compare plans on the pricing page, or run a scenario in the incident simulator to see a plan like this one execute.