Skip to content
IncidentBot

Incident response plan template for engineering teams

An incident response plan is the document your team reads before an incident, so that nobody has to invent the process during one. It should be short enough to read in ten minutes and specific enough that a new engineer on their first on-call shift knows what to do. Below is a template for software and platform teams, section by section, with notes on what to write and what to leave out.

Plan, playbook and runbook are three different documents

  • Incident response plan: the policy. Who is on call, how incidents are declared, which roles exist, how the team communicates and what happens afterwards. One per organisation or engineering department.
  • Incident response playbook: the procedure for running an incident, step by step, with checklists per role. See our incident response playbook guide.
  • Runbook: technical steps for one service or one failure mode, such as failing over a database or clearing a stuck queue. Many per service.

Keep the plan at the policy level and link out to playbooks and runbooks. A plan that tries to contain every procedure goes stale within a quarter.

Where the structure comes from

The template follows the lifecycle most frameworks share. NIST Special Publication 800-61, the widely used guide on computer security incident handling, describes preparation, detection and analysis, containment and recovery, and post-incident activity. The Google SRE book adapts the Incident Command System for production incidents, with clear roles and a single coordinator. The template below combines both for operational incidents such as outages, degradation and data problems.

Section 1. Purpose and scope

Write two or three sentences on what counts as an incident here. For example: "An incident is any unexpected event that degrades a customer-facing service, puts data at risk, or blocks the team from deploying fixes. Security incidents follow this plan plus the security addendum." List the services and teams in scope, and link to the service catalog if you have one.

Section 2. Severity levels

Define three to five levels with checkable criteria and the response each triggers. Our guides on incident severity levels and incident priority levels give ready definitions.

LevelCriteriaPagingStatus pagePostmortem
SEV1Core journey down or data at riskImmediately, day or nightRequiredRequired
SEV2Significant degradation, limited scopeImmediately, day or nightIf customers noticeRequired
SEV3Minor impact, workaround existsBusiness hoursNoOptional
SEV4Cosmetic or internal onlyNone, ticketNoNo

Section 3. Detection and declaring an incident

  • Which alert sources page a human, and which only open tickets.
  • Who can declare an incident. The best answer is anyone, including support.
  • How to declare: the command, channel or button, and the information required (service, severity, short description).
  • The rule of thumb: if you are wondering whether to declare, declare. Closing a false alarm costs little.

Section 4. On-call and escalation

  • Rotations per team, their length and handover time.
  • Escalation tiers and timeouts, such as primary for 5 minutes, then secondary for 10 minutes, then the engineering manager.
  • Paging channels by severity, such as push first, then SMS and voice for SEV1 and SEV2.
  • How overrides and swaps are recorded, so the schedule always matches reality.

Section 5. Roles and responsibilities

RoleResponsibilityWho fills it
Incident commanderCoordinates the response, makes decisions, delegates. Does not debug.On-call engineer or a trained commander
Operations leadInvestigates and applies mitigationsEngineers from the owning team
Communications leadStatus page updates, stakeholder and support updatesEngineer, support lead or product manager
ScribeKeeps the timeline: decisions, actions, timestampsAny available team member

Write down that the commander role can be handed over, and how: an explicit statement in the incident channel, recorded on the timeline.

Section 6. Communication

  • Internal: one incident channel per incident, named consistently, with a pinned summary.
  • Stakeholders: who gets updates, how often per severity, and in which format.
  • Customers: when a status page update is required, who approves the wording, and the target interval between updates.
  • Language: facts and next update time, no speculation about cause during the incident.

Section 7. Mitigation principles

State the principles, not the procedures: restore service first and find root cause later, prefer rollback over forward fixes, and ask before any action that cannot be undone. Link to runbooks per service.

Section 8. Resolution and closing

Define when an incident is resolved (customer impact over and monitoring confirms it) and who closes it. Require a final status page update and a note in the channel with the resolution time.

Section 9. Post-incident review

  • Which severities need a postmortem, and the deadline, such as within five business days.
  • The blameless principle: the review looks at systems and decisions, not at who made a mistake.
  • Action items have an owner and a due date, and live in the issue tracker, not in the document.

Section 10. Maintenance of this plan

Name an owner, set a review interval (quarterly works for most teams), and test the plan with a game day or a tabletop exercise. Record the date of the last review at the top.

Common reasons plans are ignored

  • Too long. If it takes an hour to read, nobody reads it during onboarding.
  • Not wired into the tools. A severity table in a document does nothing unless the paging and escalation rules match it.
  • Never rehearsed. The first time a new commander runs an incident should not be a real SEV1.
  • Stale contacts. Names and phone numbers in the plan go out of date. Point to schedules instead of listing people.

From document to working process

IncidentBot turns the plan into behaviour: escalation policies with tiers and timeouts, /incident in Slack opening a dedicated channel with commander, comms and scribe roles, SEV1 to SEV4 levels, runbooks attached to services, status page updates drafted from the incident and a postmortem draft built from the timeline. See how the pieces fit on the incident response tools page, compare plans on the pricing page, or run a scenario in the incident simulator to see a plan like this one execute.

More guides

Incident response playbook with roles, steps and checklists

what the commander, comms lead and scribe do from the first alert to resolution.

On call rotation with 6 schedule patterns and when to use them

weekly, split, follow-the-sun and other rotations compared by team size and time zones.

On call schedule template with sample weekly and follow-the-sun rotations

ready schedules you can adapt, with hand-off times and override rules.