Skip to content
IncidentBot

Major incident management checklist for the first 30 minutes

The first half hour of a major incident decides most of what follows. Teams that lose it usually lose it the same way: several people debugging in parallel, nobody in charge, customers finding out from social media, and a timeline rebuilt from memory the next day. This checklist splits the first 30 minutes into four stages, with the specific actions and owners for each, so the process runs even when everyone is under pressure.

What must already exist

A checklist cannot create structure that is missing. Confirm these are in place before you rely on it.

  • A severity scale that everyone uses the same way. Our guide to incident severity levels gives a SEV1 to SEV4 starting point.
  • An on-call schedule and an escalation policy for every customer-facing service, with paging that does not depend on chat.
  • A rotation or named list of people who can act as incident commander.
  • A status page with components that match what customers recognise, and pre-approved wording for the first update.
  • Runbooks linked to services, so the first responder has a starting point for the most common failures.

Acknowledge, assess, declare

The goal of the first five minutes is to confirm the problem is real and start the incident process. Resist the urge to fix things before anyone knows an incident is running.

  • Acknowledge the page so escalation stops and others know someone is on it.
  • Check impact from the customer's side: error rates, synthetic checks, support queue. Is this affecting customers now?
  • Declare an incident and set a severity. If in doubt, choose the higher level; downgrading later is cheap.
  • Open a dedicated incident channel with a predictable name, such as #inc-2481-checkout-api-latency, and post a one-line summary.
  • Page the incident commander if the responder should stay hands-on.

Google's SRE guidance recommends declaring early for exactly this reason: the cost of an unnecessary incident is a few minutes of process, the cost of a late one is an hour of uncoordinated work.

Assign roles and set the rhythm

Once the incident is declared, structure beats speed. The roles below follow the incident command approach described in Google's Site Reliability Engineering book, adapted to typical software teams.

RoleOwnsDoes not
Incident commanderDecisions, priorities, role assignments, overall stateDebug or type commands
Operations leadHands-on investigation and changes to systemsTalk to customers or leadership
Communications leadStatus page, internal updates, support and leadership briefingsMake technical decisions
ScribeTimeline, decisions, hypotheses tested, linksChange systems
  • Incident commander announces in the channel that they have command.
  • Assign operations, communications and scribe. In a small team, one person can hold comms and scribe.
  • Set an update interval, for example every 15 or 30 minutes for a SEV1, and keep to it even when there is nothing new.
  • Move side discussions out of the channel. One channel, one thread of decisions.

Tell customers, then stabilise

Customers should hear about a major incident from you, not discover it on their own. The first update does not need a cause; it needs to show that you know and are working on it.

  • Communications lead publishes the first status page update: affected component, status (degraded performance, partial outage or major outage) and a short plain-language description.
  • Notify internal stakeholders: support, account managers, leadership. Point them to the incident channel or status page rather than answering one by one.
  • Operations lead checks recent changes first: deploys, configuration, feature flags, infrastructure changes, certificate or credential expiries.
  • Prefer mitigation over diagnosis. Roll back, fail over, shed load or disable a feature to restore service; find the root cause afterwards.
  • Commander decides whether more help is needed and pages other teams through their own escalation policies rather than direct messages.

Confirm, update, plan the next hour

By now the incident is either mitigated or clearly going to take longer. Either way, the commander resets the plan.

  • Confirm the effect of any mitigation from the customer's side, not only from internal dashboards.
  • Post the second status page update on schedule, with what changed and when the next update will come.
  • Reassess severity. Upgrade if impact grew, downgrade if mitigation holds.
  • If the incident will run long, plan handoffs: who takes over as commander and operations lead, and when. Tired responders make mistakes.
  • Scribe checks that the timeline is complete up to now: alert time, acknowledgement, declaration, each change made and its result.

Close properly

  • Post a final status page update stating that service is restored, and keep monitoring for a period agreed in advance.
  • Mark the incident resolved and record the resolution time.
  • Schedule the postmortem review within a week and assign an owner for the draft. The postmortem template lists the sections that matter.
  • Thank the people involved in the channel. It costs nothing and it matters to the next on-call shift.

Status pages and incident roles in one flow

IncidentBot runs most of this checklist as part of the incident itself. /incident in Slack opens the dedicated channel, sets severity and assigns commander, comms and scribe roles; runbooks attached to the service appear in the channel; the timeline records every step. The communications lead gets a status page update drafted from the incident and publishes it in one click to a public or private page, with subscriber notifications and a custom domain. See how customer updates work on the status page app page, compare plans on pricing, or run a sample incident on the demo to watch the first 30 minutes play out.

More guides

AlertOps pricing and AlertOps cost for teams of 5, 15 and 40

Free up to 5 users, Standard at 8 USD up to 10 users, Premium at 18 USD and Enterprise at 28 USD, the Status Hub, stakeholder and OpsIQ add-ons, the per-user message allowance and the monthly cost for 5, 15 and 40 users next to PagerDuty, xMatters and IncidentBot.

Grafana IRM pricing and Grafana OnCall pricing per active user

the free tier for 3 active users, Pro at 20 USD per active IRM user plus the 19 USD platform fee, the 25,000 USD Enterprise minimum, who counts as active and the monthly cost for 5, 15 and 40 users next to PagerDuty, incident.io, Datadog On-Call and IncidentBot.

xMatters pricing and xMatters cost per user by plan

Free, Starter at 15 USD up to 25 users, Base at 39 USD and Advanced by quote, the SMS and voice allowances, what happens at 26 users and the monthly cost for 5, 15 and 40 users next to PagerDuty, Opsgenie, AlertOps and IncidentBot.