Skip to content
IncidentBot

Incident severity levels from SEV1 to SEV4 with definitions and examples

Severity levels exist so that nobody has to decide from scratch, in the middle of an outage, how loud to be. A good scale tells the on-call engineer within a minute whether to wake people, open an incident channel, post on the status page and write a postmortem. This guide gives a four-level scale with definitions, examples and response expectations, then covers the rules that keep it consistent.

Severity measures impact, priority measures order of work

The two words are often used interchangeably, and that causes confusion. Severity describes how bad the impact is right now: how many people are affected, and how badly. Priority describes how urgently the work should be done relative to other work, which can depend on things beyond impact, such as a contractual deadline or a launch. Many teams run incidents on severity alone and use priority for tickets. If you need both, our incident priority levels guide shows a P1 to P4 matrix built from impact and urgency.

SEV1 to SEV4 definitions

The table below is a starting point. Replace the examples with your own services and write your thresholds in terms your team already measures.

LevelDefinitionExamples
SEV1, criticalCore functionality unavailable or data at risk for all or most customersCheckout fails for everyone, login down, data loss or corruption, confirmed security breach
SEV2, majorSignificant degradation or a core feature down for a subset of customers, no acceptable workaroundPayments failing in one region, API error rate well above normal, severe latency on a key path
SEV3, minorLimited impact, a workaround exists, or a non-core feature is affectedReport exports delayed, one integration failing, an internal tool unavailable
SEV4, lowNo current customer impact, but something needs attentionA redundant component down, disk usage trending towards a limit, a cosmetic defect

Keep the definitions about impact, not cause. "Database is down" is a cause. "Customers cannot place orders" is an impact, and it is what everyone outside the engineering team understands.

What each severity triggers

The level is only useful if it triggers a known response. Write these expectations down next to the definitions so that declaring a SEV1 automatically answers most of the process questions.

LevelPagingCoordinationCommunicationFollow-up
SEV1Page immediately, any hour, escalate fastIncident channel, dedicated incident commander, comms and scribe rolesStatus page update early, then at a fixed interval; leadership informedPostmortem required
SEV2Page immediately, any hourIncident channel, incident commanderStatus page update if customers noticePostmortem required
SEV3Business hours, or page if it worsensHandled by on-call, channel optionalInternal updatePostmortem if it reveals something new
SEV4No page, ticketNormal work queueNoneNone

Paging rules per level are the bridge between severity and your escalation policy: SEV1 and SEV2 use short timeouts and voice calls, SEV3 waits for working hours, SEV4 never pages.

Rules that keep the scale consistent

  • When in doubt, pick the higher level. It is cheap to downgrade a SEV2 to a SEV3 after ten minutes and expensive to discover an hour in that a SEV3 was really a SEV1. Google's SRE guidance on incident management makes the same point about declaring incidents early rather than late.
  • Anyone can declare, the incident commander decides. The first responder sets a level so the process starts; the commander confirms or changes it.
  • Severity changes are recorded with a reason. Upgrades and downgrades belong on the incident timeline, because they explain why the response changed.
  • Severity reflects current impact. If a SEV1 is mitigated but the root cause is not fixed, the incident can move to SEV3 while engineers continue work.
  • Security incidents follow the same scale but add their own confidentiality rules, such as a private channel and a restricted list of participants.

Criteria for declaring an incident at all

Google's SRE book suggests a simple test for whether an issue should become a formal incident: it needs a second team involved, it is visible to customers, or it remains unsolved after about an hour of focused work. Any one of these is enough. Use it alongside your severity scale so small issues stay in the normal queue and real incidents get structure quickly.

Why severity scales drift

  • Too many levels. Five or six levels invite arguments about the difference between neighbours. Four is enough for most teams.
  • Definitions that depend on the service. Keep one scale for the whole organisation and add service-specific examples, so a SEV2 means the same thing in every team.
  • Severity used to get attention. If teams inflate severity to get help faster, the underlying problem is response time on lower levels, and that is what to fix.
  • No link to follow-up. If a SEV1 does not reliably produce a postmortem, the scale is only a label. Tie each level to the postmortem template your team uses.

Review the scale twice a year with a few recent incidents on the table: were they classified the same way by different people? If not, the definitions need sharper examples.

Severity levels built into incident response

In IncidentBot, every incident carries a severity from SEV1 to SEV4, set when it is declared from an alert or with /incident in Slack and changed from the incident channel. Severity drives what happens next: who is paged and on which channel, which roles are assigned (commander, comms, scribe), and which status the status page draft proposes, from degraded performance to major outage. Every change lands on the timeline and flows into the postmortem draft. See the full workflow on the incident response tools page, compare plans on pricing, or run a sample incident on the demo and pick a severity to watch the response change.

More guides

Major incident management checklist for the first 30 minutes

a minute-by-minute checklist for declaring, staffing and communicating a major incident.

AlertOps pricing and AlertOps cost for teams of 5, 15 and 40

Free up to 5 users, Standard at 8 USD up to 10 users, Premium at 18 USD and Enterprise at 28 USD, the Status Hub, stakeholder and OpsIQ add-ons, the per-user message allowance and the monthly cost for 5, 15 and 40 users next to PagerDuty, xMatters and IncidentBot.

Grafana IRM pricing and Grafana OnCall pricing per active user

the free tier for 3 active users, Pro at 20 USD per active IRM user plus the 19 USD platform fee, the 25,000 USD Enterprise minimum, who counts as active and the monthly cost for 5, 15 and 40 users next to PagerDuty, incident.io, Datadog On-Call and IncidentBot.