Severity measures impact, priority measures order of work
The two words are often used interchangeably, and that causes confusion. Severity describes how bad the impact is right now: how many people are affected, and how badly. Priority describes how urgently the work should be done relative to other work, which can depend on things beyond impact, such as a contractual deadline or a launch. Many teams run incidents on severity alone and use priority for tickets. If you need both, our incident priority levels guide shows a P1 to P4 matrix built from impact and urgency.
SEV1 to SEV4 definitions
The table below is a starting point. Replace the examples with your own services and write your thresholds in terms your team already measures.
| Level | Definition | Examples |
|---|---|---|
| SEV1, critical | Core functionality unavailable or data at risk for all or most customers | Checkout fails for everyone, login down, data loss or corruption, confirmed security breach |
| SEV2, major | Significant degradation or a core feature down for a subset of customers, no acceptable workaround | Payments failing in one region, API error rate well above normal, severe latency on a key path |
| SEV3, minor | Limited impact, a workaround exists, or a non-core feature is affected | Report exports delayed, one integration failing, an internal tool unavailable |
| SEV4, low | No current customer impact, but something needs attention | A redundant component down, disk usage trending towards a limit, a cosmetic defect |
Keep the definitions about impact, not cause. "Database is down" is a cause. "Customers cannot place orders" is an impact, and it is what everyone outside the engineering team understands.
What each severity triggers
The level is only useful if it triggers a known response. Write these expectations down next to the definitions so that declaring a SEV1 automatically answers most of the process questions.
| Level | Paging | Coordination | Communication | Follow-up |
|---|---|---|---|---|
| SEV1 | Page immediately, any hour, escalate fast | Incident channel, dedicated incident commander, comms and scribe roles | Status page update early, then at a fixed interval; leadership informed | Postmortem required |
| SEV2 | Page immediately, any hour | Incident channel, incident commander | Status page update if customers notice | Postmortem required |
| SEV3 | Business hours, or page if it worsens | Handled by on-call, channel optional | Internal update | Postmortem if it reveals something new |
| SEV4 | No page, ticket | Normal work queue | None | None |
Paging rules per level are the bridge between severity and your escalation policy: SEV1 and SEV2 use short timeouts and voice calls, SEV3 waits for working hours, SEV4 never pages.
Rules that keep the scale consistent
- When in doubt, pick the higher level. It is cheap to downgrade a SEV2 to a SEV3 after ten minutes and expensive to discover an hour in that a SEV3 was really a SEV1. Google's SRE guidance on incident management makes the same point about declaring incidents early rather than late.
- Anyone can declare, the incident commander decides. The first responder sets a level so the process starts; the commander confirms or changes it.
- Severity changes are recorded with a reason. Upgrades and downgrades belong on the incident timeline, because they explain why the response changed.
- Severity reflects current impact. If a SEV1 is mitigated but the root cause is not fixed, the incident can move to SEV3 while engineers continue work.
- Security incidents follow the same scale but add their own confidentiality rules, such as a private channel and a restricted list of participants.
Criteria for declaring an incident at all
Google's SRE book suggests a simple test for whether an issue should become a formal incident: it needs a second team involved, it is visible to customers, or it remains unsolved after about an hour of focused work. Any one of these is enough. Use it alongside your severity scale so small issues stay in the normal queue and real incidents get structure quickly.
Why severity scales drift
- Too many levels. Five or six levels invite arguments about the difference between neighbours. Four is enough for most teams.
- Definitions that depend on the service. Keep one scale for the whole organisation and add service-specific examples, so a SEV2 means the same thing in every team.
- Severity used to get attention. If teams inflate severity to get help faster, the underlying problem is response time on lower levels, and that is what to fix.
- No link to follow-up. If a SEV1 does not reliably produce a postmortem, the scale is only a label. Tie each level to the postmortem template your team uses.
Review the scale twice a year with a few recent incidents on the table: were they classified the same way by different people? If not, the definitions need sharper examples.
Severity levels built into incident response
In IncidentBot, every incident carries a severity from SEV1 to SEV4, set when it is declared from an alert or with /incident in Slack and changed from the incident channel. Severity drives what happens next: who is paged and on which channel, which roles are assigned (commander, comms, scribe), and which status the status page draft proposes, from degraded performance to major outage. Every change lands on the timeline and flows into the postmortem draft. See the full workflow on the incident response tools page, compare plans on pricing, or run a sample incident on the demo and pick a severity to watch the response change.