Skip to content
IncidentBot

Escalation policy design with tiers, timeouts and fallbacks

An escalation policy decides what happens when the first person paged does not respond. It is the difference between an alert that is handled in five minutes and one that sits unacknowledged until a customer complains. This guide covers the parts of a policy, sensible timeouts per severity, the fallbacks that close the gaps, and the mistakes that show up in almost every first draft.

The parts of an on-call escalation policy

A policy is an ordered list of tiers. Each tier names who is paged and how long the system waits for an acknowledgement before moving on. Around the tiers sit a few rules that decide what happens at the end of the list.

PartWhat it definesTypical choice
Tier 1 targetWho is paged firstThe primary layer of the service's on-call schedule
Tier 1 timeoutWait before escalating5 minutes for urgent alerts
Tier 2 targetWho is paged nextThe secondary layer of the same schedule
Tier 3 targetLast human lineTeam lead or engineering manager on call
Repeat ruleWhat happens after the last tierRestart from tier 1, one or two more times
ChannelsHow each person is reachedPush first, then SMS and voice for urgent alerts

Point tiers at schedule layers rather than at named people. A policy that says "page Priya" breaks the first time Priya is on holiday. A policy that says "page the primary on-call for payments" follows the on call schedule template your team already maintains, overrides included.

Timeouts per urgency

The timeout is a trade-off. Too short and the secondary is woken for alerts the primary was about to acknowledge. Too long and a real outage waits while a phone vibrates in another room. The numbers below are starting points to tune against your own acknowledgement data, not a standard.

UrgencyTier 1 timeoutTier 2 timeoutChannels
SEV1, customer-facing outage5 minutes5 minutesPush, SMS and voice at once, then voice on escalation
SEV2, major degradation10 minutes10 minutesPush, then SMS, then voice
SEV3, limited impact30 minutes, business hours onlyNext business dayPush and email
SEV4, informationalNo pagingNo pagingTicket or channel message

Two rules help. First, low urgency alerts should not page at night at all; they wait for working hours or become tickets. Second, if people regularly acknowledge just before the timeout, the timeout is too short or the alert is too noisy. Check your time to acknowledge figures before changing either one. The severity definitions behind this table are in our guide to incident severity levels.

Sample escalation policy for a customer-facing service

Here is a complete policy for a checkout service, written as it would appear in a runbook.

  • Tier 1: primary on-call for checkout. Push immediately, SMS after 1 minute, voice call after 3 minutes. Escalate after 5 minutes without acknowledgement.
  • Tier 2: secondary on-call for checkout. Same channel sequence. Escalate after 5 minutes.
  • Tier 3: engineering manager on call for the commerce group, plus the whole checkout team by push. Escalate after 10 minutes.
  • Repeat: restart from tier 1 twice. After the second full cycle, page the incident commander rotation.
  • Acknowledgement stops escalation. Reassignment or explicit escalation by the responder moves the incident to the chosen tier immediately.

The last line matters. An engineer who acknowledges and then realises the problem belongs to the database team should be able to escalate or reassign in one step, rather than resolving the page and hoping someone else picks it up.

Fallbacks that close the gaps

Every policy has holes. These are the ones that cause missed pages most often.

  • An empty schedule layer. If nobody is assigned for a time window, the tier has no target. The policy should skip to the next tier rather than wait out the timeout on nobody.
  • Out-of-office without an override. Holidays belong in the schedule as overrides, entered before the person leaves.
  • Paging that depends on chat. If Slack or Teams is part of the incident, a chat-only page never arrives. Keep push, SMS and voice in the chain.
  • Notification rules on the phone. A primary with do-not-disturb and no exception for the paging app is effectively not on call. Ask new responders to send a test page to themselves during onboarding.
  • The end of the list. A policy that stops after the last tier leaves the incident unowned. Repeat, then fall back to a rotation that is always staffed.

Escalation policy mistakes to review for

  • Paging a whole team at tier 1. Everyone assumes someone else has it, and acknowledgement gets slower, not faster.
  • Escalating straight to managers. Managers belong at the end, as a signal that the process did not work, not as the second responder.
  • One policy for every service. A batch job and a payment API do not need the same urgency.
  • No review. Look at escalations every month: which policies escalated past tier 1, and why. Frequent tier 2 pages usually point to a noisy alert or an overloaded primary, both fixable.

If your policies escalate often at night, the root cause is usually alerting, not people. Grouping, deduplication and moving non-urgent alerts to working hours reduce pages more than any timeout change, and the on call rotation guide covers how rotation length affects fatigue.

Build and test policies in on-call management

In IncidentBot, escalation policies have tiers with their own targets and timeouts, point at schedule layers, repeat a set number of times and page by mobile push, SMS, voice, email and Slack, so paging keeps working when chat is down. Responders acknowledge from the notification itself, which stops the escalation. Existing policies can be imported from Opsgenie and PagerDuty. The full picture is on the on call management page and plans are on pricing. To see a policy in action, build one with up to three tiers in the incident simulator and watch who gets paged when the primary does not answer.

More guides

Incident severity levels from SEV1 to SEV4 with definitions and examples

clear definitions with examples, and what each level triggers in paging and communication.

Major incident management checklist for the first 30 minutes

a minute-by-minute checklist for declaring, staffing and communicating a major incident.

AlertOps pricing and AlertOps cost for teams of 5, 15 and 40

Free up to 5 users, Standard at 8 USD up to 10 users, Premium at 18 USD and Enterprise at 28 USD, the Status Hub, stakeholder and OpsIQ add-ons, the per-user message allowance and the monthly cost for 5, 15 and 40 users next to PagerDuty, xMatters and IncidentBot.