The parts of an on-call escalation policy
A policy is an ordered list of tiers. Each tier names who is paged and how long the system waits for an acknowledgement before moving on. Around the tiers sit a few rules that decide what happens at the end of the list.
| Part | What it defines | Typical choice |
|---|---|---|
| Tier 1 target | Who is paged first | The primary layer of the service's on-call schedule |
| Tier 1 timeout | Wait before escalating | 5 minutes for urgent alerts |
| Tier 2 target | Who is paged next | The secondary layer of the same schedule |
| Tier 3 target | Last human line | Team lead or engineering manager on call |
| Repeat rule | What happens after the last tier | Restart from tier 1, one or two more times |
| Channels | How each person is reached | Push first, then SMS and voice for urgent alerts |
Point tiers at schedule layers rather than at named people. A policy that says "page Priya" breaks the first time Priya is on holiday. A policy that says "page the primary on-call for payments" follows the on call schedule template your team already maintains, overrides included.
Timeouts per urgency
The timeout is a trade-off. Too short and the secondary is woken for alerts the primary was about to acknowledge. Too long and a real outage waits while a phone vibrates in another room. The numbers below are starting points to tune against your own acknowledgement data, not a standard.
| Urgency | Tier 1 timeout | Tier 2 timeout | Channels |
|---|---|---|---|
| SEV1, customer-facing outage | 5 minutes | 5 minutes | Push, SMS and voice at once, then voice on escalation |
| SEV2, major degradation | 10 minutes | 10 minutes | Push, then SMS, then voice |
| SEV3, limited impact | 30 minutes, business hours only | Next business day | Push and email |
| SEV4, informational | No paging | No paging | Ticket or channel message |
Two rules help. First, low urgency alerts should not page at night at all; they wait for working hours or become tickets. Second, if people regularly acknowledge just before the timeout, the timeout is too short or the alert is too noisy. Check your time to acknowledge figures before changing either one. The severity definitions behind this table are in our guide to incident severity levels.
Sample escalation policy for a customer-facing service
Here is a complete policy for a checkout service, written as it would appear in a runbook.
- Tier 1: primary on-call for checkout. Push immediately, SMS after 1 minute, voice call after 3 minutes. Escalate after 5 minutes without acknowledgement.
- Tier 2: secondary on-call for checkout. Same channel sequence. Escalate after 5 minutes.
- Tier 3: engineering manager on call for the commerce group, plus the whole checkout team by push. Escalate after 10 minutes.
- Repeat: restart from tier 1 twice. After the second full cycle, page the incident commander rotation.
- Acknowledgement stops escalation. Reassignment or explicit escalation by the responder moves the incident to the chosen tier immediately.
The last line matters. An engineer who acknowledges and then realises the problem belongs to the database team should be able to escalate or reassign in one step, rather than resolving the page and hoping someone else picks it up.
Fallbacks that close the gaps
Every policy has holes. These are the ones that cause missed pages most often.
- An empty schedule layer. If nobody is assigned for a time window, the tier has no target. The policy should skip to the next tier rather than wait out the timeout on nobody.
- Out-of-office without an override. Holidays belong in the schedule as overrides, entered before the person leaves.
- Paging that depends on chat. If Slack or Teams is part of the incident, a chat-only page never arrives. Keep push, SMS and voice in the chain.
- Notification rules on the phone. A primary with do-not-disturb and no exception for the paging app is effectively not on call. Ask new responders to send a test page to themselves during onboarding.
- The end of the list. A policy that stops after the last tier leaves the incident unowned. Repeat, then fall back to a rotation that is always staffed.
Escalation policy mistakes to review for
- Paging a whole team at tier 1. Everyone assumes someone else has it, and acknowledgement gets slower, not faster.
- Escalating straight to managers. Managers belong at the end, as a signal that the process did not work, not as the second responder.
- One policy for every service. A batch job and a payment API do not need the same urgency.
- No review. Look at escalations every month: which policies escalated past tier 1, and why. Frequent tier 2 pages usually point to a noisy alert or an overloaded primary, both fixable.
If your policies escalate often at night, the root cause is usually alerting, not people. Grouping, deduplication and moving non-urgent alerts to working hours reduce pages more than any timeout change, and the on call rotation guide covers how rotation length affects fatigue.
Build and test policies in on-call management
In IncidentBot, escalation policies have tiers with their own targets and timeouts, point at schedule layers, repeat a set number of times and page by mobile push, SMS, voice, email and Slack, so paging keeps working when chat is down. Responders acknowledge from the notification itself, which stops the escalation. Existing policies can be imported from Opsgenie and PagerDuty. The full picture is on the on call management page and plans are on pricing. To see a policy in action, build one with up to three tiers in the incident simulator and watch who gets paged when the primary does not answer.