Skip to content
IncidentBot

Incident management blog for on-call and SRE teams

Practical guides for engineers and engineering leaders who carry the pager: how to define severity, design rotations and escalation policies, run the first thirty minutes of a major incident and write postmortems people actually read. Each post stands on its own and links to the part of incident management software it relates to.

Latest posts

  1. 01

    AlertOps pricing and AlertOps cost for teams of 5, 15 and 40

    Free up to 5 users, Standard at 8 USD up to 10 users, Premium at 18 USD and Enterprise at 28 USD, the Status Hub, stakeholder and OpsIQ add-ons, the per-user message allowance and the monthly cost for 5, 15 and 40 users next to PagerDuty, xMatters and IncidentBot.

  2. 02

    Grafana IRM pricing and Grafana OnCall pricing per active user

    the free tier for 3 active users, Pro at 20 USD per active IRM user plus the 19 USD platform fee, the 25,000 USD Enterprise minimum, who counts as active and the monthly cost for 5, 15 and 40 users next to PagerDuty, incident.io, Datadog On-Call and IncidentBot.

  3. 03

    xMatters pricing and xMatters cost per user by plan

    Free, Starter at 15 USD up to 25 users, Base at 39 USD and Advanced by quote, the SMS and voice allowances, what happens at 26 users and the monthly cost for 5, 15 and 40 users next to PagerDuty, Opsgenie, AlertOps and IncidentBot.

  4. 04

    Datadog On-Call pricing and Datadog Incident Management pricing per seat

    On-Call at 20 USD, Incident Management at 30 USD and the 40 USD bundle, who counts as a seat, committed and on-demand billing and the monthly cost for 5, 15 and 40 responders next to PagerDuty, incident.io, Rootly, Better Stack and IncidentBot.

  5. 05

    Betterstack pricing and Better Stack pricing per responder for on-call teams

    the 29 USD responder license, status page add-ons from 12 to 208 USD per page, subscriber and monitor packs and the monthly cost for 5, 15 and 40 responders next to PagerDuty, incident.io, Rootly, Squadcast, Grafana Cloud IRM and IncidentBot.

  6. 06

    Squadcast pricing per user, SolarWinds Incident Response plans and what a team really pays

    Pro, Premium and Enterprise after the SolarWinds acquisition, SMS overage, the 10 stakeholder cap and the monthly cost for 5, 15 and 40 users next to PagerDuty, incident.io, Rootly, Better Stack and IncidentBot.

  7. 07

    Rootly pricing per user, Rootly plans and the real cost of incident response plus on-call

    Incident Response and On-Call Essentials, the quoted Enterprise tiers, renewal terms and the monthly cost for 5, 15 and 40 users next to incident.io, PagerDuty, FireHydrant and IncidentBot.

  8. 08

    incident.io pricing per user, incident.io plans and what the on-call add-on really costs

    Basic, Team, Pro and Enterprise with their limits, the on-call add-on and the monthly cost for 5, 15 and 40 users next to FireHydrant, PagerDuty and IncidentBot.

  9. 09

    FireHydrant pricing and plans with cost per responder and what the add-ons change

    Free, Pro and Enterprise with their limits, the SMS and voice add-on, what the Freshworks deal means for renewals and the monthly cost for 5, 15 and 40 responders.

  10. 10

    Statuspage pricing per plan with Atlassian Statuspage limits and the real status page cost

    every public, private and audience-specific plan with its limits, the costs Statuspage leaves out and the full stack price for 5, 10 and 40 responders.

  11. 11

    Opsgenie pricing per user after end of sale with plans, renewals and what replacing it costs

    list prices for every Opsgenie plan, what renewals allow before April 5, 2027 and what Compass, Jira Service Management or IncidentBot cost for 5, 15 and 40 users.

  12. 12

    PagerDuty pricing per user with plans, add-ons and what a team really pays

    list prices for every plan, the add-ons sold separately and the yearly cost for 5, 15 and 40 responders.

  13. 13

    MTTR explained as mean time to resolve, repair and recover

    what each version of MTTR measures, how to calculate it and how to use it without gaming it.

  14. 14

    Incident priority levels in a practical P1 to P4 matrix

    a matrix that sets priority from impact and urgency, so the first responder decides in seconds.

  15. 15

    Incident response plan template for engineering teams

    a complete plan you can copy, covering roles, severity, communication and review.

  16. 16

    Incident response playbook with roles, steps and checklists

    what the commander, comms lead and scribe do from the first alert to resolution.

  17. 17

    On call rotation with 6 schedule patterns and when to use them

    weekly, split, follow-the-sun and other rotations compared by team size and time zones.

  18. 18

    On call schedule template with sample weekly and follow-the-sun rotations

    ready schedules you can adapt, with hand-off times and override rules.

  19. 19

    Postmortem template in a blameless format your team will actually fill in

    a short blameless structure built around the timeline and concrete action items.

  20. 20

    Escalation policy design with tiers, timeouts and fallbacks

    how to choose tiers and timeouts so every alert reaches someone and nobody is woken without need.

  21. 21

    Incident severity levels from SEV1 to SEV4 with definitions and examples

    clear definitions with examples, and what each level triggers in paging and communication.

  22. 22

    Major incident management checklist for the first 30 minutes

    a minute-by-minute checklist for declaring, staffing and communicating a major incident.

What we write about

The posts follow the incident lifecycle. For scheduling and paging, start with the rotation and escalation guides and then see on call scheduling software. For running incidents, read the playbook and severity guides alongside incident response tools. For learning, the postmortem template and the MTTR guide pair with incident postmortems.

See the practices in a working incident

The incident simulator runs a sample incident with the escalation policy you build, a Slack channel, a status page draft and a postmortem skeleton.