Skip to content
IncidentBot

Incident management platform for the whole incident lifecycle

An incident moves through five stages: a signal arrives, a person is woken up, a team works the problem, customers are told what is happening, and someone writes down what was learned. IncidentBot covers all five in one product, with one data model and one price per responder seat. This page walks through every feature in the order an incident actually uses it.

INC-2481 checkout-api p95 latency above 2s

SEV2 Resolved
TimeActorEvent
14:02:11DatadogMonitor "checkout-api p95 > 2s" fired
14:02:12IncidentBotGrouped 3 related alerts into INC-2481
14:02:12IncidentBotPaged primary on-call, Payments (push)
14:07:12IncidentBotNo acknowledgement in 5 min, escalated to secondary (SMS, voice)
14:08:40Secondary on-call, PaymentsAcknowledged
14:09:02IncidentBotOpened #inc-2481-checkout-api-latency, assigned commander
14:12:30CommsStatus page: "Degraded performance: Checkout" published
14:41:05CommanderResolved, rollback of release 2026.09.3
14:41:06IncidentBotPostmortem draft created with 8 timeline events

One incident management platform, five stages

Most stacks split this work across separate products for on-call, incident response, status communication and postmortems. Each one has its own login, its own copy of the timeline and its own invoice. In IncidentBot the alert that pages the engineer is the same record that opens the Slack channel, feeds the status page draft and becomes the first row of the postmortem.

  • Alert: receive, route, group and deduplicate signals from your monitoring.
  • Page: reach the right person through schedules and escalation policies.
  • Coordinate: run the incident in a dedicated Slack channel with clear roles.
  • Inform: keep customers and stakeholders updated through hosted status pages.
  • Learn: turn the captured timeline into a blameless postmortem and tracked action items.

Alert routing, grouping and deduplication

IncidentBot receives alerts from Datadog, Prometheus Alertmanager, Grafana, CloudWatch, Sentry, email and any system that can send a webhook. Routing rules read the payload and send each alert to the right service, so a latency alarm on checkout-api pages the payments team and not the whole company.

The details are on the incident alerting page.

  • Integrations with Datadog, Prometheus Alertmanager, Grafana, CloudWatch, Sentry, inbound email and generic webhooks.
  • Routing rules by service, tag, source and payload field.
  • Grouping and deduplication, so fifty identical alerts from a flapping check become one incident with a counter.
  • Maintenance windows that hold alerts during scheduled work and release them when the window closes.

On-call schedules and escalation policies

Every service has an owner and every owner has a schedule. IncidentBot keeps rotations, overrides and swaps in one calendar, and escalation policies decide who is paged next when nobody acknowledges.

Read more about on call scheduling software and on call management.

  • Weekly, daily and custom rotations, follow-the-sun handoffs across time zones, overrides and swaps.
  • Calendar sync, so engineers see their shifts next to their meetings.
  • Escalation policies with tiers and timeouts, repeated until someone acknowledges.
  • Paging by mobile push, SMS, voice call, email and Slack. Paging does not depend on Slack, so an outage in chat never silences an alert.
  • Import of schedules and escalation policies from Opsgenie and PagerDuty.

Incident coordination in Slack

Type /incident in Slack and IncidentBot opens a dedicated channel such as #inc-2481-checkout-api-latency, invites the on-call responders and posts the alert context. The team works where it already talks, and every decision lands on the timeline.

See the full workflow on the incident response tools page.

  • Roles for incident commander, communications lead and scribe, assigned in one command.
  • Severity levels SEV1 to SEV4 with custom fields for impact, customer segment or region.
  • A live timeline that records messages you pin, role changes, severity changes and status updates.
  • Runbooks attached to services and posted into the channel when the incident opens.
  • Microsoft Teams support on the Business plan.

Status pages drafted from the incident

Customers should hear about an outage from you, not from social media. IncidentBot drafts the status update from the incident itself: affected component, impact level and a plain language summary. The communications lead edits it if needed and publishes in one click.

  • Hosted status pages, public or private.
  • Subscriber notifications by email when an incident opens, changes or resolves.
  • Custom domain such as status.yourcompany.com.
  • Audience-specific status pages on the Business plan, for example one page per enterprise customer.

Postmortems and reporting

The timeline is captured while the incident runs, so the postmortem starts with facts instead of memory. IncidentBot builds a blameless draft with the timeline, impact window and response times, and the team fills in contributing factors and follow-ups.

  • Automatic timeline capture from alerts, pages, acknowledgements and Slack.
  • Blameless postmortem draft with impact, timeline and sections to complete.
  • Action items synced to Jira and Linear, so follow-ups live in the backlog they belong to.
  • MTTA and MTTR reports by service, team and severity.
  • On-call load report that shows who was paged, when, and how often out of hours.
  • SLA and uptime per service on the Business plan.

Who uses the incident management platform

The same product fits several ways of running operations. Teams choose the parts they need and grow into the rest.

  • SaaS product teams: engineers carry the pager for the services they build, with a fair rotation and one place to run incidents.
  • SRE and platform teams: a shared on-call and incident process across many teams, with consistent severities, roles and reporting.
  • IT operations and MSPs: infrastructure alerts routed to on-call technicians, escalation to team leads and a private status page for internal users.
  • Customer-facing incidents: support and communications teams follow the incident and update customers through the status page without chasing engineers.

Questions

Do we need every stage from day one?

No. Many teams start with alerting and on-call, then turn on Slack incident channels and the status page once schedules are settled. Everything is included in the plan, so there is nothing extra to buy when you are ready.

Who counts as a paid seat?

A responder seat is a user who can be on call or act on incidents. Stakeholder viewers with read-only access are included in every plan and are not billed. Plan details are on the pricing page.

Can we try it before we buy?

There is no free plan. Every plan is a monthly or annual subscription. You can try the incident simulator on the demo page without an account.

See the full lifecycle in two minutes

Run a sample incident in the browser, then create your account and connect your first alert source.