Skip to content
IncidentBot

Incident alerting software for on call alerting that reaches the right engineer

IncidentBot is incident alerting software for on-call teams: it turns raw monitoring alerts into one page to the right engineer. Monitoring tools are good at noticing that something is wrong and bad at deciding who should care. IncidentBot sits between your monitoring and your people: it receives every alert, works out which service it belongs to, collapses duplicates, holds back noise from scheduled work and pages the person on call for that service. Fewer pages, and the ones that arrive matter.

Alert grouping for checkout-api

SEV2 Triggered
TimeActorEvent
14:02:11DatadogMonitor "checkout-api p95 > 2s" fired, INC-2481 opened
14:02:34SentryError rate on checkout-api above threshold, grouped into INC-2481
14:03:02CloudWatchLoad balancer 5xx alarm on checkout, grouped into INC-2481
14:03:40DatadogSame monitor fired again, dropped as a duplicate
14:03:41IncidentBot4 alerts, 1 incident, 1 page

Alert sources you already run

Every service in IncidentBot has its own integration endpoint. Point your monitoring at it and alerts start flowing within minutes, with the fields each tool sends parsed into a common format.

Resolved notifications from the source close the alert automatically, so a check that recovers does not leave a stale page behind.

  • Datadog monitors
  • Prometheus Alertmanager
  • Grafana alerts
  • Amazon CloudWatch alarms
  • Sentry issues
  • Inbound email, for tools that only send mail
  • Generic webhook, for anything that can make an HTTP request

Alert management software that routes by service

One Alertmanager or one Datadog account usually covers many services and many teams. Routing rules read the payload and send each alert to the service that owns it, which decides the escalation policy and the people paged.

  • Rules by source, tag, label, host, environment or any payload field.
  • Severity mapping from the source's own priority field to SEV1 through SEV4.
  • Low severity alerts routed to a channel or email instead of a phone call.
  • A catch-all route so an unmatched alert still reaches someone.

Grouping and deduplication

A flapping check or a failing dependency can send hundreds of alerts in a few minutes. IncidentBot groups related alerts into one incident and deduplicates repeats by key, so the on-call engineer gets one page with a counter, not a phone that will not stop buzzing.

Grouping and deduplication rules are included in the Team plan and above.

  • Deduplication by alert key, with the count and last seen time on the incident.
  • Grouping by service, time window or shared fields, for example every alert from one cluster.
  • A single page per group, with every underlying alert listed on the timeline.

Maintenance windows

Scheduled work should not wake anyone. Schedule a maintenance window for a service, and alerts during that window are held and recorded instead of paged. When the window closes, anything still firing is paged as normal.

  • Windows for one service or several, one-off or recurring.
  • Held alerts stay visible on the service, so nothing is lost.
  • A reminder to the owner before a window ends.

Maintenance windows are included in the Team plan and above.

Incident alerting for IT operations and MSPs

The same routing works for infrastructure: servers, network devices, backups and cloud accounts. IT teams route alerts by customer or site, page on-call technicians, escalate to a team lead and keep internal users informed on a private status page.

  • Routing by client, site or environment for managed service providers.
  • Email ingestion for legacy monitoring that cannot send webhooks.
  • Private status pages for internal users on Team and above.

From alert to incident

An alert is where the lifecycle starts. Once routed, it triggers the escalation policy for that service, covered on the on call management page, which pages whoever is on the schedule from on call scheduling software. If the alert is serious, one /incident command opens a Slack channel with the alert context already posted, as shown on the incident response tools page.

Why alert noise is an on-call problem

Every unnecessary page costs attention, sleep and trust in the pager. Engineers who are woken for alerts that need no action start to treat every alert as optional, and that is when a real outage waits. Routing, grouping, deduplication and maintenance windows exist to make each page worth answering, and the on-call load report shows which sources still page too often.

Questions

Can we send alerts from more than one tool to the same service?

Yes. A service can have any number of integrations. Deduplication works across sources when the alerts share a key.

What if our monitoring tool is not on the list?

Use the generic webhook or inbound email. Both accept any payload, and you map the fields for service, severity and summary once.

Does alerting work without Slack?

Yes. Alerting and paging run independently of Slack. Slack is where teams coordinate incidents, not a dependency for being paged.

Which plan includes alert grouping?

Every plan includes all listed integrations and routing. Grouping and deduplication rules and maintenance windows are included in Team and above. See pricing.

See an alert become an incident

Pick an alert source in the incident simulator, paste a payload and watch it route, page and open a channel. Then connect your own monitoring.