Skip to content
IncidentBot

MTTR explained as mean time to resolve, repair and recover

MTTR is the metric every engineering leader is asked about and few teams measure the same way twice. The letters stand for at least four different things, the clock can start and stop at several points, and a single average hides most of what matters. This guide sets out the definitions, shows how to measure them consistently, and explains how to use the number without letting it mislead you.

MTTR meaning, one acronym and four metrics

The R in MTTR has no single official meaning. Before you compare numbers between teams, tools or quarters, agree which one you are reporting.

MetricClock startsClock stopsWhat it tells you
Mean time to respondAlert fires or incident is openedFirst responder starts working on itHow quickly people engage
Mean time to repairWork on the fix beginsThe fix is deployedHow long hands-on repair takes
Mean time to recoverImpact beginsUsers are no longer affectedHow long customers feel the problem
Mean time to resolveImpact begins or incident is openedIncident is closed, including follow-up needed to stop an immediate repeatThe full cost of the incident to the team

Mean time to resolve is the broadest of the four and the one most incident management teams mean when they say MTTR. Mean time to recover is the one customers care about. Mean time to repair is a hardware maintenance term that crossed over into software, and it is most useful when you want to separate diagnosis from the fix itself.

The companion metrics

  • MTTA, mean time to acknowledge: from the page to the moment someone acknowledges it. This is the metric your escalation policy and paging channels influence most directly.
  • MTTD, mean time to detect: from the start of impact to the first alert. It measures your monitoring, not your responders.
  • MTBF, mean time between failures: the average gap between incidents on the same service. It measures reliability, while MTTR measures recovery.

How MTTR fits the DORA metrics

The DORA research programme, which grew out of the State of DevOps reports and the book Accelerate, measures software delivery performance with a small set of key metrics: deployment frequency, lead time for changes, change failure rate and time to restore service. In more recent DORA reports the last one is framed as failed deployment recovery time, the time it takes to recover from a deployment that caused a failure in production.

The DORA framing is useful because it pairs speed with stability. A team that ships often and recovers quickly is in a better position than a team that ships rarely and treats every release as a risk. It also narrows the definition: DORA looks at failures caused by change, while an incident MTTR usually includes everything, from a cloud provider outage to a certificate that expired on a Sunday.

Mean time to resolve measurement checklist

Most MTTR disputes are really disputes about timestamps. Write down the rules once and apply them to every incident.

  • Define the start. Choose the first alert, the start of customer impact, or the moment the incident was declared. Start of impact is the most honest choice, but it is often only known after the postmortem, so record it then.
  • Define the end. Recovered means customers are no longer affected. Resolved means the incident is closed. Keep both timestamps and report them separately.
  • Decide what counts. Exclude test pages and duplicate alerts. Decide whether SEV4 incidents are in or out, and keep that rule stable.
  • Segment before you average. Report MTTR by severity and by service. A single number across SEV1 outages and SEV4 annoyances describes nothing.
  • Capture timestamps as they happen. Reconstructing a timeline from memory a week later moves every number by minutes, sometimes hours.

Why the average misleads

Incident durations are not evenly spread. Most incidents are short, and a few are very long, so one long outage can move a quarterly mean more than every other incident combined. Google engineers made this point in the O'Reilly report Incident Metrics in SRE, which argues that mean-based metrics such as MTTR are poorly suited to judging whether reliability is improving, because the distribution is too skewed and the sample too small for the mean to be stable.

In practice that means three things. Report the median and a high percentile such as the 90th alongside the mean. Look at the distribution, not just the headline. And treat a quarter with fewer incidents as a small sample, not as proof that something changed.

Where the minutes actually go

When you break an incident timeline into phases, the repair itself is rarely the longest part. Detection, getting the right person engaged, and working out what changed usually take longer.

PhaseTypical cause of delayWhat shortens it
DetectAlert on symptoms missing, thresholds too looseAlert on user-facing symptoms and service level objectives
AcknowledgePage went to the wrong person, silent phone, no fallbackOn-call schedules with escalation tiers and timeouts
AssembleHunting for the right channel, the right people, the dashboard linkA dedicated incident channel with roles assigned on open
DiagnoseNo record of recent changes, tribal knowledgeRunbooks attached to services, deploy events on the timeline
RecoverRisky manual fixesRollback as the default first move, feature flags
ResolveLoose ends, no owner for follow-upAction items tracked in the issue tracker

Practices that consistently help

  • Mitigate first, investigate later. The Google SRE book puts it plainly: in an incident, the priority is to stop the bleeding and restore service, and root cause work comes after.
  • Roll back before you debug. If a deploy correlates with the impact, reverting is usually faster and safer than a forward fix.
  • Use an incident commander. One person coordinates while others fix. The Incident Command System used by Google SRE separates command, operations and communication for exactly this reason.
  • Write the timeline while it happens. A live timeline makes handovers fast and turns the postmortem into editing rather than archaeology.
  • Fix the cause of slow acknowledgement, not the person. If MTTA is high at night, the escalation policy is usually the problem.

What not to do with MTTR

Do not set an MTTR target for individual engineers, and do not rank teams by it. Teams that are judged on the number learn to close incidents early, split them, or stop declaring them. Use MTTR as a trend to ask better questions in reviews, next to customer impact and the quality of follow-up.

Track MTTA and MTTR without reconstructing timelines

IncidentBot records every step of an incident as it happens: the alert, the page, the acknowledgement, the Slack channel, status page updates, and resolution. Those timestamps feed MTTA and MTTR reports per service and per severity, so the numbers come from the incident itself rather than from memory. See how it works on the incident tracking software page, read about incident severity levels to segment your reports, and compare plans on the pricing page. To see a timeline and the computed time to acknowledge and time to resolve, run the incident simulator in your browser.

More guides

Incident priority levels in a practical P1 to P4 matrix

a matrix that sets priority from impact and urgency, so the first responder decides in seconds.

Incident response plan template for engineering teams

a complete plan you can copy, covering roles, severity, communication and review.

Incident response playbook with roles, steps and checklists

what the commander, comms lead and scribe do from the first alert to resolution.