MTTR meaning, one acronym and four metrics
The R in MTTR has no single official meaning. Before you compare numbers between teams, tools or quarters, agree which one you are reporting.
| Metric | Clock starts | Clock stops | What it tells you |
|---|---|---|---|
| Mean time to respond | Alert fires or incident is opened | First responder starts working on it | How quickly people engage |
| Mean time to repair | Work on the fix begins | The fix is deployed | How long hands-on repair takes |
| Mean time to recover | Impact begins | Users are no longer affected | How long customers feel the problem |
| Mean time to resolve | Impact begins or incident is opened | Incident is closed, including follow-up needed to stop an immediate repeat | The full cost of the incident to the team |
Mean time to resolve is the broadest of the four and the one most incident management teams mean when they say MTTR. Mean time to recover is the one customers care about. Mean time to repair is a hardware maintenance term that crossed over into software, and it is most useful when you want to separate diagnosis from the fix itself.
The companion metrics
- MTTA, mean time to acknowledge: from the page to the moment someone acknowledges it. This is the metric your escalation policy and paging channels influence most directly.
- MTTD, mean time to detect: from the start of impact to the first alert. It measures your monitoring, not your responders.
- MTBF, mean time between failures: the average gap between incidents on the same service. It measures reliability, while MTTR measures recovery.
How MTTR fits the DORA metrics
The DORA research programme, which grew out of the State of DevOps reports and the book Accelerate, measures software delivery performance with a small set of key metrics: deployment frequency, lead time for changes, change failure rate and time to restore service. In more recent DORA reports the last one is framed as failed deployment recovery time, the time it takes to recover from a deployment that caused a failure in production.
The DORA framing is useful because it pairs speed with stability. A team that ships often and recovers quickly is in a better position than a team that ships rarely and treats every release as a risk. It also narrows the definition: DORA looks at failures caused by change, while an incident MTTR usually includes everything, from a cloud provider outage to a certificate that expired on a Sunday.
Mean time to resolve measurement checklist
Most MTTR disputes are really disputes about timestamps. Write down the rules once and apply them to every incident.
- Define the start. Choose the first alert, the start of customer impact, or the moment the incident was declared. Start of impact is the most honest choice, but it is often only known after the postmortem, so record it then.
- Define the end. Recovered means customers are no longer affected. Resolved means the incident is closed. Keep both timestamps and report them separately.
- Decide what counts. Exclude test pages and duplicate alerts. Decide whether SEV4 incidents are in or out, and keep that rule stable.
- Segment before you average. Report MTTR by severity and by service. A single number across SEV1 outages and SEV4 annoyances describes nothing.
- Capture timestamps as they happen. Reconstructing a timeline from memory a week later moves every number by minutes, sometimes hours.
Why the average misleads
Incident durations are not evenly spread. Most incidents are short, and a few are very long, so one long outage can move a quarterly mean more than every other incident combined. Google engineers made this point in the O'Reilly report Incident Metrics in SRE, which argues that mean-based metrics such as MTTR are poorly suited to judging whether reliability is improving, because the distribution is too skewed and the sample too small for the mean to be stable.
In practice that means three things. Report the median and a high percentile such as the 90th alongside the mean. Look at the distribution, not just the headline. And treat a quarter with fewer incidents as a small sample, not as proof that something changed.
Where the minutes actually go
When you break an incident timeline into phases, the repair itself is rarely the longest part. Detection, getting the right person engaged, and working out what changed usually take longer.
| Phase | Typical cause of delay | What shortens it |
|---|---|---|
| Detect | Alert on symptoms missing, thresholds too loose | Alert on user-facing symptoms and service level objectives |
| Acknowledge | Page went to the wrong person, silent phone, no fallback | On-call schedules with escalation tiers and timeouts |
| Assemble | Hunting for the right channel, the right people, the dashboard link | A dedicated incident channel with roles assigned on open |
| Diagnose | No record of recent changes, tribal knowledge | Runbooks attached to services, deploy events on the timeline |
| Recover | Risky manual fixes | Rollback as the default first move, feature flags |
| Resolve | Loose ends, no owner for follow-up | Action items tracked in the issue tracker |
Practices that consistently help
- Mitigate first, investigate later. The Google SRE book puts it plainly: in an incident, the priority is to stop the bleeding and restore service, and root cause work comes after.
- Roll back before you debug. If a deploy correlates with the impact, reverting is usually faster and safer than a forward fix.
- Use an incident commander. One person coordinates while others fix. The Incident Command System used by Google SRE separates command, operations and communication for exactly this reason.
- Write the timeline while it happens. A live timeline makes handovers fast and turns the postmortem into editing rather than archaeology.
- Fix the cause of slow acknowledgement, not the person. If MTTA is high at night, the escalation policy is usually the problem.
What not to do with MTTR
Do not set an MTTR target for individual engineers, and do not rank teams by it. Teams that are judged on the number learn to close incidents early, split them, or stop declaring them. Use MTTR as a trend to ask better questions in reviews, next to customer impact and the quality of follow-up.
Track MTTA and MTTR without reconstructing timelines
IncidentBot records every step of an incident as it happens: the alert, the page, the acknowledgement, the Slack channel, status page updates, and resolution. Those timestamps feed MTTA and MTTR reports per service and per severity, so the numbers come from the incident itself rather than from memory. See how it works on the incident tracking software page, read about incident severity levels to segment your reports, and compare plans on the pricing page. To see a timeline and the computed time to acknowledge and time to resolve, run the incident simulator in your browser.