DORA metrics are the four, arguably five, measures Google's DevOps Research and Assessment program uses to describe how well a software team ships and operates code: deployment frequency, lead time for changes, change failure rate, and time to restore service. Whether there's a fifth metric depends on which page you land on, and the sources split into two camps that don't cite each other. dora.dev and IBM both call it deployment rework rate, and IBM dates its arrival to 2024; getdx and a 2021 editor's note on Google's own Four Keys post both call it reliability instead.
What are DORA metrics?
DORA started as an independent research program studying software delivery performance, DevOps Research and Assessment. Its findings are now published at dora.dev, under Google Cloud, and the four (or five) metrics below are the operational core of that research: measurable proxies for how fast and how safely a team ships software.
One disambiguation worth stating plainly: DORA metrics and the EU's Digital Operational Resilience Act, Regulation (EU) 2022/2554, a financial-sector ICT-risk rule in force since 17 January 2025, share an acronym and nothing else.
The four DORA metrics, defined
DORA's own guide to the four metrics sets the baseline definition every vendor page on this list copies, with only minor wording differences.
Deployment frequency
How often a team successfully releases to production. Elite teams deploy multiple times a day; a team still batching releases into weekly or monthly windows is measuring something closer to release management than deployment frequency.
Lead time for changes
The time from a commit landing in the main branch to that code running in production. This is the metric most sensitive to process rather than skill: a slow approval queue inflates lead time as much as slow code does.
Change failure rate
The percentage of deployments that cause a failure in production, whether that shows up as an outage, a rollback, or a hotfix. It's the one metric that punishes speed without corresponding quality.
Failed deployment recovery time
How long it takes to restore service once something breaks in production. dora.dev's current term for this metric is failed deployment recovery time, which replaced MTTR; older material and many searches still call it time to restore service. The goal hasn't changed: restoring what the user sees, not necessarily fixing the underlying cause on the same timeline.
Is there a 5th DORA metric?
Ask this question and you'll get two different answers depending on which page you land on, and the two sides don't cite each other.
dora.dev's fifth-metric entry calls it deployment rework rate: the proportion of unplanned deployments triggered by a production incident rather than planned work. IBM's own DORA metrics page uses the same name and the same definition, and dates the metric's arrival to 2024. getdx's DORA metrics guide disagrees on both the name and the definition: it calls the fifth metric Reliability, measured through uptime, response time, and correctness, closer to a system-health score than a rework count. Google's 2020 Four Keys post carries a 2021 editor's note that lands on the same side as getdx, naming a fifth metric called reliability without defining it further.
One side ties the fifth metric to deployment rework, the other to system reliability, and neither cites the other. GitLab's DORA analytics documentation is the one still-ranking page that sticks to four metrics with no fifth mentioned at all. Until dora.dev, IBM, getdx, and Google converge on one definition, treat the fifth DORA metric as a live disagreement, not a settled addition.
DORA performance tiers and benchmarks
IBM's DORA metrics benchmarks put the elite tier at multiple deployments a day, lead time under an hour, failed deployment recovery time under an hour, and a change failure rate of 0-15%.
| Metric | Elite tier |
|---|---|
| Deployment frequency | Multiple deploys per day |
| Lead time for changes | Under 1 hour |
| Failed deployment recovery time | Under 1 hour |
| Change failure rate | 0-15% |
High tier is where the two most-cited benchmark pages stop matching. IBM puts the high-tier change failure rate at 16-30%; Datadog's DORA metrics knowledge base entry puts the same tier at 15-30%. One point apart isn't much, but neither page cites a specific table in DORA's underlying research report, so treat both as approximate rather than exact. IBM and Datadog both publish medium and low tier bands too. This piece only tables the elite tier, where the two agree, and the high-tier change failure rate, where they don't, since that's the more useful contrast for benchmarking your own team. For the full grid, see the IBM benchmarks linked above.
How to start tracking DORA metrics
1. Instrument your CI/CD pipeline
Deployment frequency and lead time both come from the same data: your CI/CD system needs to log a timestamp for every commit merged to main and every successful production deployment. If your pipeline doesn't already emit that, this is the actual first step, not picking a dashboard tool.
2. Link deployments to incidents
Change failure rate and time to restore service both require your incident tracker and your deployment log to reference the same deployment ID. Without that link, you can count incidents and you can count deployments, but you can't compute the ratio between them, which is the point of the metric.
3. Pick a way to compute it
GitLab's dashboard, introduced above, computes all four metrics directly from CI/CD data already sitting in the platform: deployment frequency as the mean, not median, of successful production deployments, lead time as merge-time to production, time to restore as the median duration an incident stays open, and change failure rate as incidents divided by deployments. Google's Four Keys project, also mentioned above, does the ETL work itself, pulling GitHub or GitLab deployment and incident data into BigQuery and visualizing it in a DataStudio dashboard. It's been running since 2020, computes the four original metrics only, and still ranks on page one for this search, ahead of pages published far more recently. If you'd rather not build or maintain either, paid platforms like getdx and Jellyfish package the same calculations into a dashboard you don't have to own.
4. Decide who owns the pipeline before you start measuring
Someone has to own the instrumentation itself, not just read the resulting dashboard. On a single in-house team, that's usually whoever already owns CI/CD. On a team split across time zones or vendors, that ownership question doesn't have a default answer, and skipping it is how teams end up with dashboards nobody trusts.
None of this is worth doing if the numbers end up gamed instead of measured. A team under pressure to hit an elite deployment frequency can ship trivial commits just to inflate the count, activity that looks productive on a dashboard without being productive at all. HighCircl has made the case for why busy isn't the same as productive elsewhere, and DORA metrics reward exactly the wrong behavior if nobody's watching for it. They're one set of engineering KPIs among several, but they're the rare ones with a shared, external definition, which is worth protecting.
DORA metrics for a distributed or nearshore engineering team
Lead time for changes is defined as commit-to-production time, and that definition assumes review happens close to when the commit lands. On a team split across time zones, it often doesn't. A commit opened at 4pm in one zone can sit unreviewed for twelve hours before someone in another zone picks it up. That's a handoff delay, and it inflates the lead-time number the same way slow code would. Read it without accounting for the time-zone gap, and you'll conclude the team is slower than it actually is.
Measuring the handoff separately from actual coding and review time fixes this without needing a different metric at all, so a long lead time caused by a time-zone gap doesn't get mistaken for a process problem.
That's one reason step four above, deciding who owns the pipeline, matters more on a blended team than on a single-location one. The engineer who owns your deploy pipeline end to end is also the person best placed to say whether a given lead-time number reflects real delay or just a time-zone handoff. On a team where product ownership, architecture opinions, and CI/CD instrumentation are split between in-house and augmented engineers, how a blended team's roles are usually split needs to be decided explicitly, not assumed, before the DORA numbers mean much.
Does AI change what counts as elite?
Some of it, yes. DORA's 2024 report found that individual productivity, flow, and job satisfaction went up with AI adoption while software delivery stability and throughput went down. How AI adoption is already showing up in DORA's own surveys follows how that picture has moved since. A deployment frequency or change failure rate benchmark set before your team adopted AI tooling may already be measuring the wrong baseline.
FAQ
What are DORA metrics?
DORA metrics are four measures of software delivery performance defined by Google's DevOps Research and Assessment program: deployment frequency, lead time for changes, change failure rate, and time to restore service. Some sources, including DORA's own guide, now describe a fifth metric, though there's no agreement yet on what it's called or what it measures.
Is "DORA metrics" the same as the EU DORA regulation?
No. DORA metrics come from DevOps research. The EU's Digital Operational Resilience Act, Regulation (EU) 2022/2554, is a financial-sector ICT-risk rule that's been in force since 17 January 2025. They share an acronym and nothing else.
What counts as a good deployment frequency for a small team?
DORA's elite tier is multiple deployments a day, but that benchmark comes from organizations with mature CI/CD and a large enough team to support it. A small team shipping once a day or a few times a week is probably just matched to its own size and risk tolerance, not behind the industry. Track the trend for your own team over time rather than chasing the elite number directly.
Do DORA metrics work for a nearshore or distributed team?
Yes, with one adjustment: measure the review handoff between time zones separately from actual coding and review time, so a long lead time caused by a time-zone gap doesn't get counted as a process problem. Decide who owns the CI/CD instrumentation before you start measuring, since that ownership question doesn't have a default answer on a blended team.
What is the 5th DORA metric?
There isn't a settled answer yet. dora.dev and IBM both call it deployment rework rate, the share of unplanned deployments triggered by a production incident, and IBM dates its arrival to 2024. getdx calls it Reliability instead, measured through uptime, response time, and correctness, and a 2021 editor's note on Google's Four Keys post lands on the same side. GitLab's docs are the one still-ranking page that sticks to four metrics with no fifth at all. Treat the fifth metric as an open disagreement between sources, not an established addition.
