Two teams can report cycle time and measure different intervals, because one starts the clock at the first commit while the other starts when a pull request opens. Unless those boundaries are explicit, comparing the numbers tells a leader little about where work slows down.
This guide explains how to calculate 15 engineering metrics, which decisions they support, and where their interpretation can go wrong.
Engineering metrics for leadership decisions in 2026
| # | Engineering metric | What it measures | Primary use |
|---|---|---|---|
| 1 | Deployment frequency | How often an application or service reaches production | Examine release cadence and batch size |
| 2 | Change lead time | Elapsed time from commit to successful production execution | Find delays after code is committed |
| 3 | Failed deployment recovery time | Time to recover from a deployment-caused failure | Assess recovery capability for change-related failures |
| 4 | Change fail rate | Share of deployments that require remediation | Watch delivery instability |
| 5 | Deployment rework rate | Share of deployments that are unplanned responses to production incidents | Quantify reactive deployment work |
| 6 | PR cycle time | Elapsed time across an explicitly defined pull-request interval | Locate delay across the change workflow |
| 7 | Time to first review | Time from PR opening to the first human review | Examine reviewer pickup and capacity |
| 8 | Build and test feedback time | Time from starting a build or test to receiving a usable result | Prioritize slow feedback loops |
| 9 | Throughput | Completed items of one defined type per period | Understand delivery rate and support forecasting |
| 10 | Work in progress | Active items in a workflow or stage at a point in time | Expose queues, overload, and aging work |
| 11 | Engineering allocation | Share of effort assigned to stable work categories | Compare actual investment with priorities |
| 12 | Planning accuracy | Share of forecast Sprint work completed under a stable counting rule | Calibrate Scrum forecasts and reveal unplanned work |
| 13 | SLO compliance and error budget | Whether user-relevant service behavior meets its objective | Balance reliability and release decisions |
| 14 | Developer satisfaction | Recurring survey responses about the work experience of developers | Find friction that workflow telemetry cannot observe |
| 15 | AI-assisted delivery impact | Delivery, reliability, and review costs for comparable AI-assisted work | Evaluate whether AI improves the whole delivery system |
Define the metric before you compare the number
Before adding an engineering metric to a dashboard, record the following six fields.
- Specify the start and end events, states, or survey responses used.
- Define the population, including services, teams, work types, or respondents.
- State the unit counted, such as deployments, work items, or requests.
- Record the reporting window and aggregation, such as a rate, median, or percentile.
- Name the decision the metric can inform.
- Explain what it cannot determine alone.
For time-based measures, preserve the distribution because an average can hide unusually slow work. Google SRE explains this limitation for latency, where measurement windows and percentiles affect what the number reveals.
Place the definition beside the chart so someone else can reproduce the calculation. If the events, population, or counting rule change, record that change before comparing results across periods.
Software delivery performance metrics
The current DORA software delivery model groups five metrics into throughput and instability. As the model evolved from the older four-metric set, failed deployment recovery time replaced the broader MTTR framing and deployment rework rate became a distinct measure.
DORA recommends applying these measures to a single application or service, because combining services with different release paths and operating risks can obscure their performance.
1. Deployment frequency
Deployment frequency counts production deployments for a defined application or service during a reporting window. Exclude staging releases and document how canaries, rollbacks, and repeated pipeline runs affect the count.
Use it to examine release cadence and batch size. If frequency falls, investigate release approvals, test duration, integration queues, or larger delivery batches.
Frequent deployments do not establish customer value, since configuration changes can raise the count without improving the product. Read frequency alongside stability and a relevant product outcome.
2. Change lead time
Change lead time measures elapsed time from code committed in version control until that change runs successfully in production. With sufficiently complete data, report a median and a higher percentile for the defined application, work type, and window.
Use it to investigate delays in integration, testing, approval, and deployment. A rising tail can reveal slow changes even when the median remains stable.
Because the clock starts at commit, it excludes discovery, design, backlog wait, and part of coding. To measure from a customer request instead, add a work-item lead time with a separately defined starting point.
3. Failed deployment recovery time
Failed deployment recovery time measures the interval from a deployment-caused impairment requiring immediate intervention to restored service. Define the exact failure and restoration events used in the calculation.
Use the result to investigate detection, ownership, rollback or fix-forward paths, and diagnostic information. Examine the slow tail because difficult recoveries may carry more risk than the typical incident.
Keep it separate from a general mean time to restore or resolve, since infrastructure outages and third-party failures may require different owners and recovery procedures.
4. Change fail rate
Change fail rate is the number of production deployments that require immediate intervention divided by total production deployments in the same scope and window. Define remediation before calculating the rate, including whether it covers rollback, fix-forward, hotfix, patch, or another intervention.
If delivery accelerates while this rate rises, inspect test strategy, batch size, release controls, and recovery readiness. Also track incident impact, since a low rate can still include severe failures.
The rate does not count every defect, because bugs that do not trigger immediate remediation may fall outside the numerator.
5. Deployment rework rate
Deployment rework rate is the share of deployments that were unplanned responses to production incidents. Use the same deployment population as deployment frequency and document how the organization attributes an unplanned deployment to an incident.
Where change fail rate counts deployments requiring intervention, rework rate counts the unplanned deployments made in response to incidents. If rework rises, investigate recurring incidents, ineffective fixes, and their pressure on planned work.
Refactoring a branch or revising a pull request is code rework, so it enters this numerator only if it produces an incident-driven production deployment.
Flow and feedback metrics
Flow and feedback measures help locate where work waits. Kanban defines delivery rate and WIP through explicit workflow boundaries, while DevEx research examines friction from slow tool and human feedback.
6. PR cycle time
PR cycle time measures elapsed time between two chosen events in the pull-request workflow. Subtract the start timestamp from the end timestamp and display that boundary alongside the result.
For review, the interval may run from PR opened to merged. A broader change-flow definition runs from first commit to successful production deployment, split into coding, pickup, review, merge, and deploy stages. Since this includes work outside the PR itself, the two definitions produce values that are not directly comparable.
Investigate the stage where delay accumulates, then check stability and review signals to assess whether a shorter interval preserved quality.
7. Time to first review
Time to first review measures elapsed time from a pull request opening to the first substantive human review action. Exclude automated bot comments or report them separately, and document whether an approval, requested change, or line-level comment qualifies.
Use it to examine review ownership, time-zone coverage, and capacity. DORA research from 2023 associated fast code reviews with better delivery and operational performance, without establishing a universal threshold.
Fast pickup does not prove a useful review, so also examine completion time, repeated rounds, stability signals, and developer feedback.
8. Build and test feedback time
Build and test feedback time measures the interval from starting a run to receiving an actionable result. Separate local from CI feedback, and build from test duration when the owners or remedies differ.
The DevEx framework identifies feedback loops as one of three core dimensions and recommends combining workflow observations with developer perceptions. Track typical and slow runs, then ask where the wait interrupts work.
A fast suite may still have weak coverage or unreliable results, while a slower suite may belong at a later pipeline stage. Prioritize earlier feedback where it helps developers without weakening the checks.
9. Throughput
Throughput counts completed work items of one defined type per period. The Kanban guide calls this delivery rate, with examples such as features per week. Keep the unit stable, because pull requests, issues, and customer features are not interchangeable.
Use historical throughput for forecasting, separating work types when their delivery paths differ. If the count rises, check WIP, lead time, stability, and rework to understand the change.
Smaller items can raise throughput without increasing the amount or importance of delivered work, so a higher count alone does not prove greater productivity.
10. Work in progress
Work in progress, or WIP, counts active items inside a defined workflow or stage at a point in time. Kanban distinguishes this observation from a WIP limit, which is a policy constraining how much work may be active.
WIP helps expose queues, multitasking pressure, and work that has started without finishing. Break it down by stage and age so that five newly opened items do not look identical to five items that have waited for weeks.
Low WIP can also reflect too little incoming work, an upstream blockage, or an overly restrictive policy. Check throughput and lead time before deciding to change a limit.
Investment and planning metrics
11. Engineering allocation
Engineering allocation divides effort in each work category by total classified effort during a period. Define locally relevant categories, such as product development, reliability, maintenance, technical debt, and support.
Compare allocation with priorities to inform planning. If reliability is a priority but receives little capacity, investigate that gap while keeping categories, classification coverage, and the effort proxy consistent across periods.
A Microsoft survey of 484 developers in June and July 2024 associated larger gaps between ideal and actual workweek allocation with lower perceived productivity and satisfaction. However, it used self-reports from one company in India and the United States, with an 8.06 percent response rate, and did not establish causation.
The research does not establish an ideal mix for every team. Interpretation also depends on classification coverage, since mentoring, collaboration, and interruptions can consume capacity without fitting cleanly into a category.
12. Planning accuracy
For a Scrum team, an optional planning accuracy measure divides completed forecast work by forecast Sprint work, using the same unit in both parts. Document how added, removed, split, or renegotiated scope affects the calculation.
Use the trend to calibrate forecasts by examining unplanned work, scope changes, dependencies, and capacity assumptions before attributing carryover to poor estimation.
The Scrum Guide treats selected work as a forecast, permits scope renegotiation as the team learns, and makes the Sprint Goal the commitment. Completing 100 percent of the forecast should therefore not become a universal performance target.
For continuous Kanban flow without a Sprint forecast, use throughput, WIP, aging, and lead-time distributions instead.
Reliability and developer experience metrics
13. SLO compliance and error budget
SLO compliance compares a service level indicator, or SLI, with its service level objective, or SLO, over a defined window. For a request-based SLI, divide requests meeting the success or latency criterion by all eligible requests, then compare that proportion with the target.
For this request-based SLO, the error budget is the permitted share of requests that miss the criterion, calculated as one minus the target. Use budget consumption to inform release and reliability decisions, without treating absolute reliability as the goal.
Check whether the indicator reflects user experience. Server-side latency can miss browser or client delays, so user reports may reveal a gap that requires revising the SLI.
14. Developer satisfaction
Track developer satisfaction with recurring, specific survey questions, keeping the scale, population, and cadence stable. Ask about work the team can improve, such as build feedback, documentation, cognitive load, focus, or goal clarity.
SPACE and DevEx support reading developer perceptions alongside workflow data, since each can reveal friction the other misses.
Research by Storey and colleagues describes a bidirectional relationship between satisfaction and perceived productivity, influenced by technical, social, and contextual factors. This does not establish that one causes the other, so satisfaction should not substitute for delivery performance.
Interpret changes with the team to choose follow-up work, and aggregate responses into groups large enough to protect anonymity. An eNPS score alone cannot explain what to change.
15. AI-assisted delivery impact
AI-assisted delivery impact compares outcomes for similar work, including subsequent review effort, rework, and reliability costs. Usage and generated-code counts describe adoption, but do not establish productivity or return on investment.
Define the measurement before rollout by following these steps.
- Name a task class and hypothesis, such as reducing cycle time for bounded maintenance work.
- Capture a baseline and record repository, staffing, task mix, and process conditions.
- Compare similar AI-assisted and non-assisted work when attribution is reliable, or use a carefully bounded comparison before and after adoption when it is not.
- Read speed or throughput with review load, rework, change failures, service reliability, and developer feedback.
The 2025 DORA research drew on nearly 5,000 technology professionals and more than 100 hours of qualitative data. It associated higher AI adoption with higher delivery throughput and instability, but its aggregate, correlational findings do not prove causation for a particular team.
Results can vary by task, and faster authoring can shift the queue to review or testing. For evidence and evaluation designs, see the analysis of AI and developer productivity.
How to choose the engineering metrics to start with
Choose a delivery outcome for the decision at hand, process measures to investigate it, and checks for costs to quality or developer experience. Google SRE similarly recommends a small set of representative indicators, without prescribing a fixed number.
| Leadership decision | Outcome | Diagnostic signal | Quality check or context |
|---|---|---|---|
| Investigate slower production changes | Change lead time | PR cycle stages, build and test feedback | Change fail rate |
| Identify whether review is the current queue | PR cycle time | Time to first review, WIP in review | Review quality signals, developer feedback |
| Assess whether active work exceeds capacity | Throughput and lead time | WIP by stage and age | Work-type mix |
| Check whether planned investment reaches the intended work | Engineering allocation | Unplanned-work share | Developer satisfaction, product priorities |
| Decide whether the service can accept more release risk | SLO compliance and error-budget consumption | Failed deployment recovery time | User impact and incident severity |
| Evaluate whether AI improves delivery | Comparable cycle time or throughput | Review load and rework | Change fail rate, SLOs, developer satisfaction |
Assign an owner, review cadence, and action rule to each metric. A hypothetical team might investigate after the upper tail rises for three periods, but that trigger must fit its context.
Remove metrics when they no longer inform a decision, since a migration or AI pilot may justify temporary tracking without creating a permanent KPI.
Engineering metrics that should not evaluate individual developers
Do not turn commits, lines of code, pull-request counts, story points, AI usage, or the metrics in this guide into individual rankings. Activity counts omit work such as mentoring, coordination, and risk reduction, which the multidimensional SPACE model and the DevEx approach help account for.
Keep delivery and flow metrics at team, service, or workflow level to investigate constraints. Individual comparisons cannot support an evaluation when they fail to account for task type, complexity, role, collaboration, and opportunity.
Using metrics as personal targets can encourage people to split work, inflate activity, or avoid difficult tasks. Treat an unexpected number as a reason to investigate the work and its context.

