On a hypothetical Tuesday morning, six pull requests are open while two engineers who usually review them handle a production incident. A feature planned for the sprint also depends on an answer from another team, so even though everyone has legitimate work in front of them, very little is reaching “done”. This is how slow development can hide inside a busy week.

In this article, 100% utilization means that all of the team’s available capacity has already been committed to planned work. The term refers to those commitments rather than hours logged, visible activity, or the traffic intensity used in a queueing model. Response capacity, in turn, is the room left to review, unblock, coordinate, recover, and finish work when demand changes.

Architecture, technical debt, unclear priorities, and weak tooling can all slow development, but this article focuses on what happens when a team pre-assigns all available capacity. Because ordinary variation then becomes harder to absorb, work can accumulate in queues, leaving more items in progress and extending cycle time until commitments slip.

Full utilization creates a queue, not extra delivery capacity

Once all available capacity is spoken for, an interruption goes into a queue or into someone else’s half-finished work.

Queueing theory offers a useful analogy for this congestion. In Kingman’s work on a single-server queue in heavy traffic, waiting grows as traffic intensity approaches one. An engineering team involves several people, different skills, uneven work, shifting priorities, and external dependencies, which makes it more complicated than that model. Kingman therefore helps explain the direction of congestion under variability without providing a staffing formula for software teams.

The analogy is useful because reviews arrive at uneven intervals, incidents interrupt planned work, and dependencies can clear quickly or sit for days. When the people able to respond are already committed, new demand either waits or displaces something underway. In the first case, elapsed time grows as work sits in a queue, while in the second, more work stays open and the team has more context to recover later.

The official Kanban guide addresses this problem by recommending that teams manage flow, limit WIP, and pull new work when capacity becomes available. This gives managers a way to decide when the system can accept more work and when it needs help finishing what is already underway.

Capacity left unassigned helps only when the team can direct it to the part of the flow that needs support. A spare hour cannot fix a queue if the person available lacks the skills or authority to review, respond to an incident, or unblock a dependency, so the planning decision also needs to account for skill coverage and decision rights.

Busy developers can hide a slow delivery system

On that Tuesday, the engineers handling the incident are busy, as are the authors of the six pull requests, who may have started more work while waiting. Meanwhile, the dependency owner is occupied elsewhere, leaving the planned feature blocked. Each person’s activity can make sense on its own while the shared delivery outcome gets worse.

Assignments are easy to see in a planning meeting, which makes utilization an appealing way to describe capacity. Yet the queue often sits between those assignments, after coding, before review, or beside a blocked dependency, without a single owner or a line on the capacity spreadsheet. By the time that waiting becomes visible as a missed date, adding more tasks to individual plans has made it harder to see where delivery first slowed.

While utilization tells managers whether capacity is assigned, flow shows what is moving, what is waiting, and what can finish next. Once delivery slows, that second view gives the team a more useful starting point for deciding where help is needed.

This difference also changes how the team should read its metrics. The SPACE framework treats developer productivity as multidimensional, so a single activity measure cannot represent the whole system. Reading cycle time and WIP at the team or process level helps locate review queues and blocked handoffs while keeping the investigation focused on delivery. For a deeper treatment, see why cycle time is useful when broken into stages and the guardrails for using metrics for teams rather than individuals.

When a delivery date slips, asking where the item waited gives the team a concrete part of the process to investigate. Opening with “who needs to move faster?” can instead direct attention toward individuals before anyone understands the source of the delay.

What consumes the team’s room to respond?

Planned features share capacity with code review and the work required to clear blockers, often drawing on a small set of people with the right context. Even a blocker that needs only an approval, an access change, an environment fix, or five minutes with a specialist can wait when that person’s calendar is already full.

Other demands change the plan after work has started. Production incidents take priority immediately, while bugs, customer escalations, and urgent requests compete with existing commitments regardless of whether the sprint has space. At the same time, a dependency can leave an active item stranded even when its owner is busy with legitimate work elsewhere.

When unplanned work arrives, making the trade-off visible helps the team decide whether to absorb a small item with available capacity, replace a lower-priority commitment, or defer the request. The choice depends on the situation, but adding everything requires acknowledging how the original plan will change. The DevStats guide to handling unplanned work develops that operating model further.

Allocation describes where capacity is intended to go, while utilization describes how much appears consumed. Keeping those meanings separate helps explain why a roadmap can be fully staffed even as bugs, maintenance, and support draw on the same people and change where their capacity actually goes.

Extra buffer may help the team absorb a rare incident, but recurring demand calls for a closer look at its source. A review queue that returns every week suggests investigating review ownership or batch size, while a service that repeatedly interrupts the sprint needs reliability work. Slack can accommodate variation as the team investigates, but leaving the underlying problem unresolved allows it to keep consuming that capacity.

Use flow signals to choose the response

The visible symptom should guide the first intervention, so the table below connects flow signals with plausible explanations and actions to try. These are working hypotheses that narrow the investigation, and the team still needs to check them against its own process.

Signal Plausible explanation First move Check afterward
Review queue keeps growing Review capacity or ownership is constraining flow Pause some new starts, make review work explicit, and share context where possible Pickup and review wait, plus the queue trend
WIP rises across several stages More work is entering than the system can finish Stop pulling new items temporarily and finish open work WIP and end-to-end cycle time
Unplanned work returns every sprint The plan understates demand, or one source keeps recurring Triage, trade scope explicitly, then investigate the source Planned versus unplanned work and commitment changes
Cycle time rises while WIP stays level One stage, dependency, or work type may have become more variable Inspect stage timing and blocked items before changing team load Which stage expanded, for which work
Incidents and rework rise Instability is consuming downstream capacity Stabilize the affected service and revisit near-term scope Throughput alongside instability

WIP limits put this policy into practice by allowing new work to enter only when the relevant part of the system has capacity. The number on the board is therefore a constraint on how much work can remain active. If changing it reduces waiting in one stage but moves the queue elsewhere, the team needs to investigate that new constraint before treating the experiment as a success.

When an urgent item arrives

Before changing the current commitment, check whether the new item really needs immediate action. If it can wait, place it in the next commitment, but if it cannot, review the open work and identify what will move to make room.

A small request may fit when someone can take it without abandoning reviews or blocked work. If that capacity is unavailable, replacing a lower-priority item makes the trade-off explicit, and recording the change lets stakeholders see why the original commitment shifted.

Taking a few minutes to record the decision keeps the forecast aligned with the work the team has agreed to do. Otherwise, the sprint can grow while stakeholders continue to rely on a forecast based on the earlier scope.

When review is controlling delivery

Code waiting for review is already inventory, so starting another feature adds more work to the system while the existing queue remains unresolved.

Once the review queue is visible, the team can decide how to help the work move through it, whether by pairing on a difficult change, giving review explicit priority, sharing context with another reviewer, or reducing the next batch. Follow the effect across several items, because one fast merge can be luck. A sustained reduction in pickup and review wait gives a better indication that the change helped.

Decide what the clock measures

Cycle time becomes ambiguous when different clocks share the same name. For the staged PR view in this article, PR cycle time covers the elapsed path across coding, pickup, review, merge, and deployment, although a team may choose different start and end events. Documenting those boundaries makes it possible to compare equivalent measures and distinguish them from issue cycle time, which follows the issue workflow, or DORA change lead time, which runs from commit to production.

Start with WIP and stage-level cycle time to find accumulation and waiting, then read throughput alongside instability to see how delivery outcomes are changing. DORA groups software delivery performance into those two factors and recommends applying the measures in the context of a particular application or service, as explained in DORA’s current metrics guide.

In a second hypothetical scenario, the team splits work into smaller changes and coding time falls, but pickup wait doubles because reviewers now receive more pull requests. Aggregate cycle time may barely move, while the stage view reveals how the improvement upstream increased demand downstream. This points the next decision toward reviewer coverage, review policy, or the rate at which new work starts, because asking authors to code faster would leave the review constraint unresolved.

Reading the stages and outcomes together helps the team judge whether a local improvement benefits delivery as a whole. A shorter coding stage is less useful when the review queue expands, just as higher throughput needs closer investigation when rework and incidents rise alongside it. The metrics reveal the pattern, but the team still needs to investigate what caused it.

To interpret those changes, compare the team’s current flow with its own history across similar work before borrowing a threshold from another service or organization.

The relationship between these views is covered in more depth in flow metrics versus DORA metrics.

Slack is a policy you tune

Intentional slack preserves some ability to respond when the week stops matching the plan. Since incident load, review concentration, dependencies, and incoming-work volatility vary across teams, the amount of response capacity each one needs cannot be reduced to a universal percentage.

Treat that policy as an experiment by choosing the queue that is hurting delivery, changing how work enters or how help reaches that stage, and watching the effect on WIP, cycle time, throughput, and stability. If the queue moves elsewhere, follow it to understand the next constraint. If nothing changes, investigate whether the restriction was somewhere else to begin with.

At the next planning meeting, ask what must remain finishable when this plan changes.