Man wearing a black cap, round glasses, and a dark green t-shirt standing with hands on hips against a wall with vertical black slats illuminated by teal and purple lights.

Code got cheaper. Decisions didn't.

If you run a delivery org, the dashboard probably looks healthier than the team feels. PR throughput is up. Time-to-first-draft has come down. Stand-ups feel calmer. And yet the same sprint keeps ending with three items that should have been done last Tuesday and weren't, for reasons nobody can summarise in one sentence.

The interesting reason is structural, and a year of AI rollout has made it sharper, not gentler.

What the numbers actually show

Plandek's 2026 Engineering Productivity Benchmarks pulled delivery data from 2,000+ engineering teams. Vendor source, methodology behind a download form, not an audited benchmark. Top-quartile teams complete more than two-thirds of the work they plan in a sprint. Bottom-quartile teams complete under half. The gap is not subtle. The interesting part is that it is wider than the AI-acceleration gap, which Plandek reports at roughly four-fold between strong and weak performers on lead time gains.

DORA's 2025 State of AI-Assisted Software Development, based on ~5,000 respondents, lands in the same neighbourhood: AI amplifies whatever the underlying engineering system already is. Mature systems convert AI into throughput. Less mature systems lose stability.

The common reading is that less mature systems carry technical debt, weaker testing, and shakier integration. All true. The reading that gets less attention is that less mature systems usually also have a more clogged decision queue, and AI hits that queue harder than it hits the code.

The decision queue is a real queue

Reinertsen's queueing argument from 2009 still does most of the work here. Wait time at any resource grows non-linearly with utilisation. Once a stakeholder, product owner, architect, or governance reviewer crosses around 80% utilisation, queue length and wait time grow disproportionately. Pre-AI, this was already the dominant driver of right-tail cycle time in many teams. AI changes three things at the decision queue at once.

Arrival rate goes up. Cheap code generation lets engineers propose more options, more variants, more PRs, more design alternatives per unit time. Each of these is a potential decision.

Service rate stays roughly flat. A product owner still makes product decisions at human speed. A security reviewer still reads the change. A senior architect still asks the same question they would have asked in 2022.

Variance goes up. AI-generated work has wider scope and wider review-effort variance, so the input to the decision queue is more uneven than it used to be.

That is the textbook recipe for a queue that grows.

What does an empirical signal of this look like

Wiesche's Interruptions in Agile Software Development Teams (Project Management Journal, 2021) gives the closest empirical anchor we currently have. Grounded-theory study of four agile software teams across Scrum and Kanban. Three categories of interruption came out of the analysis: programming-related work impediments, interaction-related interruptions, and external-environment interruptions.

Inside the first category, alongside malfunctioning code and missing information, sits a recognisable mechanism: "distributed decision-making causes discrepancies and additional work." Teams coped through better information retrieval practices, reduced team dependencies, and a scrum master role explicitly used as a buffer against unnecessary interruptions.

Wiesche's study is pre-AI, qualitative, and four teams. Treat it as a mechanism source, not as a quantified causal claim. What it does establish is that decision-flow problems are an empirically observed interruption class in agile teams, not a freelance-coach talking point. The mechanism predates AI. AI raises its weight.

Ali et al.'s 2024 systematic review of waiting times in business processes (Business & Information Systems Engineering) makes a related point from the BPM side: waiting-time analysis has a workable taxonomy with at least three dimensions (purpose, cause, measure), and one of the recognised causes is exactly prioritisation and decision wait, extractable from execution-log data. The translation to a software-delivery context is by analogy, not by replication, but it confirms that "decision wait" is a measurable category, not a vibe.

Why your dashboard doesn't show this

The decision queue is the easiest queue in the system to lose track of. Three structural reasons.

PR review tools rarely capture it. The item is parked outside the review platform once it hits "waiting on PO" or "needs architectural decision." So it stops accruing measurable review-queue age. It accrues unmeasured decision-queue age instead.

Jira and ADO bury it. "Blocked," "needs input," "on hold," "follow-up question," and "rework" rarely sum to a single queue-age distribution unless someone explicitly extracts them.

Standups don't surface it. "We're waiting on Markus" gets said at standup for three days and disappears into the third week of the quarter. Nobody on the team sees the histogram.

The net result is that the cycle-time distribution carries the cost of a decision queue that nobody plots.

The five moves a Delivery Manager can make this sprint

None of this requires a new tool. All of it can start tomorrow.

1. Pull the decision-wait histogram. From Jira or ADO, extract every state change into states that mean "waiting for a person to decide": blocked, needs input, awaiting product decision, awaiting architecture review, pending governance approval. Sum the time in those states per item over the last 50 delivered items. Look at the median, the 85th, and the 95th. The 85th is your real predictability cost.

2. Identify the binding decision-maker. Usually one to three people per team. The product owner, an architect, a domain expert who has to sign off, a stakeholder who owns the "what does this even need to do" question. Their utilisation is the cap on your team's decision throughput. If they are at 90%+ on Mondays and Tuesdays, the queue you measured in step 1 will not shrink without changing how decisions reach them.

3. Reduce the arrival rate at the decision queue. Most teams underestimate how many "do you want this in or out" questions could be batched into a single weekly slot with a written default answer. Cheap code generation pushes the cost of asking down, which means more questions get asked. That is fine for code. It is not fine for the stakeholder. Intake discipline at the decision queue is the same idea as intake discipline at the AI-cost queue from three weeks ago, applied one layer up.

4. Increase decision throughput honestly. Two real options. Delegate, by giving the team explicit decision rights for a category of decisions, with the conditions written down. Or schedule, by reserving decision blocks on the calendar of the binding stakeholder. "Be faster" is not a third option. It does not survive the queueing equation.

5. Audit the slow items, not the fast ones. Pick the three items with the longest decision-wait in the last 30 items. Read the history. What stretched them: an unavailable stakeholder, a contested scope question, a missing decision right, a mismatched authority between product and engineering, a late-arriving compliance question? Each cause maps to a different next move. The audit takes thirty minutes. Most teams skip it because the role that owns "decision flow as a metric" is not staffed in most orgs.

When the bottleneck isn't decisions

A useful red-team. The decision-queue framing fits some teams much better than others. In a platform team with stable mandates, the binding queue is more often integration or release. In a regulated environment, the binding queue may legitimately be governance, and the answer is policy change, not faster decisions. In a team where the codebase is the limiting factor, AI sits downstream of an architectural mess that no decision-flow work will fix.

The diagnostic is simple. Pull the slowest fifteen items of the last sixty and tag each one with where it spent the most time: coding, review, decision wait, integration, governance, customer feedback. The category with the largest sum is the binding queue. If decision wait wins, the framing in this article applies. If it doesn't, the team has a different problem, and the right answer is to fix that one.

What this is and isn't

This is not an argument that AI is the problem. The DORA data is fairly clear that mature systems convert AI into real throughput. Pulling AI out blanket-style loses that gain.

It is also not an argument for a heavier decision process, more meetings, or another framework. The opposite. The point of measuring the decision queue is so that the conversation moves from "we need more alignment," a phrase that has never improved a histogram, to "the 85th percentile of decision-wait is twelve days; the binding stakeholder is at 95% utilisation on Mondays; we are going to delegate this category of decision and book three decision slots on Tuesday afternoons."

That is a Delivery Manager conversation. AI did not invent it. AI made it the conversation that is hardest to avoid, because the cheap part of the work is now done and the queue moved.

Sources