Man wearing a black cap, round glasses, and a dark green t-shirt standing with hands on hips against a wall with vertical black slats illuminated by teal and purple lights.

Your delivery metric is lying to you

The adoption gap is real. The measurement gap is worse.

You've been in this room. The sprint review starts, someone shares a screen, and there it is: velocity holding steady at 42 points, burndown a clean diagonal, the dashboard glowing green. The product owner nods. The stakeholder on the call says "great progress." And you know, because you sat in three of the standups that week, that two of the "done" stories are done in the sense that they were demoed, not in the sense that anyone can use them. The thing on the screen and the thing that happened are two different stories.

That gap is the whole problem. Many organizations have closed the adoption gap. They run Scrum or some honest variant of it, they have a board, they track velocity, they may even report DORA numbers upward. What they haven't closed is the measurement gap: the distance between what the number says and what was actually delivered. And that gap doesn't close on its own. Left alone, it widens, because the act of measuring changes the thing being measured.

The proxy is not the thing

Every delivery metric is a stand-in. Velocity stands in for capacity. Burndown stands in for progress. Coverage stands in for test quality. Deployment frequency stands in for delivery maturity. None of these is the thing you actually care about, which is working software that solves a real problem for a real person. They're proxies, and proxies are fine, necessary even. You can't manage what you can't see, and you can't see "value delivered" directly, so you watch the shadow it casts.

The trouble starts when the shadow becomes the target. Charles Goodhart, writing about monetary policy in 1975, put it this way: any observed statistical regularity tends to collapse once pressure is placed on it for control purposes. The plain-English version most of us know: when a measure becomes a target, it stops being a good measure. Donald Campbell, a social psychologist, said the sharper version four years later. The more a quantitative metric is used for social decision-making, the more it gets gamed, and the more it corrupts the process it was meant to monitor. Campbell's Law is the one that bites in delivery, because delivery metrics almost always end up feeding social decisions: who looks good in the quarterly review, which team gets the headcount, whether the consultancy renews the contract.

That's the mechanism. It isn't malice. It's gravity.

What it looks like in delivery

Velocity is the obvious one. The moment velocity leaves the team and becomes something reported upward or compared across teams, the estimates start to drift. Point inflation: the same work that was a 3 last quarter is a 5 now, and the velocity line climbs while nothing about the team's actual output changed. Sandbagging in the other direction, where a team commits to less than it can do to keep the line flat and predictable. Mid-sprint re-estimation so the points completed match the points committed. The CodePulse practitioner write-up on Goodhart's Law in software catalogues exactly these moves, and if you've run sprint reviews you've watched at least two of them happen in real time.

Burndown gets gamed more quietly. Work gets parked or reclassified rather than finished, so the chart stays clean. Scope shrinks without anyone calling it a scope change, and the line slopes down on schedule. The chart looks like progress because someone adjusted what counts as remaining.

DORA metrics are not immune just because they're better metrics. Deployment frequency goes up when you slice changes thinner, not when you actually ship more value. Lead time improves when "started" gets redefined to a later point in the process. Change failure rate drops when failures get filed as planned maintenance. The Accelerate research behind DORA is survey data about what high-performing organizations correlate with, and it explicitly warns against using these numbers to evaluate individual teams. Which is, of course, exactly what organizations do, and the gaming follows.

Coverage is the cleanest demonstration of the whole law. A test that calls a function and asserts nothing raises the coverage number and tests nothing. The proxy goes up, the thing it stood for goes down, and the dashboard can't tell the difference.

Be honest about what we actually know

Here's where I have to be straight with you, because the anti-pattern in posts like this is to dress practitioner folklore up as science. The mechanism, Goodhart and Campbell, is well-established in economics and organizational psychology. That part is solid. The specific claim that software teams game velocity, burndown, DORA, and coverage in the ways I just described is mostly practitioner-reported. It comes from blog posts, Scrum Master forums, and delivery practitioners describing recurring patterns. It is not, as far as I can find, peer-reviewed.

What the academic literature does support is the upstream condition. Pasuksmit and colleagues' 2024 systematic review in ACM Computing Surveys on effort estimation in agile identifies five sources of estimation inaccuracy, and one of them is business influences: external pressure shaping the estimate. That matters, because it means estimation isn't only a technical calibration problem; it's a social one. Pressure distorts estimates. Goodhart's Law is what you get when you apply pressure systematically and call it a target. The review documents the pressure. The gaming is the predictable consequence, but the consequence itself hasn't been measured at the level a skeptic would want. Hold both of those at once.

What to do in your next review

The instinct is to find an ungameable metric. There isn't one. Any number used for control eventually bends. So the move isn't a better proxy. It's changing how the proxy is used. A few things that actually hold up:

  • Keep team metrics inside the team. The single biggest driver of gaming is a metric that leaves the room and feeds a performance or budget decision. Velocity is a planning tool for the people doing the planning. The moment it goes on a slide for someone who can't see the work, it starts to lie.
  • Separate estimation from commitment, and both from evaluation. If the number you estimate with is also the number you're judged on, you've built the incentive to distort it. Break that link and most of the pressure Pasuksmit describes has nowhere to go.
  • Watch outcomes, not output. "Did this change reduce the support tickets" is harder to game than "did we ship 40 points," because you can't fake a problem going away.
  • Trust the leading indicators you can feel. Work-in-progress, blocked time, how long a review sits open, escape rate to production. These are harder to dress up than a velocity total and they tell you something sooner.
  • When a number stops moving, get suspicious, not pleased. A suspiciously stable velocity is often a managed velocity.

The practical test for your next planning session is one question, asked out loud: what decision does this number feed, and who makes it? If the answer is "the team, to plan its own next two weeks," keep it. If the answer is "someone upstairs, to decide who's performing," you already know which way it's going to bend. The dashboard was green. That was always the easy part.

Sources

  • Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90. doi.org/10.1016/0149-7189(79)90048-X
  • Fernández-Diego, M., Méndez, E. R., González-Ladrón-De-Guevara, F., Abrahão, S., & Insfran, E. (2020). An update on effort estimation in agile software development: A systematic literature review. IEEE Access, 8, 166768–166800. doi.org/10.1109/ACCESS.2020.3021664
  • Forsgren, N., Humble, J., & Kim, G. (2018). Accelerate: The science of lean software and DevOps: Building and scaling high performing technology organizations. IT Revolution Press. itrevolution.com/product/accelerate
  • Goodhart, C. A. E. (1975). Problems of monetary management: The U.K. experience. In Papers in monetary economics (Vol. 1, pp. 1–20). Reserve Bank of Australia. econbiz.de/10002525062
  • Pasuksmit, J., Thongtanunam, P., & Karunasekera, S. (2024). A systematic literature review on reasons and approaches for accurate effort estimations in agile. ACM Computing Surveys, 56(11), 1–37. doi.org/10.1145/3663365
  • Russell, A. (2026, January 8). Goodhart’s law in software: Why your metrics get gamed. CodePulse. codepulsehq.com/guides/goodharts-law-engineering-metrics
  • Usman, M., Mendes, E., Weidt, F., & Britto, R. (2014). Effort estimation in agile software development: A systematic literature review. In Proceedings of the 10th International Conference on Predictive Models in Software Engineering (PROMISE ’14) (pp. 82–91). ACM. doi.org/10.1145/2639490.2639503