Man wearing a black cap, round glasses, and a dark green t-shirt standing with hands on hips against a wall with vertical black slats illuminated by teal and purple lights.

Every team hit its target. The product still arrived late.

The quarterly report says analysis finished on schedule. Development kept utilization high. Security review stayed inside its SLA. Operations protected stability and took no unplanned outage. The change reached the customer six weeks late.

Nobody there is incompetent. Nobody is gaming numbers. Each function did what it was measured on, and the measurement was accurate. The problem sits one level up. The targets were attached to units of output. The work and the customer's outcome were interdependent.

What "performance" means depends on how the work is coupled

Courtright et al. (2015) meta-analyzed 107 independent samples covering 7,563 teams. They separated two kinds of coupling. Task interdependence means one unit's work physically requires another's, and it runs mainly through task-focused coordination. Outcome interdependence means units share the consequences, and it runs through interpersonal processes and cohesion. Different levers. Different interventions.

The consequence is narrower than the usual advice to use team goals. If product, engineering, security and operations must combine outputs before anything reaches a user, each function's local throughput number is an incomplete performance measure. That number describes component motion. It says nothing about system completion.

Where the target level and the work level diverge

Kleingeld et al. (2011) meta-analyzed group goal setting across 49 comparisons. Specific group goals were associated with meaningfully better group performance. Inside interdependent groups, individual goals aimed at maximizing personal performance showed a strong negative relationship with group performance. Individual goals aimed at maximizing one's contribution to the group showed a strong positive relationship.

Those subgroup findings are the most striking part of the paper and the least secure. They rest on only six and four comparisons. The evidence base is small, the uncertainty is substantial, and the paper is from 2011. The authors themselves note a shortage of recent organizational field studies. Read the findings as strong directional evidence for a mechanism, not as a number to carry into a business case.

A recent experiment points the same way. Shi et al. (2025) tested how relative performance information interacts with task interdependence. Without interdependence, both individual and team comparison beat no comparison. With interdependence, individual relative performance information produced competing coordination and competition mindsets and reduced performance, while team-level information improved coordination. Treat this as supporting evidence rather than the foundation of the argument: it is one recent experiment, while the broader evidence comes from the two meta-analyses above.

The mechanism is queues, not motivation

None of this requires anyone to behave badly. It only requires each function to respond sensibly to what it is measured on.

Analysts asked to maximize readiness produce a larger ready queue. Development measured on utilization increases the work waiting for review. A review group protecting its own SLA takes the small, well-formed slices quickly, while awkward exceptions age outside its clock. Operations minimizing change exposure batches releases, which lengthens the wait for everything in the batch.

Every local report stays green because each target stops at a handoff. Customer lead time is made of the queues between handoffs, the intervals nobody owns and nobody is measured on. Local optimization is invisible to a scorecard whose observation ends exactly where the delay begins.

Two different fixes, often confused

The first fix reduces avoidable coupling. DORA's guidance on loosely coupled teams recommends cutting dependencies that force teams to request permission from outside groups or coordinate in fine detail for routine changes (DORA, n.d.). Some coupling exists because the product requires it. Much of it comes from architecture decisions, permission design, or policy nobody has revisited. That part is removable.

The second fix aligns goals where coupling remains and cannot be removed: regulated approvals, shared platforms, physical hardware dependencies. Attach the target to the smallest application or service outcome that contains both the work and the feedback loop. DORA's metrics guidance recommends applying delivery metrics to one application or service at a time and warns about Goodhart effects and cross-team ranking (DORA, 2026).

Not every boundary is a problem. Bento et al. (2020) reviewed 40 studies, 20 of them empirical, and found that clusters can provide useful local reinforcement and support idea generation. They turn into silos when the barrier weakens cooperation the organization's goals depend on. The design question is whether a boundary contains the work required for an outcome, or cuts across a dependency that gets renegotiated every week.

A trace you can run on Monday

Pick one change from the last month that arrived late or came back for rework. One. Use event timestamps from Jira or Azure DevOps, not recollection. Memory compresses waiting time and inflates working time.

For each function that touched it, write down the target it was expected to protect, the event that marked the work locally "done," the queue that event created, and how long the item waited before the next useful action. Add the rework that crossed the boundary. Add the outcome that closed the loop for the customer.

Then read the trace against three questions. Did any local target reward starting, handing off, or refusing work where the system needed finishing? Did any function shorten its own elapsed time by pushing uncertainty downstream? Which of the dependencies you documented are real, and which are artifacts of architecture or permission design?

Where a target ends at a handoff, the replacement is rarely a single number. A service-level lead-time distribution paired with an instability or quality measure keeps speed honest. DORA's throughput and instability pairing is the worked example. A queue-age number is a good diagnostic and a bad ranking.

The honest limit: direct field evidence linking function-level targets to cycle time, handoff delay and rework in current software delivery is thin. What exists is strong general evidence that interdependence changes what performance means, moderate evidence on goal level, and DORA guidance pointing the same way. Enough to justify running the trace on your own system. Not enough to promise a number.

Most organizations already have the data. What they lack is one report that puts a function's target and the customer's waiting time on the same page.

References

Bento, F., Tagliabue, M., & Lorenzo, F. (2020). Organizational silos: A scoping review informed by a behavioral perspective on systems and networks. Societies, 10(3), Article 56. https://doi.org/10.3390/soc10030056

Courtright, S. H., Thurgood, G. R., Stewart, G. L., & Pierotti, A. J. (2015). Structural interdependence in teams: An integrative framework and meta-analysis. Journal of Applied Psychology, 100(6), 1825–1846. https://doi.org/10.1037/apl0000027

DORA. (n.d.). Loosely coupled teams. Retrieved August 12, 2026, from https://dora.dev/capabilities/loosely-coupled-teams/

DORA. (2026, January 5). DORA's software delivery performance metrics. https://dora.dev/guides/dora-metrics/

Kleingeld, A., van Mierlo, H., & Arends, L. (2011). The effect of goal setting on group performance: A meta-analysis. Journal of Applied Psychology, 96(6), 1289–1304. https://doi.org/10.1037/a0024315

Shi, L., Tafkov, I. D., & Zhou, F. H. (2025). Task interdependence in teams: How to structure relative performance information to improve team performance. European Accounting Review. Advance online publication. https://doi.org/10.1080/09638180.2025.2537650