Man wearing a black cap, round glasses, and a dark green t-shirt standing with hands on hips against a wall with vertical black slats illuminated by teal and purple lights.

Your retrospective isn't broken. Your improvement system is.

If you have spent any time on an engineering floor in the last two years, you have heard a version of the same complaint. "We do retros every sprint. Nothing actually changes." The reflex response from coaches and managers is to fix the meeting. Try a new format. Run a Liberating Structure. Use a retrospective game. Bring in a fresh facilitator.

The peer-reviewed evidence on what actually happens when teams do that is unflattering. Format changes produce small gains, sometimes none. The constraint sits downstream of the meeting, in a system most teams do not have. That system is what this article is about.

The one finding that reframes the conversation

In 2024, Ng and colleagues published a paper at the ACM Symposium on Applied Computing with a title that telegraphs its argument: Implementing Action Items Over Improving the Format of Retros. They ran the established Przybyłek 2022 retrospective-game protocol in two Scrum teams at Carrier Poland. Both teams had documented low retrospective enthusiasm going in. The intervention is well-evidenced. Przybyłek's original action research across six Scrum teams showed measurable improvements in conversation quality. Ng's teams were running the same recipe.

The result, in their words: "the observed benefits were considerably more modest compared to the results of previous studies. Our key conclusion is that without active involvement in implementing action items, improvements will not materialize, regardless of the meeting format."

That sentence is the article. The format is not where the loop breaks. The implementation of what came out of the format is where the loop breaks.

You can corroborate this through a completely different evidence channel. Dantas and colleagues, at CHASE 2025 (co-located with ICSE), interviewed 15 practitioners across diverse organisations about how retrospectives handle community smells: Lone Wolf, Organizational Silo, Radio Silence, Black Cloud. Identification of the smells in retrospectives works well. Strategies for refactoring are formulated. Their implementation and monitoring remain inconsistent.

Two different research designs, two different question framings, the same shape of finding. The detection layer works. The conversion layer breaks.

Why the meeting still matters, and where its ceiling is

This is not an argument against retrospectives. The Verwijs & Russo Theory of Scrum Team Effectiveness, validated against around 5,000 developers and 2,000 Scrum teams in ACM Transactions on Software Engineering and Methodology, places continuous improvement as one of the five structural factors that drive Scrum team effectiveness. Continuous improvement does not happen without the inspect-and-adapt step. The retrospective is the load-bearing meeting for that step. Cutting it is not the move.

But the meeting has a ceiling. Hundhausen and colleagues (ICSE-SEET 2024) coded 963 retrospective statements from 32 undergraduate software-engineering teams across four courses at two universities. Their headline number is uncomfortable: only 13% of the teams' reflections provided justification for a strategy to be stopped, continued, or started. The setting is undergraduate, so do not quote this as an industry parameter. Read it as a ceiling indicator. Even when reflection runs, the supporting reasoning behind the resulting action is thin in the large majority of cases.

An action item without a justification has nothing to defend itself with when next sprint's delivery pressure arrives. Przybyłek named the same dynamic from the other end: the Sprint Retrospective is "an agile practice most likely to be implemented improperly or sacrificed when teams perform under pressure to deliver." The retrospective is fragile by design. So is its output.

So the meeting matters. Better facilitation matters. Better inputs matter. Erdoğan and colleagues showed in 2018 that bringing statistical analysis of velocity, estimation, defect density and quality-effort distribution into the retrospective measurably improved estimation accuracy and product quality on a real industrial site. None of that, on its own, closes the loop.

The loop, drawn small

Sprint ends. Six steps to observable improvement two sprints out.

  1. Input data exists from the closed sprint: Jira or ADO state, cycle-time distribution, escaped defects, decision-wait age, blocked-state durations, review queue age.
  2. Retrospective runs. Reflection is produced.
  3. Action items are formulated, with or without justification.
  4. Action items are owned, sized, and accepted into the next sprint's capacity.
  5. Action items are executed and their effect monitored against the data from step 1.
  6. The result re-enters the next retrospective.

If you map the evidence onto this loop, the breaks are clear. Erdoğan and colleagues point at step 1 to 2. Hundhausen points at step 2 to 3. Ng et al. point at step 3 to 4. Dantas points at step 4 to 5. Format-of-the-meeting interventions, the kind most teams reach for first, live almost entirely between steps 2 and 3. That is why they produce small gains alone.

The interesting work is at steps 3, 4, and 5. That is also the work that lives outside the meeting.

The small downstream system that closes it

You do not need a programme to fix this. You need four small commitments that almost no team I have seen in flow-audit work actually carries.

One. The action item arrives in the next sprint backlog with a name, a size, and a measurable closure condition. Not "we will improve our PR review process." A real Jira/ADO item, sized like any other piece of work, owned by one person (not "the team"), with a definition of done that does not depend on a feeling. If the closure condition cannot be written down, the action item is not ready. Park it; it will come back. The Beek et al. 2023 design-science work already proposes "number of solved retrospective items after a new sprint" as one of the candidate objective measures of the Continuous Improvement effectiveness factor in Scrum. The signal is trackable. The teams that look at it are rare.

Two. The action item gets capacity in the sprint that follows the retrospective, not "if we have time." This is where most loops actually die. The team agrees on three action items. The next sprint is full of feature work. The action items inherit zero capacity. By sprint end they are still open. By sprint 3 they are quietly dropped and the retrospective writes new ones. The conversation continues; nothing changes. The minimum viable commitment here is a small reserved slice. Pick a number you can defend; some teams do 5% of capacity, some less, for retrospective work. Not aspiration. Capacity in the plan.

Three. The decision-rights for the action item are pre-agreed. Most retrospective action items end with "we will talk to the PO" or "we will bring this up with architecture." That is not an action item. That is an intent to start a meeting. The decision-rights question, who can approve this change without escalation and what is the budget envelope, has to be settled before the action item leaves the room, or it joins the decision-wait queue this same blog called out a fortnight ago, and your action conversion rate will collapse for the same reason your cycle-time right tail did.

Four. The measurement re-enters the next retrospective as an input, not as a vibe. "Did we improve PR review wait time?" is answered with the Jira/ADO query, not with a show of hands. If the team cannot measure it, the action item was probably not real to begin with.

None of this requires a tool. None of it requires a framework change. It requires somebody, a Scrum Master, a Delivery Manager, a senior engineer with appetite, to own the downstream loop the way a Product Owner owns the upstream backlog. In most teams nobody owns it. That is the gap.

The Monday-morning instrument

You can read your own loop in an hour. No special tooling.

  • Pull the last six retrospectives. Look at the action items written.
  • For each action item, classify: was it accepted into the next sprint with a size and owner; was it closed by the next retrospective; did its result re-enter the next retrospective as evidence.
  • Count three numbers. Total action items written. Action items closed by the next retrospective. Action items whose effect was measured in the data of the next sprint.
  • The first ratio (closed / written) is your action-conversion rate. The second ratio (measured / closed) is your evidence-loop rate.

You will know the answer before you finish counting. Most teams' action-conversion rate over six sprints lands somewhere between "we don't know" and a number low enough that you would not bet on it surviving disclosure. Most teams' evidence-loop rate is, in practice, zero.

That is not a moral indictment of those teams. It is a delivery-system fact, in the same family as cycle-time variance or decision-wait age. It is also one of the cheapest signals to move once it is visible. A team that goes from a 20% action-conversion rate to a 60% rate in two months is doing work the rest of its delivery metrics will eventually feel.

What the evidence does not yet say

Three caveats worth holding while reading any current claim about retrospective effectiveness, including this one.

The strongest mechanism-level paper here, Ng et al. 2024, is two teams at one Polish industrial firm. It is peer-reviewed and the mechanism statement is clearly worded, but the effect size is qualitative. Dantas et al. 2025 is 15 practitioner interviews. Hundhausen et al. 2024 is 963 statements from undergraduate teams, an illustrative ceiling, not an industry parameter. Przybyłek 2022 is six teams across action-research cycles. There is no quantitative panel-data study of industrial Scrum action-item conversion rates that I have been able to find. Practitioner content circulating a "70 to 80% of action items never get implemented" number traces to no primary study and should be ignored.

All cited studies are international; none specifically DACH. The mechanism should transfer. The specific numbers should not be quoted at a DACH client.

This is also a problem the Theory of Scrum Team Effectiveness (Verwijs & Russo 2023) supports structurally, continuous improvement is one of five validated factors, but does not measure at the action-item level. The link between the action-conversion rate and downstream delivery metrics (cycle time, escaped defects, decision-wait age) is mechanistically plausible and Alami et al. 2022 supports it through the inspection-and-adaptation route, but it has not yet been quantified in a peer-reviewed industrial panel.

If anyone reading this has access to a few hundred industrial retrospective records, the empirical study that closes that gap is sitting there waiting. It would land.

The Monday-morning question

Not "is our retrospective good." That is a question for an essayist. The question for a Delivery Manager or freelance Coach this week is small and specific.

Of the last six retrospectives my team ran, how many of the action items written were closed before the next retrospective, and how many had their effect measured against the data from the sprint that produced them?

The two numbers are an hour of work to pull. They will tell you, with more honesty than any meeting-format experiment, where your team's improvement loop is actually breaking. And they will give you a small, defensible lever to move it that has nothing to do with whether your retrospective is "engaging."

The retrospective was never the problem most teams thought it was. The system around it was.

Sources

  • Alami, A., & Krancher, O. (2022). How Scrum adds value to achieving software quality? Empirical Software Engineering, 27(7), Article 165. doi.org/10.1007/s10664-022-10208-4
  • Beek, K., Wagenaar, G., Kester, L., Overbeek, S., & de Rooij, E. (2023). Measuring team effectiveness in Scrum. In J. M. Fernandes, G. H. Travassos, V. Lenarduzzi, & X. Li (Eds.), Quality of information and communications technology (pp. 233–247). Springer. doi.org/10.1007/978-3-031-43703-8_17
  • Dantas, C., Massoni, T., Sarmento, C., Rocha, R., & Gualberto, D. (2025). The role of the retrospective meetings in detecting, refactoring and monitoring community smells: A qualitative study. In 2025 IEEE/ACM 18th International Conference on Cooperative and Human Aspects of Software Engineering (CHASE) (pp. 102–113). IEEE. doi.org/10.1109/CHASE66643.2025.00021
  • Erdoğan, O., Pekkaya, M. E., & Gök, H. (2018). More effective sprint retrospective with statistical analysis. Journal of Software: Evolution and Process, 30(5), Article e1933. doi.org/10.1002/smr.1933
  • Hundhausen, C., Conrad, P., Tariq, A., Pugal, S., & Zamora Flores, B. (2024). An empirical study of the content and quality of sprint retrospectives in undergraduate team software projects. In Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training (pp. 104–114). ACM. doi.org/10.1145/3639474.3640074
  • Ng, Y. Y., & Kuduk, R. (2024). Implementing action items over improving the format of retros. In Proceedings of the 39th ACM/SIGAPP Symposium on Applied Computing (pp. 853–855). ACM. doi.org/10.1145/3605098.3636174
  • Przybyłek, A., Albecka, M., Springer, O., & Kowalski, W. (2022). Game-based Sprint retrospectives: Multiple action research. Empirical Software Engineering, 27(1), Article 1. doi.org/10.1007/s10664-021-10043-z
  • Verwijs, C., & Russo, D. (2023). A theory of Scrum team effectiveness. ACM Transactions on Software Engineering and Methodology, 32(3), Article 74. doi.org/10.1145/3571849
  • Wolpers, S. (2024, October 21). Ditch the unfinished action items. Scrum.org. scrum.org/resources/blog/ditch-unfinished-action-items