
What Should Your Agile Coach Be Accountable For?
A contract ends in six weeks. Someone in finance asks what the coaching line item bought. The Head of Delivery looks at two quarters of teams that run every ceremony, name their impediments out loud, and still miss committed dates. Nobody thinks the coach has been lazy. Nobody can say precisely what changed either.
That gap produces unfair judgments in both directions. Some coaches get renewed because the teams like them. Others get cut because delivery numbers did not move, even though the causes sat with funding decisions, dependency queues, or a product owner who was never given decision rights. Neither call is grounded in anything you could defend in a steering meeting. The fix belongs at the start of the contract rather than the end.
The Accountability Question, Answered Directly
Agile coach accountability should connect a delivery symptom to an operating mechanism, an observable signal, and a capability the organization can retain without the coach.
The delivery condition is the symptom you actually care about. Not "improve agility." Something like: releases slip by two to three weeks with no warning until the final sprint, or production defects have doubled since the last platform migration.
The operating mechanism is the specific way of working the coach will change to address that condition. A refinement practice that produces sized, testable work. A dependency forum that resolves cross-team blockers within a fixed window. A forecasting method the product owner can run alone.
The observable signal is what management watches to see whether the mechanism is working. It is a leading indicator, visible on a normal cadence, not a quarterly self-report.
The capability transfer is what stays when the coach leaves: a named person who runs the mechanism, a written decision rule, a facilitation skill practiced without supervision.
Hold the coach accountable for the mechanism, the signal, and the transfer. Hold the organization accountable for the delivery condition, because the coach cannot control funding, staffing, architecture debt, or the executive who reprioritizes mid-quarter.
Ceremonies Are Instruments, Not Evidence
There is a fashionable move in coaching reviews where someone points at the meeting load and calls it theatre. Resist it, at least as a starting assumption. Kadenic et al. (2023) surveyed 182 Scrum practitioners and found that team maturity, composition, values, roles, and events were all associated with perceived Scrum success. Events were part of the picture, not noise around it. That study measures perception rather than audited delivery, and association is not causation, but it is enough to stop you from cutting ceremonies as a reflex.
A ceremony is an instrument. It can produce something or produce nothing, and attendance tells you which.
A Sprint Review produces stakeholder feedback when a stakeholder changes their mind in the room and that change reaches the backlog. If nobody outside the team spoke, the Review was a status meeting with a different name. A Retrospective produces a tested improvement when the team runs one experiment, checks it, and keeps or drops it. If the same three items appear in the notes every fortnight, the Retrospective is generating a list rather than a change.
So the review question is not whether the teams are doing the ceremonies. It is what the last four Reviews changed, and which improvement was tested after the last four Retrospectives.
The broadest available evidence points the same way. A seven-year research programme spanning thirteen field studies, 4,940 professionals, and 1,978 teams built a model linking team effectiveness to responsiveness and stakeholder concern, enabled by continuous improvement, team autonomy, and management support (Verwijs & Russo, 2023). That is a large body of work, and it highlights continuous improvement and management support as enabling conditions. It is also cross-sectional survey and modelling evidence. It does not isolate what a coach contributes, and it does not prove that improving these conditions causes delivery outcomes to move. Use it to decide what deserves attention, not as proof that a specific intervention will pay off.
Effective Coaching Includes Directing
A common contract assumption holds that a good coach facilitates and gradually withdraws. Sometimes that is right. As a universal rule the evidence points elsewhere.
An exploratory survey of 301 respondents found a mismatch worth naming: team members often want more directive help and tangible outcomes, while coaches tend to emphasize facilitation and systemic impact (Koumaditis et al., 2026). This is new and exploratory work, so treat the size of the gap cautiously. The practical implication is low risk. Write down which mode you are buying. Facilitation of a specific forum. Direct support such as drafting the first release checklist or pairing on estimation. System work such as renegotiating a handoff with another department. Deliberate transfer of a leadership responsibility.
That last mode has the clearest research behind it. A grounded theory study with 75 practitioners across 11 divisions traced nine leadership roles moving from the Scrum Master to the team as maturity developed, shaped by the leadership gap, supportive climate, trust, freedom, and role conflict (Spiegler et al., 2021). Transfer was not automatic. It depended on conditions that management partly controls. Interviews with 13 professionals at ten companies similarly found leadership to be dynamically shared, tied to a sense of belonging, and shaped by competing organizational cultures (Gren & Ralph, 2022). Both studies are qualitative and neither measured delivery performance, so read them as descriptions of how transfer happens rather than as a promise of what it produces.
If your organization has a leadership gap the team cannot fill yet, a coach who steps back on schedule will leave a hole. Sequence matters more than ideology.
Signals to Watch, and What They Are Worth
Pick two or three observable signals per delivery condition. Candidates management can see without a special report:
- Review-wait time: how long finished work sits before someone with authority looks at it
- Blocked-item age: the oldest currently blocked item, in days
- Forecast error: committed scope versus delivered scope, tracked over several sprints
- Release frequency
- Escaped defects reaching production
- Decision latency: time from a question being raised to a decision being made
- Completed improvement experiments: how many were run and concluded, not proposed
Be clear about their status. These are proposed measures for your local context. None of the studies above tested them as coaching outcomes, and none establish that a coach moves any of these numbers. They are instruments for a conversation, chosen because they are hard to fake and visible on a normal cadence.
Running the Review Conversation
Sit down with the coach and work through four questions in order.
Which delivery condition were we working on, and can we both state it the same way? Which mechanism did we change, and can you show me it running without you in the room? What signal did we agree to watch, and what did it do over the last quarter? Who now owns the mechanism, what have they run alone, and what would break if you stopped next month?
Where the signal moved, ask what else changed at the same time. Where it did not, separate the two possible causes: the mechanism was never really installed, or it was installed and the constraint sits somewhere the coach cannot reach. The second answer is often the most useful thing the engagement produces, because it tells you the next problem is yours.
A coach can prepare for that review honestly, and a sponsor can defend it. It also makes renewal easier to decide, because you stop arguing about whether the coaching felt valuable and start looking at four specific things, deciding which of them still needs work.
References
Gren, L., & Ralph, P. (2022, May 21). What makes effective leadership in agile software development teams? In Proceedings of the 44th International Conference on Software Engineering (pp. 2402–2414). ACM. https://doi.org/10.1145/3510003.3510100
Kadenic, M. D., Koumaditis, K., & Junker-Jensen, L. (2023). Mastering scrum with a focus on team maturity and key components of scrum. Information and Software Technology, 153, 107079. https://doi.org/10.1016/j.infsof.2022.107079
Koumaditis, K., Harder, P., Hestbæk, B. S., Kadenic, M. D., & Lui, L. F. (2026, April 12). The agile coach role decoded: What agile coaches deliver versus what team members experience. Journal of Software: Evolution and Process, 38(4), e70107. https://doi.org/10.1002/smr.70107
Spiegler, S. V., Heinecke, C., & Wagner, S. (2021, March 22). An empirical study on changing leadership in agile teams. Empirical Software Engineering, 26(3), 41. https://doi.org/10.1007/s10664-021-09949-5
Verwijs, C., & Russo, D. (2023, April 27). A theory of Scrum team effectiveness. ACM Transactions on Software Engineering and Methodology, 32(3), 1–51. https://doi.org/10.1145/3571849