The Kelbridge Service Centre is a fictional customer-contact operation with two teams of fourteen on the same floor. Over six months, Team B has lost three staff, its mean score on a validated exhaustion scale has reached 74 out of 100, and supervisor performance ratings have fallen. Team A, across the corridor, has not. A review is commissioned, and it releases what it finds in two rounds.
Nothing about Team B changes between the rounds. Only what you have been shown does.
There is no penalty for changing your mind, and no penalty for refusing to choose between the levels. Both are recorded.
Six interventions, three levels, and ten units to spend on twenty units' worth of them. The model that turns your combination into a twelve-month projection is printed in the teaching notes; the numbers are illustrative and are not forecasts of anything.
The number that matters is not the exhaustion score. It is the list underneath it.
| Combination | Cost | Levels | Exhaustion |
|---|
Five conclusions somebody might draw from the Kelbridge review. Say whether each is supported, partly supported, or not supported.
The correlation of 0.61 in Team B is real, orderly and would replicate. It is also entirely compatible with the whole team being at 74 because of a rota. A correlation between people measures how they vary within one setting - and a setting that is constant for everybody cannot show up in it at all. That is not a subtlety about statistics; it is the reason a workplace can produce a textbook individual-differences finding and a structural problem at the same time, and it is why the correlation survives every intervention in experiment 2.
Round 1 contains only individual-level evidence, and the individual-level explanation is the best-supported reading of it. That is not a trap. It is what happens whenever a review is scoped to one level: the evidence that would have pointed elsewhere is not absent from the world, it is absent from the file. Ask who chose the scope, and what it would have cost to widen it.
Round 2 hands you the team and the work together, and they are not the same level. Norms about asking for help live between people; a rota and a target are built into the job before anybody arrives. The distinction matters because it decides what you buy: a debrief in protected time treats the norm, and a rota redesign treats the conditions under which that norm formed. If the relational mechanisms are carrying the effect rather than causing it, treating them alone leaves the source in place - and the team gets a weekly meeting in which to discuss a rota nobody has changed.
Two teams, matched on staff characteristics at recruitment, on the same floor, doing nominally the same job, thirty-six points apart. That comparison is why the structural reading becomes compelling in round 2, and it is worth being precise about its limits: matched at recruitment is not matched now, three people have left Team B, and nobody randomly assigned anyone to a rota. It is strong evidence and it is not proof, and the difference matters if somebody is about to spend money on it.
Spend the whole budget on resilience and coaching and the split shifts, the queue and the calls-per-hour target are all still there, quietly reloading the exhaustion the training just discharged - which is why the model halves the effect of resilience training when nothing structural has changed. Spend the whole budget on the rota and the target and two people still have a documented training gap that no scheduling change will close. There is no combination in this laboratory that leaves nothing running, and finding out which residue you can live with is the actual decision.
An explanation that places the problem in the staff makes the staff the thing to be repaired: it is cheaper, faster, requires nobody else's permission, and if it fails the failure belongs to the people who were trained. An explanation that places it in the work design implicates a scheduling system, a target and a sign-off policy, all of which belong to somebody more senior. Both explanations can be partly true at once, so the choice between them is rarely settled by the evidence alone - and noticing whose interests each framing serves is part of reading a review, not a substitute for reading it.