Claude Code for PMs · Final report · 24 September 2026
The four who went quiet
Release 4.2 pushed four reliable responders to the bottom of a ranking they cannot climb out of. Nothing in Dispatch noticed, recorded it, or told anyone. This is how I proved it, and what I would build.
Antje Barth · PM, Rook Dispatch For Helen Achebe, Director of Product Rook Industries is a fictional teaching scenario
The order · release weekscore Δ
Sgt. Falkirk+0.08
Captain Vantage+0.04
The Longcast+0.04
Cindermark0.00
Corporal Ashgrove0.00
The Gale0.00
The Drift−0.04
Ironvale−0.08
Stormwrack−0.08
Halfmoon−0.12
Nightwell−0.12
Sgt. Bulwark−0.20
Vesper−0.24
Farlight−0.40
Meteor Mite−0.40
The Undertow−0.52
Net score change in the week of 10 August: 0.08 for each offer taken, minus 0.12 for each one missed. callout-history.csv, rows 98–113.
Start here
Put four people back this week, then build the record
Turning the timer back to 90 seconds helps a stranded responder catch the one offer he still gets each week. It does not get him more offers. The ranking decides who is asked, and his standing is on the floor.
1
This week
Put the four back where everyone else is
Reset Farlight, Meteor Mite, The Undertow and Vesper to the top of the range, recorded with a name and a reason. A data change, not a release. The midpoint sounds fair and is not: everyone else is at the top, so it leaves them behind everyone for seven more accepts in a row.
2
First build
Record every offer, and turn the timer back up
Who it went to, where they were in the order, whether the phone showed it, what came back. Rook keeps none of this today. It depends on nothing else, and it is how we find out from data whether scores survive a deploy. The timer goes back to 90 seconds in the same release, said openly, so the next release week does not repeat this one.
3
Then
Tell the handler, and tell the responder
Alert the handler the day a responder drops below the threshold. Show the responder that he is still active and where he stands. The Undertow asked that on 26 August, and nobody could answer.
4
Next
Answer Wen's 2019 question
Yes, the standing should ease back over time. Nobody should carry one bad week forever. That fixes it for the next person, once we know where the score is kept.
Notreverting the ranking weightsresetting anyone silentlya setting somebody flips
The weights stay because the rebalance halved what each miss costs; the shorter timer caused the misses. No silent reset, because a system that changed people's standing without telling anyone is the failure being fixed. No setting, because Helen asked for something a handler would notice.
What the dashboard showed
The dashboard said recovery.
54%
release week
→
73%
by 31 August
Weekly acceptance rate, all sixteen responders: 77.9% the week before release, 54.2% in release week, 72.7% by 31 August. Rook's glossary calls it the headline metric. callout-history.csv
Same weeks
What happened to people
Four people had disappeared.
11–14
offers a week, before
→
0 or 1
after 12 August
Offers per responder per week, the four in amber. Afterwards: Farlight 3, 1, 0. The Undertow 4, 1, 1. Vesper 5, 2, 1. Meteor Mite 4, 2, 1. Part of the recovery on the left is these four dropping out of the count: they went from 28% of all offers to under 2%.
4 of 16
responders stopped being asked
callout-history.csv, lines 114–161
79–89%
fewer offers each, in the three weeks after release
62% on the most cautious slicing
−4%
total offers across all sixteen
172.3 a week before, 165 on 31 Aug
0
alerts, records or messages about any of it
no log call in the routing code we have
Two of the six people hurt never wrote in
Vesper and Meteor Mite never filed a ticket; they surfaced because their handlers were interviewed about something else. Losing offers produces no event, so it only reaches a ticket if someone happens to notice.
Wrote in
Silent
Hurt
4Farlight, The Undertow. Plus Ashgrove and Halfmoon, sliding 20 to 27%, a softer case.
2Vesper, Meteor Mite.
Not hurt
8At record volume, still reporting silence. Still unexplained.
2Sgt. Bulwark, The Gale.
Tickets against the data file, from 03-rewind/prompts.md, Round 2.
How I found it
Each module checked the one before it
Each module added a different source, and a different way of not taking the last answer on trust. For each: the technique, one finding, and the prompt I wrote that produced it, as I typed it.
01
Module 1 · Onboard
Context from Rook's own documents
The ranking change alone was too small to explain it
4.2 bundled a ranking change and a shorter timer, and my predecessor had said the dip was mostly seasonal. The prompt below separated the two changes. Modelled in Module 1, the ranking change on its own was not big enough to cause a collapse, and the likelier trigger was the timer cutting into responders' records in release week. Module 4 later confirmed that from the code.
TechniquePoint Claude at the folder and write a CLAUDE.md, then ask what a fresh chat could not.
Too small alone
the ranking change, modelled
My prompt · 01-origin-story/prompts.md #1In 00-rook/code/dispatch-routing/, simulate routing.score() for Farlight, Meteor Mite, The Undertow, and Vesper using the pre-4.2 weights (0.45/0.40/0.15) versus the 4.2 weights (0.60/0.25/0.15), holding proximity and capability match constant. How much of their collapse in callout-history.csv is explained by the weight rebalance alone, separate from the shorter timeout?
Modelled, not measured. 4.2-investigation.md, Round 1.
02
Module 2 · Listen
Interviews against tickets
The people interviewed cannot see the problem
All four interviews were with handlers. When I asked what none of them mentioned, the answer was the responder's phone. None of them uses it, and that is where every offer arrives and disappears. Nobody complained about slow support either, in a month when tickets ran high, because two of them said filing something is where it stops.
TechniqueGroup, count and quote, instead of summarising. Then ask what nobody said.
0 of 4
handlers who use the responder's phone
My prompt · 02-super-hearing/prompts.md #6What would people in this situation be expected to complain about that none of the four mentioned at all? Not the things one person raised and the others didn't, but the things nobody raised. For each, say whether their silence means it isn't happening, or means it wouldn't be visible to someone in their role.
02-super-hearing/prompts.md, Round 2 findings.
03
Module 3 · Verify
The number, the rows, and a second method
The same four names, however you measure it
I recomputed the finding twenty-seven ways: three baselines, with and without release week, three metrics, three thresholds. The same four responders came out every time. Nobody joined the list and nobody left it.
The number I put in front of Helen in Module 3 was the cautious one: four of sixteen responders receive 62–100% fewer offers than before 12 August, counting release week as after, while total offers are down 4%. On the three weeks after release alone the drop is 79–89%, the figure in the tiles above.
TechniqueAsk for the rows behind a number, then work it out another way.
27 ways
same four names every time
My prompt · 03-rewind/prompts.md #2Recompute the starvation finding under every reasonable alternative definition and show me how much it moves. Baseline: 6 weeks pre-release vs 3 weeks vs the single last week before. After-window: including vs excluding release week. Metric: pings sent, pings taken, or share of total weekly volume. Threshold: what counts as "starved" — under 2/week, under 25% of prior, or zero? Then tell me the version of the number that's most conservative and still true, and whether any choice flips the four-responder list — does Ashgrove or Halfmoon join it, does anyone drop out?
04
Module 4 · Inspect
What the code does, and doesn't
A ratchet with no way back up
A missed offer costs 0.12 and an accepted one earns 0.08, so a responder must accept 60% just to hold level. A timeout counts as a decline. Nothing ever eases the score back, and the floor is zero. In release week the whole fleet accepted 54%, under break-even, because the timer had just dropped from 90 seconds to 60.
I wrote four predictions before opening the file. All four held. Ranking everyone by what they lost that week puts exactly the four at the bottom, which is the order on the cover.
Marcus's question from 14 August, whether the change treated people already turning jobs down differently: no. Same rules for everyone. The four were not turning work down beforehand; Vesper had the second-best record in the fleet.
TechniqueWrite down what the data must show if you are right, then check.
4 of 4
predictions held
My prompt · 04-x-ray-vision/prompts.md #4My hypothesis is that the four starved responders fell to the floor of the recent-acceptance score because the 60s window made them miss offers in release week, and the absorbing floor kept them there. Before looking, write down what `callout-history.csv` must show if that's true: (1) the four were *not* below the 60% break-even before 12 Aug, so they carried no pre-existing penalty; (2) in release week they lost more score than anyone else, computed as 0.08 × taken − 0.12 × (sent − taken); (3) ranking all 16 by that release-week loss puts exactly those four at the bottom; (4) their offers fall *after* the score falls, not before. Then check each, and say where the numbers stop supporting it. Along the way: were any of the 16 "already turning jobs down" before 4.2, and did the code treat them differently?
05
Module 5 · Prototype
The brief, then the build
Every fact in his month was in the system on 16 August
I built the brief and the prototype for one person: The Undertow, who wrote to us from his phone on 26 August, "starting to wonder if im still even in the system." The prompt below produced the screen that puts his August side by side: what happened, and what would have happened with this built. It is the last screen of the prototype in the fix.
TechniqueWrite the brief for one real person, push on it, then build something to click.
19 days → same day
two tickets, against one alert and one decision
My prompt · 05-super-speed/prompts.md #7Put the two Augusts side by side for The Undertow, same dates, same rows. On the left, what actually happened: 12 August he is second in the order, 16 August he is fourteenth, Okafor files T-005 as Low on the 19th, The Undertow writes to us himself on the 26th, T-019 goes in as High on the 31st, and nothing changes at any point. On the right, the same month with this built: the alert on the 16th, Okafor acting, his phone answering him on the 26th. One person, one month, twice. That is the screen I want Helen looking at.
06
Module 6 · Automate
A check that runs without me
It caught the most dangerous line in my own brief
review-checklist checks any brief for six things. Four came from the course. Two are mine, because four modules of corrections came from numbers that cited nothing and a guess written as a finding. On my own Module 5 brief it flagged "decay is slow, roughly 0.02 a day". That rate was an assumption from a simulation, written as a fact.
TechniqueWrite a habit down once as a skill, run it twice, schedule it.
5 flags
in my own brief, first run
My prompt · 06-sidekicks/prompts.md #1I want to save a habit as something I can reuse. When I review a brief before it goes any further, here's what I actually check for: - It names who owns it - It says how we'll know it worked - The scope at the end matches the scope at the start - It explains the problem before it proposes a fix - Every number used as evidence says where it came from, so someone could reopen the source - Anything not yet known is written as a question, not as a fact
The last two are mine. Four modules of this course were spent catching numbers that sounded right and cited nothing, and a guess written as a finding that ended up in a committed file. Turn that into a skill called review-checklist, so I can point it at any brief and get the same check every time, without me explaining it again. Report one line per criterion and quote the words behind every flag.
The fix
Tell the handler the day it happens, and tell the responder why
Click first on the switch at the top of the console. Today shows the morning Okafor actually had: three cards, nothing wrong, nothing to open. Proposed shows the alert. Open it, then put The Undertow back and watch his phone change.
Live prototype, five screens. A mock, not a build.Open full screen
The brief behind it, 05-super-speed/brief.md, also says what a lift costs. When Meteor Mite goes back up, The Gale, who took most of her callouts, is asked less. Every reranking has somebody on the other side of it.
Keeping it from happening again
The number everyone watched is the kind my skill now rejects
What Rook watched
54% → 73%
Weekly acceptance, which Rook's glossary calls the headline metric. It fell in release week and was back most of the way by 31 August. It looked like recovery.
What was happening
28% → under 2%
The four's share of all offers over the same weeks. Part of the recovery was them no longer being counted.
My skill's second criterion flags a success measure that can move for another reason. This one moved back up because four people stopped being asked. Had 4.2 been judged by a measure like that, the skill would have flagged it before anyone relied on it.
The skill keeps the next plan honest. The offer record and the alert are what keep the next four from going unseen.
review-checklist · seven runs
1
My Module 5 brief
5 flags
No owner, no success measure, decay never scoped, unsourced numbers, a simulated rate stated as fact.
1b
The same brief, fixed
0 flags
1c
The brief after later edits
4 flags
Two broken by edits made after it passed, and a modelled number the first two runs missed. Fixed: 0.
1d
Cut to one page, stricter criteria
0 flags
2
Four briefs with an answer key
matched, with a caveat
I ran it after reading the key, so the match proves little.
3
A brief written to fail
exact match
Expectation written first. It caught all three planted problems, including one the skill's first wording would have passed.
Peer
Two classmates' briefs
3 flags each
Both missing a named owner.
Every run so far applied the skill's written instructions by hand from another session. It loads on its own in a session opened in this repo, which is where a scheduled run would happen. The weekly schedule is documented in 06-sidekicks/schedule.md rather than left running for a fictional scenario.
Credibility
The numbers held. The sourcing is where I slipped.
Every number I recomputed held up. What failed was the sourcing and the reasoning around the numbers. I left each correction visible in the repo rather than cleaning it away.
Module 1
A prompt cited a "duplicate push notification on re-offer" fix in the 4.2 notes, and built a whole theory on it.
Caught
The changelog lists three items and none of them is that. The theory was then refuted on data: Nightwell accepted 14 callouts in the week she reported hearing nothing. Corrected in place, dated.
Module 2
The acceptance "recovery" was called mostly an artifact of the four dropping out of the count.
Caught
Recomputed without them: it explains about 2.6 of an 18.5 point climb. Real, but a minority. The other twelve genuinely recovered.
Module 2
Total offers were written as down 6%.
Caught
Only if you measure from release week itself. Against the six weeks before, it is 4.3%.
Module 5
The brief said decay is slow, "roughly 0.02 a day", as if it were a fact.
Caught
Decay does not exist. 0.02 was a rate I assumed. My own review-checklist flagged it, which is what it was built for.
How sure I am
Two numbers from Ravi and one answer from Wen decide the rest
The four responders and the mechanism are confirmed by the data, the code and their handlers independently. These four questions decide how much further the rest goes.
Would change the plan, due before the regroup
If Ravi's count of callouts created shows about a hundred went unanswered, this stops being a fairness fix and goes to Helen as a safety issue the same day.
If Wen says the score survives a deploy, the alert can ship before the offer record, because the number it watches is stable.
If the tagged-incident pull shows specialists losing to people without the tag, capability becomes a gate on the list and that work goes ahead of the phone screen.
Wen
Does the score survive a deploy?
The code keeps it in memory and never saves it. If production does the same, every release resets everyone to the midpoint. Modelled that way, the ranking alone picks out exactly these four in 95% of cases, against 26% otherwise.
Changes the build order
Ravi
What is this data file?
Eight busy responders report silence in weeks the file shows them at record volume. No reading makes both true. Nothing independent confirms the file's high numbers.
Decides what the rest rests on
Ravi
Did a hundred callouts go unanswered?
Accepted callouts fell from 132 a week to 96, 104, 108 and 120 while offers held steady. If an accepted offer is a filled callout, about a hundred incidents found nobody. Nobody at Rook has mentioned one.
A safety question, not a fairness one
Wen and Marcus
Should capability gate the list?
Someone with none of the needed skills now outranks a fully qualified responder 46 minutes away if they are within 24 minutes. Before 4.2 it was 10. The Undertow and Farlight are both specialists.