Claude Code for PMs · Final report · 24 September 2026

The four
who went
quiet

Release 4.2 pushed four reliable responders to the bottom of a ranking they cannot climb out of. Nothing in Dispatch noticed, recorded it, or told anyone. This is how I proved it, and what I would build.

The order · release weekscore Δ
Sgt. Falkirk+0.08
Captain Vantage+0.04
The Longcast+0.04
Cindermark0.00
Corporal Ashgrove0.00
The Gale0.00
The Drift−0.04
Ironvale−0.08
Stormwrack−0.08
Halfmoon−0.12
Nightwell−0.12
Sgt. Bulwark−0.20
Vesper−0.24
Farlight−0.40
Meteor Mite−0.40
The Undertow−0.52

Net score change in the week of 10 August: 0.08 for each offer taken, minus 0.12 for each one missed. callout-history.csv, rows 98–113.

Start here

Put four people back this week, then build the record

Turning the timer back to 90 seconds helps a stranded responder catch the one offer he still gets each week. It does not get him more offers. The ranking decides who is asked, and his standing is on the floor.

1
This week

Put the four back where everyone else is

Reset Farlight, Meteor Mite, The Undertow and Vesper to the top of the range, recorded with a name and a reason. A data change, not a release. The midpoint sounds fair and is not: everyone else is at the top, so it leaves them behind everyone for seven more accepts in a row.

2
First build

Record every offer, and turn the timer back up

Who it went to, where they were in the order, whether the phone showed it, what came back. Rook keeps none of this today. It depends on nothing else, and it is how we find out from data whether scores survive a deploy. The timer goes back to 90 seconds in the same release, said openly, so the next release week does not repeat this one.

3
Then

Tell the handler, and tell the responder

Alert the handler the day a responder drops below the threshold. Show the responder that he is still active and where he stands. The Undertow asked that on 26 August, and nobody could answer.

4
Next

Answer Wen's 2019 question

Yes, the standing should ease back over time. Nobody should carry one bad week forever. That fixes it for the next person, once we know where the score is kept.

Notreverting the ranking weightsresetting anyone silentlya setting somebody flips

The weights stay because the rebalance halved what each miss costs; the shorter timer caused the misses. No silent reset, because a system that changed people's standing without telling anyone is the failure being fixed. No setting, because Helen asked for something a handler would notice.

What the dashboard showed

The dashboard said recovery.

60%80%29 Jun12 Aug31 Aug60%: below this, scores fall54%73%
54%
release week
→
73%
by 31 August

Weekly acceptance rate, all sixteen responders: 77.9% the week before release, 54.2% in release week, 72.7% by 31 August. Rook's glossary calls it the headline metric. callout-history.csv

Same weeks

What happened to people

Four people had disappeared.

0102029 Jun12 Aug31 Augthe four
11–14
offers a week, before
→
0 or 1
after 12 August

Offers per responder per week, the four in amber. Afterwards: Farlight 3, 1, 0. The Undertow 4, 1, 1. Vesper 5, 2, 1. Meteor Mite 4, 2, 1. Part of the recovery on the left is these four dropping out of the count: they went from 28% of all offers to under 2%.

4 of 16
responders stopped being asked
callout-history.csv, lines 114–161
79–89%
fewer offers each, in the three weeks after release
62% on the most cautious slicing
−4%
total offers across all sixteen
172.3 a week before, 165 on 31 Aug
0
alerts, records or messages about any of it
no log call in the routing code we have

Two of the six people hurt never wrote in

Vesper and Meteor Mite never filed a ticket; they surfaced because their handlers were interviewed about something else. Losing offers produces no event, so it only reaches a ticket if someone happens to notice.

Wrote in
Silent
Hurt
4Farlight, The Undertow. Plus Ashgrove and Halfmoon, sliding 20 to 27%, a softer case.
2Vesper, Meteor Mite.
Not hurt
8At record volume, still reporting silence. Still unexplained.
2Sgt. Bulwark, The Gale.

Tickets against the data file, from 03-rewind/prompts.md, Round 2.

How I found it

Each module checked the one before it

Each module added a different source, and a different way of not taking the last answer on trust. For each: the technique, one finding, and the prompt I wrote that produced it, as I typed it.

01
Module 1 · Onboard
Context from Rook's own documents

The ranking change alone was too small to explain it

4.2 bundled a ranking change and a shorter timer, and my predecessor had said the dip was mostly seasonal. The prompt below separated the two changes. Modelled in Module 1, the ranking change on its own was not big enough to cause a collapse, and the likelier trigger was the timer cutting into responders' records in release week. Module 4 later confirmed that from the code.

TechniquePoint Claude at the folder and write a CLAUDE.md, then ask what a fresh chat could not.
Too small alone
the ranking change, modelled
My prompt · 01-origin-story/prompts.md #1In 00-rook/code/dispatch-routing/, simulate routing.score() for Farlight, Meteor Mite, The Undertow, and Vesper using the pre-4.2 weights (0.45/0.40/0.15) versus the 4.2 weights (0.60/0.25/0.15), holding proximity and capability match constant. How much of their collapse in callout-history.csv is explained by the weight rebalance alone, separate from the shorter timeout?

Modelled, not measured. 4.2-investigation.md, Round 1.

02
Module 2 · Listen
Interviews against tickets

The people interviewed cannot see the problem

All four interviews were with handlers. When I asked what none of them mentioned, the answer was the responder's phone. None of them uses it, and that is where every offer arrives and disappears. Nobody complained about slow support either, in a month when tickets ran high, because two of them said filing something is where it stops.

TechniqueGroup, count and quote, instead of summarising. Then ask what nobody said.
0 of 4
handlers who use the responder's phone
My prompt · 02-super-hearing/prompts.md #6What would people in this situation be expected to complain about that none of the four mentioned at all? Not the things one person raised and the others didn't, but the things nobody raised. For each, say whether their silence means it isn't happening, or means it wouldn't be visible to someone in their role.

02-super-hearing/prompts.md, Round 2 findings.

03
Module 3 · Verify
The number, the rows, and a second method

The same four names, however you measure it

I recomputed the finding twenty-seven ways: three baselines, with and without release week, three metrics, three thresholds. The same four responders came out every time. Nobody joined the list and nobody left it.

The number I put in front of Helen in Module 3 was the cautious one: four of sixteen responders receive 62–100% fewer offers than before 12 August, counting release week as after, while total offers are down 4%. On the three weeks after release alone the drop is 79–89%, the figure in the tiles above.

TechniqueAsk for the rows behind a number, then work it out another way.
27 ways
same four names every time
My prompt · 03-rewind/prompts.md #2Recompute the starvation finding under every reasonable alternative definition and show me how much it moves. Baseline: 6 weeks pre-release vs 3 weeks vs the single last week before. After-window: including vs excluding release week. Metric: pings sent, pings taken, or share of total weekly volume. Threshold: what counts as "starved" — under 2/week, under 25% of prior, or zero? Then tell me the version of the number that's most conservative and still true, and whether any choice flips the four-responder list — does Ashgrove or Halfmoon join it, does anyone drop out?
04
Module 4 · Inspect
What the code does, and doesn't

A ratchet with no way back up

A missed offer costs 0.12 and an accepted one earns 0.08, so a responder must accept 60% just to hold level. A timeout counts as a decline. Nothing ever eases the score back, and the floor is zero. In release week the whole fleet accepted 54%, under break-even, because the timer had just dropped from 90 seconds to 60.

I wrote four predictions before opening the file. All four held. Ranking everyone by what they lost that week puts exactly the four at the bottom, which is the order on the cover.

One thing adds points, one takes them off, and nothing saves them history.py: +0.08 per accept, −0.12 per miss, floor 0.0, ceiling 1.0. No decay, no reset, no persistence, no log. Offers go to the highest score first, so the score decides who is asked at all. 0.0 floor0.5 NEUTRAL_SCORE1.0 ceiling asked last, or neverwhere a new responder startswhere everyone sat before 4.2 9 misses from the ceiling to the floor 5 misses from neutral to the floor 7 accepts back to neutral, still below everyone 13 accepts back to the ceiling, where the pack is, and you cannot accept what you are not offered Share of 10,000 release-week offer orderings in which the bottom four are exactly Farlight, Meteor Mite, The Undertow and Vesper 26% if every score started the week at 1.0 six good weeks, scores kept between deploys 95% if every score started the week at 0.5 _scores = {} is never saved; a deploy resets it If the 12 Aug deploy reset every score to 0.5, the whole fleet started release week five misses from the floor, and the ratchet alone explains all four. Whether production persists the score is a yes/no for Marcus, and the most important question of the week. Source: 00-rook/code/dispatch-routing/history.py lines 17, 22, 30–43; config.py; simulation over callout-history.csv rows 98–113. Module 4, Round 3.
Rank all sixteen by release-week score loss and the bottom four are the four who went quiet Net change to the recent-acceptance score in the week of 10 Aug: 0.08 × taken − 0.12 × missed, from callout-history.csv rows 98–113. Prediction, written before opening the file: if the 60s window is what starved them, the four should show the largest score loss of anyone. Result: exactly those four, no overlap. The Undertow -0.52 Farlight -0.40 Meteor Mite -0.40 Vesper -0.24 Sgt. Bulwark -0.20 Halfmoon -0.12 Nightwell -0.12 Ironvale -0.08 Stormwrack -0.08 The Drift -0.04 Cindermark 0.00 Corporal Ashgrove 0.00 The Gale 0.00 Captain Vantage +0.04 The Longcast +0.04 Sgt. Falkirk +0.08 -0.5 -0.4 -0.3 -0.2 -0.1 0 +0.1 net score change in release week 0.04 from Vesper, 0.75 min of travel; 10 offers next week to her 5 stopped being offered work after 12 Aug everyone else Miss count alone does not pick them out: Nightwell missed seven, the same as The Undertow, and recovered. Misses net of accepts does. Nobody was below the 60% break-even before 12 Aug, so nobody carried a penalty into release week. Source: 00-rook/data/callout-history.csv, week starting 2026-08-10; constants from config.py. Module 4, Round 2, prompt 4.

Marcus's question from 14 August, whether the change treated people already turning jobs down differently: no. Same rules for everyone. The four were not turning work down beforehand; Vesper had the second-best record in the fleet.

TechniqueWrite down what the data must show if you are right, then check.
4 of 4
predictions held
My prompt · 04-x-ray-vision/prompts.md #4My hypothesis is that the four starved responders fell to the floor of the recent-acceptance score because the 60s window made them miss offers in release week, and the absorbing floor kept them there. Before looking, write down what `callout-history.csv` must show if that's true: (1) the four were *not* below the 60% break-even before 12 Aug, so they carried no pre-existing penalty; (2) in release week they lost more score than anyone else, computed as 0.08 × taken − 0.12 × (sent − taken); (3) ranking all 16 by that release-week loss puts exactly those four at the bottom; (4) their offers fall *after* the score falls, not before. Then check each, and say where the numbers stop supporting it. Along the way: were any of the 16 "already turning jobs down" before 4.2, and did the code treat them differently?
05
Module 5 · Prototype
The brief, then the build

Every fact in his month was in the system on 16 August

I built the brief and the prototype for one person: The Undertow, who wrote to us from his phone on 26 August, "starting to wonder if im still even in the system." The prompt below produced the screen that puts his August side by side: what happened, and what would have happened with this built. It is the last screen of the prototype in the fix.

TechniqueWrite the brief for one real person, push on it, then build something to click.
19 days → same day
two tickets, against one alert and one decision
My prompt · 05-super-speed/prompts.md #7Put the two Augusts side by side for The Undertow, same dates, same rows. On the left, what actually happened: 12 August he is second in the order, 16 August he is fourteenth, Okafor files T-005 as Low on the 19th, The Undertow writes to us himself on the 26th, T-019 goes in as High on the 31st, and nothing changes at any point. On the right, the same month with this built: the alert on the 16th, Okafor acting, his phone answering him on the 26th. One person, one month, twice. That is the screen I want Helen looking at.
06
Module 6 · Automate
A check that runs without me

It caught the most dangerous line in my own brief

review-checklist checks any brief for six things. Four came from the course. Two are mine, because four modules of corrections came from numbers that cited nothing and a guess written as a finding. On my own Module 5 brief it flagged "decay is slow, roughly 0.02 a day". That rate was an assumption from a simulation, written as a fact.

TechniqueWrite a habit down once as a skill, run it twice, schedule it.
5 flags
in my own brief, first run
My prompt · 06-sidekicks/prompts.md #1I want to save a habit as something I can reuse. When I review a brief before it goes any further, here's what I actually check for:
- It names who owns it
- It says how we'll know it worked
- The scope at the end matches the scope at the start
- It explains the problem before it proposes a fix
- Every number used as evidence says where it came from, so someone could reopen the source
- Anything not yet known is written as a question, not as a fact

The last two are mine. Four modules of this course were spent catching numbers that sounded right and cited nothing, and a guess written as a finding that ended up in a committed file. Turn that into a skill called review-checklist, so I can point it at any brief and get the same check every time, without me explaining it again. Report one line per criterion and quote the words behind every flag.

The fix

Tell the handler the day it happens, and tell the responder why

Click first on the switch at the top of the console. Today shows the morning Okafor actually had: three cards, nothing wrong, nothing to open. Proposed shows the alert. Open it, then put The Undertow back and watch his phone change.

Live prototype, five screens. A mock, not a build.Open full screen

The brief behind it, 05-super-speed/brief.md, also says what a lift costs. When Meteor Mite goes back up, The Gale, who took most of her callouts, is asked less. Every reranking has somebody on the other side of it.

Keeping it from happening again

The number everyone watched is the kind my skill now rejects

What Rook watched
54% → 73%

Weekly acceptance, which Rook's glossary calls the headline metric. It fell in release week and was back most of the way by 31 August. It looked like recovery.

What was happening
28% → under 2%

The four's share of all offers over the same weeks. Part of the recovery was them no longer being counted.

My skill's second criterion flags a success measure that can move for another reason. This one moved back up because four people stopped being asked. Had 4.2 been judged by a measure like that, the skill would have flagged it before anyone relied on it.

The skill keeps the next plan honest. The offer record and the alert are what keep the next four from going unseen.

review-checklist · seven runs

1
My Module 5 brief
5 flags
No owner, no success measure, decay never scoped, unsourced numbers, a simulated rate stated as fact.
1b
The same brief, fixed
0 flags
1c
The brief after later edits
4 flags
Two broken by edits made after it passed, and a modelled number the first two runs missed. Fixed: 0.
1d
Cut to one page, stricter criteria
0 flags
2
Four briefs with an answer key
matched, with a caveat
I ran it after reading the key, so the match proves little.
3
A brief written to fail
exact match
Expectation written first. It caught all three planted problems, including one the skill's first wording would have passed.
Peer
Two classmates' briefs
3 flags each
Both missing a named owner.

Every run so far applied the skill's written instructions by hand from another session. It loads on its own in a session opened in this repo, which is where a scheduled run would happen. The weekly schedule is documented in 06-sidekicks/schedule.md rather than left running for a fictional scenario.

Credibility

The numbers held. The sourcing is where I slipped.

Every number I recomputed held up. What failed was the sourcing and the reasoning around the numbers. I left each correction visible in the repo rather than cleaning it away.

Module 1

A prompt cited a "duplicate push notification on re-offer" fix in the 4.2 notes, and built a whole theory on it.

Caught

The changelog lists three items and none of them is that. The theory was then refuted on data: Nightwell accepted 14 callouts in the week she reported hearing nothing. Corrected in place, dated.

Module 2

The acceptance "recovery" was called mostly an artifact of the four dropping out of the count.

Caught

Recomputed without them: it explains about 2.6 of an 18.5 point climb. Real, but a minority. The other twelve genuinely recovered.

Module 2

Total offers were written as down 6%.

Caught

Only if you measure from release week itself. Against the six weeks before, it is 4.3%.

Module 5

The brief said decay is slow, "roughly 0.02 a day", as if it were a fact.

Caught

Decay does not exist. 0.02 was a rate I assumed. My own review-checklist flagged it, which is what it was built for.

How sure I am

Two numbers from Ravi and one answer from Wen decide the rest

The four responders and the mechanism are confirmed by the data, the code and their handlers independently. These four questions decide how much further the rest goes.

Would change the plan, due before the regroup
  • If Ravi's count of callouts created shows about a hundred went unanswered, this stops being a fairness fix and goes to Helen as a safety issue the same day.
  • If Wen says the score survives a deploy, the alert can ship before the offer record, because the number it watches is stable.
  • If the tagged-incident pull shows specialists losing to people without the tag, capability becomes a gate on the list and that work goes ahead of the phone screen.
Wen

Does the score survive a deploy?

The code keeps it in memory and never saves it. If production does the same, every release resets everyone to the midpoint. Modelled that way, the ranking alone picks out exactly these four in 95% of cases, against 26% otherwise.

Changes the build order
Ravi

What is this data file?

Eight busy responders report silence in weeks the file shows them at record volume. No reading makes both true. Nothing independent confirms the file's high numbers.

Decides what the rest rests on
Ravi

Did a hundred callouts go unanswered?

Accepted callouts fell from 132 a week to 96, 104, 108 and 120 while offers held steady. If an accepted offer is a filled callout, about a hundred incidents found nobody. Nobody at Rook has mentioned one.

A safety question, not a fairness one
Wen and Marcus

Should capability gate the list?

Someone with none of the needed skills now outranks a fully qualified responder 46 minutes away if they are within 24 minutes. Before 4.2 it was 10. The Undertow and Farlight are both specialists.

The mechanism is certain; the impact is not