# Release 4.2: investigation log

> Working notes on what 4.2 did and why. Kept out of `CLAUDE.md` so the
> context file stays a context file. Rook is a fictional teaching scenario.

## Round 1, 8 Sept: the two-population split

From the routing code, callout history and tickets read together. Regroup
brief with next steps: [4.2-regroup-brief.md](4.2-regroup-brief.md).

- The acceptance-rate dip is not a gradual seasonal drift. Weekly data
  shows a sharp step down (78% to 54%) in the week 4.2 shipped, recovering
  to ~73% by month-end. Be skeptical of "it's seasonal, wait for September."
- "Phone never goes off" is two different problems. Farlight,
  Meteor Mite, The Undertow, and Vesper are starved; their offer
  volume collapsed toward zero, a real routing/ranking effect. Nightwell,
  Stormwrack, and Sgt. Falkirk are at record-high offer volume yet report
  the identical complaint. **The push-delivery explanation for this second
  group was refuted in Round 2.**
- The recovering topline acceptance number may be misleading: as the
  starved group's volume shrinks toward zero, they drag the aggregate down
  less, so the number climbing back doesn't prove the split is healing.
- The routing reweight's own effect is real but bounded (isolated as
  roughly ±0.10 on the score), not big enough alone to cause a total
  collapse. The likely trigger is the timeout cut crashing a responder's
  acceptance-history score in the release week itself, which then
  compounds through the top-ranked-only dispatch mechanism.
- Marcus's Slack question about decline-vs-timeout scoring has an answer in
  `history.py`. It is the mechanism, not a documentation gap. See Round 2.
- Capability-tag specialists (Undertow/aquatic, Farlight/crowd-management)
  may be structurally disadvantaged now that proximity dominates the score.
  "Coverage gap" wouldn't catch this, since it only fires when nobody
  matches at all.

## Round 2, 10 Sept: the ratchet, and what the feedback piles each miss

From reading the routing source, then re-testing the interviews and tickets
against it. Prompts and per-round findings:
[02-super-hearing/prompts.md](02-super-hearing/prompts.md).

- **The ranking score is a ratchet.** `history.py` has no decay. Wen's
  TODO asking whether the score should ease back toward neutral is dated
  2019 and still open.
  `DECLINE_PENALTY` (0.12) is 1.5x `ACCEPTANCE_CREDIT` (0.08), so a
  responder must accept 60% of offers just to hold level. A timeout is
  scored identically to a refusal. `SCORE_FLOOR = 0.0` is absorbing: at
  the floor you rank last on every callout, are never offered anything,
  and can never earn the credit that would lift you. In release week the
  whole cohort hit 54.2%, under the 60% break-even, so every score fell
  at once. Most climbed back. Farlight, Meteor Mite, The Undertow and
  Vesper did not, and **cannot recover without someone resetting them.**
- **No push fix shipped in 4.2.** The CHANGELOG lists three items only:
  ranking weights, timeout 90s->60s, console filter persistence. "Mobile
  push reliability" was 4.1 (16 Jun); Priya's handoff confirms mobile has
  been stable since. Any push hypothesis is misattributed.
- **The delivery-bug theory is refuted by `pings_taken`.** Nightwell
  accepted 14 callouts in the week she filed "nothing in like 10 days";
  Stormwrack's take count hit a ten-week high the week his handler called
  the quietest in two years. A silent phone can't be answered 14 times.
  What's actually wrong for that group is still unknown.
- **`glossary.docx` is wrong on two points that matter**, and everyone
  including Priya has been reasoning from it: it says the recent-acceptance
  component drops "until the component recovers" (it never recovers), and
  that a decline is "distinct from a timeout in the data" (the code treats
  them identically). This is why "wait for September" sounded reasonable.
- **Don't trust `callout-history.csv` until Ravi confirms it.** It has no
  provenance note and may be the rough pull Marcus offered on 19 Aug
  rather than real weekly reporting. The code corroborates its four
  collapsing responders; nothing corroborates its record-high numbers.
- **The ticket pile is curated, not a population.** T-001..T-025, no gaps,
  ~1/day, 100% on-topic, starting the day *after* the release. It is almost
  certainly Nadia's assembled breakdown. It has no pre-4.2 baseline, so
  "3x normal" is unverifiable and no "this is new" claim can be tested
  against it.
- **Watch distribution, not the acceptance rate.** Offer spread was flat
  for six weeks (max/min 2.1-2.7x) then went 6x, 20x, 21x and is still
  widening, while topline acceptance "recovers" from 54% to 73%. Total offer
  volume is roughly flat (-4.3% against the pre-release six-week average of
  172.3 offers/week; -6.8% only if measured from the release week itself,
  which is the wrong baseline), which by itself refutes the seasonal
  theory. Nobody computes the spread.
- **Each feedback channel caught only half the starved group.** Vesper and
  Meteor Mite appear only in the interviews (zero tickets; their handlers
  absorb it). Farlight and The Undertow appear only in the tickets. Read
  both, always. Starvation generates no event, so it only reaches a ticket
  if someone notices.
- **Halloran's Supply issue is unresolved and is a safety matter:** eleven
  days waiting on a quartermaster signature for a cracked vest plate, with
  the responder in the field on degraded armour. It appears in no ticket
  because the queue is callout-only.
- **Two steers in Priya's handoff to resist:** that this is "mostly
  seasonal," and that the filter-persistence tickets are "cosmetic and
  noise." Ambrose's version is a silent-reversion bug, not a cosmetic
  complaint.
