Trust / Failure
The front page: the coach cites a precedent that never existed and a team kills a $25,000 test on it.
The catch: a fabricated citation always reaches a human and drops the product to decline-only. Bad model, silence, not confident error. Hallucinated citations under 1%, alert at 2%.
Not softened: the golden set is 10 rows, so the contract is not enforceable until 100. The coach has already been wrong once, by 9.2 points, and that miss stays in.
Scale / Governance
What breaks at 10x: not inference. Human onboarding, 65% of cost. Year one is capped at five hand-onboarded teams so year two can automate what the founder did by hand.
The posture: the coach argues and never acts. No write path into any customer system. Two decisions stay human: ambiguous read-outs and any prompt change, gated at 90% golden pass.
Named defect: the record does not yet feed the reasoning. Horizon 2, with its own kill rule.
Competitive
Honest: there is no competitive kill criterion, and it is better to say so than dress one up.
What kills us: not Statsig shipping review. The backtest failing, because a coach that is not better than the team loses to any native step closer to the data.
The gate: flagged experiments must succeed 20 points less often than unflagged, pooled across two or three named partners, about 36 flags, by week eight. One partner is a signal; the pool is the verdict. Under that, we stop and publish the miss.