↑↓ Space navigate · HomeEnd jump · Esc hide hint
01 / 08
AI Evals · Final ProjectCPO Review

Ascend IQ, Ship/Hold Decision for the Top-50 Enterprise Launch

An AI agent that answers plain-language competitive-intelligence questions for VP-level strategists at Fortune 500 accounts, grounded in Ascend Analytics' verified data.

Author · Cohort
Antje Barth · AI Evals Certification · Oct 2026
Slide 5 · StrategyM1

Accuracy before speed: one invented figure ends a $50k account

Target user + use case

VP-level strategists and product leaders at Fortune 500 companies who pay $50k+ a year for verified market intelligence · getting a specific, verified competitive insight from a plain-language question.

Definition of “good”

An answer they can put in front of their board without checking the source, in seconds instead of hours, so that Ascend keeps its $50k+ renewals through Q4.

Top 3 trust metrics
1

Hallucination rate, per claim

share of factual claims with no supporting span in the retrieved source

2

Latency

P95 time to a complete answer

3

Robustness

answers that stay correct and sourced on messy or out-of-scope questions

Two trade-offs
1 · Hallucination rate over latency

Hallucination rate over latency: retrieval and source checking cost seconds, and a wrong answer in front of a board ends a $50k relationship.

2 · Robustness over latency

Robustness over latency for the launch cohort: ask or refuse on an ambiguous question rather than answer fast.

→AI Evaluation Strategy Canvas: 01-evaluation-strategy/strategy-canvas.md
Slide 6 · RisksM2

6 of 20 beta answers stated a detail the source does not support

FAILURE 1P0

Fabricated Specifics

a detail the source does not support or contradicts (a seat minimum, a 'confirmed' speaker, a funding stage)

6of 20 beta answers
FAILURE 2P0

Dropped Qualifier

a true fact with the condition that makes it false removed ('native' SQL export, 'seamless' HubSpot)

3of 20 beta answers
FAILURE 3P1

Retrieval miss reported as 'cannot find'

(SOC2 badge was in the footer)

1of 20 beta answers
!
Business impact

This failure matters because Ascend IQ stated a specific the source does not support in 6 of 20 audited answers, and at that rate a VP asking ten questions has a 97% chance of carrying one into a board deck, which results in the first caught fabrication ending a $50k renewal and nearly every account in the 50-account launch cohort meeting one in its first week.

→Failure Taxonomy Canvas: 02-failure-discovery/failure-taxonomy.md
Slide 7 · ProofM3

Only the LLM judge catches it: 9 of 11, against 1 for code

eval-results.png
Ascend IQ three-layer eval suite results
Eval results screenshot unavailable. See the result summary.
9 / 11
caught by LLM judge
1 / 6
usage-drop trajectory
0.824
Cohen's kappa
Result summary

Three-layer suite run on the P0 pricing case and all 20 audit rows: the LLM judge caught 9 of the 11 confirmed failures (9 hallucinations, 1 retrieval miss, 1 P2 brand-voice miss outside the top three), the two deterministic layers caught 1 between them, so the P0 class is semantic. The usage-drop trajectory scored 1 of 6 dimensions (HOLD): the agent guessed 'seasonal' without ever checking ingestion. The judge reached Cohen's kappa 0.824 against human labels after two rubric revisions (from 0.286), measured on the 12 calibration traces it was tuned on; a held-out recalibration comes before the audit.

Slide 8 · StandardsM4

Zero unsupported claims in 300, or it does not ship

Eval spec · P0 target threshold
0
human-confirmed unsupported claims in a 300-claim held-out audit
per-claim rate < 1%95% confidence

Target risk: unsupported claims, covering both P0s (fabricated specifics and dropped qualifiers). Threshold: 0 human-confirmed unsupported claims in a 300-claim held-out audit, which bounds the per-claim rate under 1% at 95% confidence.

Gate action on failureHARD GATE

Hard gate: the release does not ship to any account until the audit passes; a failed audit re-runs on a fresh 300. In CI, faithfulness blocks the merge below 95 or on a regression of more than 3 points (PR #218 was blocked at -9).

Risk-tolerance rationale

Clients pay for verified data, so one invented number in a board deck ends the account; per the M1 canvas we put hallucination rate ahead of latency. The 20-row beta log cannot prove the bar (zero in 20 only bounds the rate under 15%), which is why the gate is a 300-claim audit. Latency is a Soft gate (P95 at or under 2.0s, override band to a 10s ceiling under a staged rollout) and brand voice is advisory.

Slide 9 · Decision · Pyramid memoM5 + M6

Hold until the audit passes; go/no-go on Nov 9

HOLD
until the 300-claim zero-fabrication audit passes; first wave of 10 accounts Nov 9 if it does
The Answer

I recommend we hold the Ascend IQ launch to our top 50 accounts until it passes the zero-fabrication audit, with the go/no-go on Nov 9, because in our beta 6 of 20 answers stated a detail the source does not support, and shipping now puts more than $2.5M of annual renewals in front of an answer a VP has a 97% chance of catching out within ten questions.

The 3 Arguments
01

Brand risk. The failure we would ship is the one clients pay us to prevent: invented specifics and dropped qualifiers break 'use it without checking the source', and none of the 20 beta answers cites a source.

02

Revenue risk. The exposure lands on our most valuable accounts, all at once, where we can't take it back: an invented number goes into a client's board deck, so we learn about it from them.

03

Reliability risk. We can't yet prove it's fixed, but we know exactly what would: the judge (kappa 0.824 on its tuning set), the CI gate (PR #218 blocked) and Level 3 coverage are in place; the 300-claim audit is the missing proof.

Business risk + next step

Holding costs a few weeks, not trust. Approve by Fri Oct 2; fixes by Oct 23; audit from Oct 26; go/no-go Nov 9.

Pass: 10 accounts go live behind flags, all 50 after 14 clean days. Fail twice: back to the CPO with a new date.

Riskiest assumption: that VPs open the citations. Wave 1 measures it.

Coverage matrix (M5)

Portfolio context, a sibling product (AI-Powered Report Summaries): hallucination covered only by a proxy judge, latency covered, bias, toxicity and drift uncovered. Toxicity accepted on low expected impact with a kill criterion; drift is the critical gap, closed by a pre-send grounding check and a new-template launch gate. The same pre-send pattern is a launch condition for Ascend IQ.

Budget allocation (M5)

Level 3: Data Fabrication ($85K) + Source Attribution ($65K) = $150K of the $200K cap, 2 of 3 slots; all-in with L2 and L1 fallbacks $163,750. Context Specificity downgraded to L2 ($7K): it fails in plain sight, so the user's re-ask rate plus a weekly audit detects it, while attribution fails silently in the client's own audit. Bias L2, cost overruns L1 with a hard per-request token cap. The third slot stays empty: the $50K headroom cannot buy Bias ($55K) at Level 3.

Reflection

Good enough is a number with a sample size behind it

One realization

Twenty rows showed a fabrication problem but could not show whether a fix worked, so the hold ends in an audit, not a date. The judge is part of the eval: one rubric paragraph moved its agreement with humans from 0.286 to 0.824. One call I got wrong first: a Level 3 judge on context misses, which users catch themselves, instead of on citations, which fail where nobody looks. The bet stops for a CPO decision if the audit fails twice.

Next sprint challenge

Building and running the 300-claim held-out audit is the critical next investment: it is the measurement that ends the hold.

Submission

Ascend IQ · AI Evals Final Project

Repository
https://github.com/antje/product-school-ai-evals
Antje Barth · AI Evals Certification · Oct 2026 Submit to the learning platform →
M1 · Strategystrategy-canvas.md
M3 – M4 · Proof + Standardsproduct-school-ai-evals
M5 – M6 · DecisionHOLD