An AI agent that answers plain-language competitive-intelligence questions for VP-level strategists at Fortune 500 accounts, grounded in Ascend Analytics' verified data.
VP-level strategists and product leaders at Fortune 500 companies who pay $50k+ a year for verified market intelligence · getting a specific, verified competitive insight from a plain-language question.
An answer they can put in front of their board without checking the source, in seconds instead of hours, so that Ascend keeps its $50k+ renewals through Q4.
share of factual claims with no supporting span in the retrieved source
P95 time to a complete answer
answers that stay correct and sourced on messy or out-of-scope questions
Hallucination rate over latency: retrieval and source checking cost seconds, and a wrong answer in front of a board ends a $50k relationship.
Robustness over latency for the launch cohort: ask or refuse on an ambiguous question rather than answer fast.
a detail the source does not support or contradicts (a seat minimum, a 'confirmed' speaker, a funding stage)
a true fact with the condition that makes it false removed ('native' SQL export, 'seamless' HubSpot)
(SOC2 badge was in the footer)
This failure matters because Ascend IQ stated a specific the source does not support in 6 of 20 audited answers, and at that rate a VP asking ten questions has a 97% chance of carrying one into a board deck, which results in the first caught fabrication ending a $50k renewal and nearly every account in the 50-account launch cohort meeting one in its first week.
Three-layer suite run on the P0 pricing case and all 20 audit rows: the LLM judge caught 9 of the 11 confirmed failures (9 hallucinations, 1 retrieval miss, 1 P2 brand-voice miss outside the top three), the two deterministic layers caught 1 between them, so the P0 class is semantic. The usage-drop trajectory scored 1 of 6 dimensions (HOLD): the agent guessed 'seasonal' without ever checking ingestion. The judge reached Cohen's kappa 0.824 against human labels after two rubric revisions (from 0.286), measured on the 12 calibration traces it was tuned on; a held-out recalibration comes before the audit.
Target risk: unsupported claims, covering both P0s (fabricated specifics and dropped qualifiers). Threshold: 0 human-confirmed unsupported claims in a 300-claim held-out audit, which bounds the per-claim rate under 1% at 95% confidence.
Hard gate: the release does not ship to any account until the audit passes; a failed audit re-runs on a fresh 300. In CI, faithfulness blocks the merge below 95 or on a regression of more than 3 points (PR #218 was blocked at -9).
Clients pay for verified data, so one invented number in a board deck ends the account; per the M1 canvas we put hallucination rate ahead of latency. The 20-row beta log cannot prove the bar (zero in 20 only bounds the rate under 15%), which is why the gate is a 300-claim audit. Latency is a Soft gate (P95 at or under 2.0s, override band to a 10s ceiling under a staged rollout) and brand voice is advisory.
I recommend we hold the Ascend IQ launch to our top 50 accounts until it passes the zero-fabrication audit, with the go/no-go on Nov 9, because in our beta 6 of 20 answers stated a detail the source does not support, and shipping now puts more than $2.5M of annual renewals in front of an answer a VP has a 97% chance of catching out within ten questions.
The 3 ArgumentsBrand risk. The failure we would ship is the one clients pay us to prevent: invented specifics and dropped qualifiers break 'use it without checking the source', and none of the 20 beta answers cites a source.
Revenue risk. The exposure lands on our most valuable accounts, all at once, where we can't take it back: an invented number goes into a client's board deck, so we learn about it from them.
Reliability risk. We can't yet prove it's fixed, but we know exactly what would: the judge (kappa 0.824 on its tuning set), the CI gate (PR #218 blocked) and Level 3 coverage are in place; the 300-claim audit is the missing proof.
Holding costs a few weeks, not trust. Approve by Fri Oct 2; fixes by Oct 23; audit from Oct 26; go/no-go Nov 9.
Pass: 10 accounts go live behind flags, all 50 after 14 clean days. Fail twice: back to the CPO with a new date.
Riskiest assumption: that VPs open the citations. Wave 1 measures it.
Portfolio context, a sibling product (AI-Powered Report Summaries): hallucination covered only by a proxy judge, latency covered, bias, toxicity and drift uncovered. Toxicity accepted on low expected impact with a kill criterion; drift is the critical gap, closed by a pre-send grounding check and a new-template launch gate. The same pre-send pattern is a launch condition for Ascend IQ.
Level 3: Data Fabrication ($85K) + Source Attribution ($65K) = $150K of the $200K cap, 2 of 3 slots; all-in with L2 and L1 fallbacks $163,750. Context Specificity downgraded to L2 ($7K): it fails in plain sight, so the user's re-ask rate plus a weekly audit detects it, while attribution fails silently in the client's own audit. Bias L2, cost overruns L1 with a hard per-request token cap. The third slot stays empty: the $50K headroom cannot buy Bias ($55K) at Level 3.
Twenty rows showed a fabrication problem but could not show whether a fix worked, so the hold ends in an audit, not a date. The judge is part of the eval: one rubric paragraph moved its agreement with humans from 0.286 to 0.824. One call I got wrong first: a Level 3 judge on context misses, which users catch themselves, instead of on citations, which fail where nobody looks. The bet stops for a CPO decision if the audit fails twice.
Building and running the 300-claim held-out audit is the critical next investment: it is the measurement that ends the hold.