The numbers behind the numbers.
We maintain a standing forbidden-claims list and a claims sheet where every published figure has a dataset and a denominator. This page is the public version.
The fielded system was measured in two independent 1,000-call production evaluations. Each paired 500 human-to-AI calls with 500 human-labeled calls; the first ran end to end through the shadow pipeline, and the second was a temporally fresh, zero-overlap replication with deployment-matched preprocessing.
| Claim on site | Value | Dataset / population | n | 95% CI | Notes |
|---|---|---|---|---|---|
| Real AI-agent call detection (call level) | 99.2%; replication 98.2% | two independent 1,000-call production evaluations, one platform's live traffic, Jan–Jul 2026 | 500 human-to-AI calls per evaluation; 1,000 total calls per evaluation | [98.0, 99.7]; replication [96.6, 99.1] | frozen threshold, ≥1 high window; first evaluation end to end through the shadow pipeline; second evaluation temporally fresh, zero-overlap, and scored offline with deployment-matched preprocessing |
| False-positive rate, live production callers | 1.8% | operator-labeled human calls, one platform's production traffic, listening-adjudicated | 500 calls | [0.7, 3.7] | recommended ≥2-window alert rule; measured floor on verified pure-human calls 0.78% (3/387) |
| Raw flag rate in human-labeled production cohorts | 24.4% → 31.0% | two independent production evaluations, one platform's human-labeled traffic, Jun–Jul 2026 | 500 calls per cohort | first [20.8, 28.4]; second not reported | before listening adjudication; not an FPR; includes genuine AI screeners, voicemail, and recorded prompts |
| AI or recorded machine speech inside human-labeled calls | ≈21%; 99.3% detected across the full corpus | one platform's production traffic, Jan–Jul 2026 | 500 human-labeled calls; 1,000 calls in the full corpus | — | screeners, voicemail, and recorded prompts; case-study observation, not an industry statistic |
| Frozen production operating point | 0.935; production-derived optimum 0.9335 | adjudication-aware platform recalibration study on the first 1,000-call production evaluation | 500 synthetic calls; 387 verified pure-human calls; 71,919 trio-covered windows | not applicable — operating-point sweep | frozen before production scoring; production-derived optimum landed within 0.0015 |
| Partial-synthetic call detection | 95% | 60 s calls, ~20 s synthetic content, both conditions | 20 calls | [76.4, 99.1] | controlled splice test; final + dev partitions |
| Held-out-voice TTS recall (monitoring band) | 100% | telco condition, voices and accents absent from training | 240 | — | final partition; strict band 87.5% |
| Time to first verdict | ~5.5–6 s | measured, single host | — | — | warm per-family p95 57–474 ms |
| Window FPR (benchmark speech) | 0.4% clean / 0.6% telco | LA21-eval bonafide through real PSTN/VoIP codecs — not spontaneous platform callers; final partition | 500 windows / condition | clean [0.1, 1.4]; telco [0.2, 1.7] | frozen per-condition high bands |
| Compute | ≈$3.90 / 1,000 calls | on-demand CPU, 3-min avg calls, single-host measured | — | — | ~$2.40 reserved |
A maintainer-approved acoustic-head update completed the governed candidate → evidence → validation → applied-shadow flow. Its pilot-package change remains under review, so candidate performance is not presented here as fielded product performance.
What we won't claim — perfect detection, coverage of all synthetic voices, industry-wide statistics from one platform's traffic, field false-positive rates without an adjudication protocol, or automation that "stops fraud." The product is advisory, and says so.
What we've tested. What we haven't. Labeled.
"Measured" means numbers with confidence intervals exist. "Exercised" means demonstrated end to end without statistical claims. "Roadmap" means not yet tested — and we say so.
| Modern neural TTS, held-out voices and accents | MEASURED |
| Unseen TTS architecturedocumented gap at strict band; caught at monitoring band | MEASURED |
| AI agent calls over real telephonycall-level, two independent production evaluations; 500 human-to-AI calls per evaluation; 1,000 total calls per evaluation | MEASURED |
| Partial-synthetic injection (spliced calls)n=20 | MEASURED |
| Telco codec launderingattacks traverse the real phone network by construction | MEASURED |
| Degraded audio (noise, packet loss, reverb) | EXERCISED |
| Replay / re-recordingapproximation tested; true replay is a pilot-phase test | EXERCISED |
| Voice conversion / zero-shot cloning | ROADMAP |
| Very short utterances (1–3 s)measured limitation — confidence demotion is the validated defense | MEASURED |
Our evaluation harness manufactures labeled attack calls through the actual phone network. A new synthesis system becomes test traffic in hours, with per-vendor recall built into the gates.