The numbers behind the numbers.
We maintain a standing forbidden-claims list and a claims sheet where every published figure has a dataset and a denominator. This page is the public version.
The production-call headline was re-measured on a second, later, fully quarantined 1,000-call corpus before publication.
| Claim on site | Value | Dataset / population | n | 95% CI | Notes |
|---|---|---|---|---|---|
| Real AI-agent call detection (call level) | 99.2%; replication 98.2% | production human-to-AI calls, one platform's live traffic, Jan–Jul 2026 | 500 unseen calls; replication: 500 fresh calls | [98.0, 99.7]; replication [96.6, 99.1] | frozen threshold, ≥1 high window; two corpora with zero identifier overlap with training; operating point frozen before either was scored |
| False-positive rate, live production callers | 1.8% | operator-labeled human calls, one platform's production traffic, listening-adjudicated | 500 calls | [0.7, 3.7] | recommended ≥2-window alert rule; measured floor on verified pure-human calls 0.78% (3/387) |
| AI or recorded machine speech inside human-labeled calls | ≈21%; 99.3% detected across the full corpus | one platform's production traffic, Jan–Jul 2026 | 500 human-labeled calls; 1,000 calls in the full corpus | — | screeners, voicemail, and recorded prompts; case-study observation, not an industry statistic |
| Partial-synthetic call detection | 95% | 60 s calls, ~20 s synthetic content, both conditions | 20 calls | [76.4, 99.1] | controlled splice test; final + dev partitions |
| Held-out-voice TTS recall (monitoring band) | 100% | telco condition, voices and accents absent from training | 240 | — | final partition; strict band 87.5% |
| Time to first verdict | ~5.5–6 s | measured, single host | — | — | warm per-family p95 57–474 ms |
| Window FPR (benchmark speech) | 0.4% clean / 0.6% telco | LA21-eval bonafide through real PSTN/VoIP codecs — not spontaneous platform callers; final partition | 500 windows / condition | clean [0.1, 1.4]; telco [0.2, 1.7] | frozen per-condition high bands |
| Compute | ≈$3.90 / 1,000 calls | on-demand CPU, 3-min avg calls, single-host measured | — | — | ~$2.40 reserved |
An approved detector update is in governed shadow rollout. Its performance remains an internal candidate result until it ships to the pilot package.
What we won't claim — perfect detection, coverage of all synthetic voices, industry-wide statistics from one platform's traffic, field false-positive rates without an adjudication protocol, or automation that "stops fraud." The product is advisory, and says so.
What we've tested. What we haven't. Labeled.
"Measured" means numbers with confidence intervals exist. "Exercised" means demonstrated end to end without statistical claims. "Roadmap" means not yet tested — and we say so.
| Modern neural TTS, held-out voices and accents | MEASURED |
| Unseen TTS architecturedocumented gap at strict band; caught at monitoring band | MEASURED |
| AI agent calls over real telephonycall-level, two production cohorts; 500 unseen calls; replication: 500 fresh calls | MEASURED |
| Partial-synthetic injection (spliced calls)n=20 | MEASURED |
| Telco codec launderingattacks traverse the real phone network by construction | MEASURED |
| Degraded audio (noise, packet loss, reverb) | EXERCISED |
| Replay / re-recordingapproximation tested; true replay is a pilot-phase test | EXERCISED |
| Voice conversion / zero-shot cloning | ROADMAP |
| Very short utterances (1–3 s)measured limitation — confidence demotion is the validated defense | MEASURED |
Our evaluation harness manufactures labeled attack calls through the actual phone network. A new synthesis system becomes test traffic in hours, with per-vendor recall built into the gates.