01METHODOLOGY

The numbers behind the numbers.

We maintain a standing forbidden-claims list and a claims sheet where every published figure has a dataset and a denominator. This page is the public version.

The fielded system was measured in two independent 1,000-call production evaluations. Each paired 500 human-to-AI calls with 500 human-labeled calls; the first ran end to end through the shadow pipeline, and the second was a temporally fresh, zero-overlap replication with deployment-matched preprocessing.

Claim on siteValueDataset / populationn95% CINotes
Real AI-agent call detection (call level)99.2%; replication 98.2%two independent 1,000-call production evaluations, one platform's live traffic, Jan–Jul 2026500 human-to-AI calls per evaluation; 1,000 total calls per evaluation[98.0, 99.7]; replication [96.6, 99.1]frozen threshold, ≥1 high window; first evaluation end to end through the shadow pipeline; second evaluation temporally fresh, zero-overlap, and scored offline with deployment-matched preprocessing
False-positive rate, live production callers1.8%operator-labeled human calls, one platform's production traffic, listening-adjudicated500 calls[0.7, 3.7]recommended ≥2-window alert rule; measured floor on verified pure-human calls 0.78% (3/387)
Raw flag rate in human-labeled production cohorts24.4% → 31.0%two independent production evaluations, one platform's human-labeled traffic, Jun–Jul 2026500 calls per cohortfirst [20.8, 28.4]; second not reportedbefore listening adjudication; not an FPR; includes genuine AI screeners, voicemail, and recorded prompts
AI or recorded machine speech inside human-labeled calls≈21%; 99.3% detected across the full corpusone platform's production traffic, Jan–Jul 2026500 human-labeled calls; 1,000 calls in the full corpus—screeners, voicemail, and recorded prompts; case-study observation, not an industry statistic
Frozen production operating point0.935; production-derived optimum 0.9335adjudication-aware platform recalibration study on the first 1,000-call production evaluation500 synthetic calls; 387 verified pure-human calls; 71,919 trio-covered windowsnot applicable — operating-point sweepfrozen before production scoring; production-derived optimum landed within 0.0015
Partial-synthetic call detection95%60 s calls, ~20 s synthetic content, both conditions20 calls[76.4, 99.1]controlled splice test; final + dev partitions
Held-out-voice TTS recall (monitoring band)100%telco condition, voices and accents absent from training240—final partition; strict band 87.5%
Time to first verdict~5.5–6 smeasured, single host——warm per-family p95 57–474 ms
Window FPR (benchmark speech)0.4% clean / 0.6% telcoLA21-eval bonafide through real PSTN/VoIP codecs — not spontaneous platform callers; final partition500 windows / conditionclean [0.1, 1.4]; telco [0.2, 1.7]frozen per-condition high bands
Compute≈$3.90 / 1,000 callson-demand CPU, 3-min avg calls, single-host measured——~$2.40 reserved
FALSE-POSITIVE RATES ON LIVE CALLERSThe raw production flag rate is not the false-positive rate. Listening showed that the large majority of reviewed flags were the detector correctly identifying machine speech inside calls labeled human: AI call screeners answering before their owners, voicemail systems, and recorded prompts. The raw rate rose across the two independent cohorts and is reported separately, before adjudication. The listening-adjudicated false-positive rate on live production callers is 1.8% [0.7, 3.7] on 500 human-labeled calls at the recommended ≥2-window alert rule. We publish the raw signal, the adjudication method, and the decomposition, because field FPR without listening measures label quality as well as detector behavior.
AI-INTERMEDIARY DISCOVERY · CASE STUDYAcross the first of this one platform's two 1,000-call production evaluations, about 21% of 500 human-labeled calls contained AI or recorded machine speech, and Norvin detected AI speech in 99.3% of the full corpus's estimated AI-containing calls. This is a case study, not an industry statistic.
BENCHMARK-TO-FIELD CALIBRATIONThe operating point frozen before either production evaluation landed within a narrow margin of the production-derived optimum without post-hoc threshold tuning. The production sweep also reaffirmed the existing fusion weights rather than selecting a single-family shortcut.
WHY THE THIRD FAMILY STAYSThe incumbent third family remains because it anchors the fused score distribution. Removing it forced the gated operating point into score saturation and damaged recall; pretrained replacements inverted on real telephony. Its value is distributional stability, not standalone detection.

A maintainer-approved acoustic-head update completed the governed candidate → evidence → validation → applied-shadow flow. Its pilot-package change remains under review, so candidate performance is not presented here as fielded product performance.

What we won't claim — perfect detection, coverage of all synthetic voices, industry-wide statistics from one platform's traffic, field false-positive rates without an adjudication protocol, or automation that "stops fraud." The product is advisory, and says so.

02THREAT MODEL

What we've tested. What we haven't. Labeled.

"Measured" means numbers with confidence intervals exist. "Exercised" means demonstrated end to end without statistical claims. "Roadmap" means not yet tested — and we say so.

Modern neural TTS, held-out voices and accentsMEASURED
Unseen TTS architecturedocumented gap at strict band; caught at monitoring bandMEASURED
AI agent calls over real telephonycall-level, two independent production evaluations; 500 human-to-AI calls per evaluation; 1,000 total calls per evaluationMEASURED
Partial-synthetic injection (spliced calls)n=20MEASURED
Telco codec launderingattacks traverse the real phone network by constructionMEASURED
Degraded audio (noise, packet loss, reverb)EXERCISED
Replay / re-recordingapproximation tested; true replay is a pilot-phase testEXERCISED
Voice conversion / zero-shot cloningROADMAP
Very short utterances (1–3 s)measured limitation — confidence demotion is the validated defenseMEASURED

Our evaluation harness manufactures labeled attack calls through the actual phone network. A new synthesis system becomes test traffic in hours, with per-vendor recall built into the gates.