01RESEARCH

From the lab.

Measurements, null results, and what synthetic-voice detection actually looks like on the phone network. Written by the people who ran the evaluations.

STATUS

The field-evaluation manuscript and supporting research notes are in review. No article ships before its claims clear the same evidence discipline as the rest of this site.

RESEARCH

When a human-labeled call contains AI speech

Listening adjudication uncovered screeners, voicemail, and recorded prompts inside the negative class. Why field false-positive claims fail without ears on the audio.

IN REVIEW
RESEARCH

What replication changed — and what it did not

A fresh, zero-overlap production cohort reproduced the fielded result while the raw human-label flag rate moved. The case for frozen operating points and adjudicated metrics.

IN REVIEW
TECHNICAL

The detector that earns its slot without detecting

Removing the third family destabilized the fused score geometry. Its value is distributional anchoring, not standalone discrimination.

IN REVIEW
TECHNICAL

Opposite failure modes: why our ensemble is load-bearing, not marketing

Telephony conditioning breaks the two strongest detector architectures in opposite directions. The measurements that made a third family non-optional.

IN REVIEW
RESEARCH

Two null results that found the real domain gap

Neither synthetic codec conditioning nor real PSTN transmission reproduced our false positives. Spontaneous speech did. Domain gaps live in speech style, not codec math.

IN REVIEW
TECHNICAL

Your lab-perfect detector is wrong on the phone network

Why checkpoints selected on benchmark metrics produced telephony false positives, and what selecting on deployment audio looks like.

IN REVIEW