When a human-labeled call contains AI speech
Listening adjudication uncovered screeners, voicemail, and recorded prompts inside the negative class. Why field false-positive claims fail without ears on the audio.
IN REVIEWMeasurements, null results, and what synthetic-voice detection actually looks like on the phone network. Written by the people who ran the evaluations.
The field-evaluation manuscript and supporting research notes are in review. No article ships before its claims clear the same evidence discipline as the rest of this site.
Listening adjudication uncovered screeners, voicemail, and recorded prompts inside the negative class. Why field false-positive claims fail without ears on the audio.
IN REVIEWA fresh, zero-overlap production cohort reproduced the fielded result while the raw human-label flag rate moved. The case for frozen operating points and adjudicated metrics.
IN REVIEWRemoving the third family destabilized the fused score geometry. Its value is distributional anchoring, not standalone discrimination.
IN REVIEWTelephony conditioning breaks the two strongest detector architectures in opposite directions. The measurements that made a third family non-optional.
IN REVIEWNeither synthetic codec conditioning nor real PSTN transmission reproduced our false positives. Spontaneous speech did. Domain gaps live in speech style, not codec math.
IN REVIEWWhy checkpoints selected on benchmark metrics produced telephony false positives, and what selecting on deployment audio looks like.
IN REVIEW