Skip to content

Build1 publisher3 min readPublished

A 1984 phase-retrieval algorithm moved 12 of 13 synthetic-speech detectors on genuine recordings

A pre-registered test held speaker, words, room and microphone fixed and processed genuine studio recordings in ways that generate no speech tokens. Then it measured what 13 frozen detectors did to the same utterance.

The Engineer · Build desk

Illustration accompanying A 1984 phase-retrieval algorithm moved 12 of 13 synthetic-speech detectors on genuine recordings

What happened

  • Thirteen synthetic-speech detectors were declared and hashed before any of them was scored, then run through four pre-registered experiments on the same 47 utterances from three speakers.
  • A single codec encode-decode pass over real studio recordings moved 9 of the 13 detectors past Holm correction.
  • Three of those detectors were fully prospective, with no behaviour observed before the protocol was frozen, and all three moved on 47 of 47 utterances at a matched-pairs rank-biserial of exactly -1.000.
  • An independent adversarial audit ran three rounds against the frozen estate by parsing the stored raw scores instead of the author's code, reproduced every statistic, and led to thirteen claims being withdrawn.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Swapping in a different checkpoint relocates the false positive instead of removing it, since the two families in the panel break on opposite inputs, so procurement becomes a choice of which error you are prepared to defend.
  • exposure The person carrying the flag is whoever submitted a genuine recording that got transcoded or mastered on the way in, and the score itself does not distinguish processing from generation.
  • decision Anyone wiring a provenance score into a pipeline now has to decide whether to score before normalisation and store the processing history alongside the number, or accept that the two are mixed.
  • precedent Hashed pre-registration plus an outside audit that reparses raw scores raises the bar for the next provenance claim, and an AUC table with no paired before-and-after scores looks thin against it.

The structural claim in the post is a confound in the training data. Every generated clip in a text-to-speech corpus has been through a vocoder or a neural codec, and almost no genuine clip has [1]. A model minimising training loss has no incentive to separate what its training data never separated [2]. So part of what these scores measure is whether the waveform was ever reconstructed [13].

Running that confound backwards is the whole design. Speaker, words, performance, room and microphone stay fixed. The only change is a transformation that produces no generated speech tokens, and the measurement is the paired change in a frozen detector's score on the same utterance [3]. There was no AUC gate, because admitting only the detectors that discriminate on this corpus would have selected for the codec sensitivity under test [5]. The panel also includes a detector already known to be inverted, put there so the set is visibly not picked to agree [6].

Before any of this transfers to your queue, the dull explanations have to go. Three professional microphones captured the same physical performance at the same time, and the effect held on all three, so it is not one capture chain's artefact [9]. It held on a speaker who was not in the corpus [10]. Griffin-Lim, phase retrieval published in 1984 with no neural network in it, moved 12 of 13 detectors in the same direction [11]. The inputs were studio recordings, so the condition for transfer is that your audio reaches the scorer with some reconstruction already in its history [7].

A denoise-and-master chain of the sort routinely applied in production moved two detectors by -2.95 and -4.22 in their native units [12]. The ASVspoof-era checkpoints in the panel could not separate real speech from synthetic speech on this material at all, and that same mastering chain still moved them on 41 and 46 of the 47 genuine human recordings [14]. That is 87 and 98 percent of the corpus [21]. The post does not report where any detector's decision threshold sits.

One pre-registered prediction failed. "I predicted that a detector trained on codec audio would resist a codec pass," the author wrote [17]. That checkpoint collapsed hardest of anything in the panel, and because the prediction was frozen in the pre-registration it is in the paper [18].

Which transformation moves which detector tracks each checkpoint's documented training exposure across three architectures [20]. The author reports that as an association among frozen checkpoints, not a causal effect of training data, because the checkpoints differ in frontend construction and optimisation as well as in corpus, and no isolating intervention was performed [20].

What to watch

  • Whether the paper, the frozen hashes and the thirteen checkpoint names are published in a form another team can rescore.
  • Whether any detector vendor publishes paired before-and-after scores for a codec pass on genuine audio next to its AUC table.
  • Whether the independent auditor publishes its three-round report separately from the author's withdrawals.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories