Build1 distinct publisher3 min readPublished
A one-author benchmark held the audio and the trial list fixed and varied only the model, and the spread is wider than the architecture arguments people usually have about speaker verification. Voice-auth thresholds sit downstream.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The worst instrument in the panel fails in one direction only. Its genuine scores, the same speaker twice, sit close to the rest of the pack; the impostor pairs are where the damage shows up [8]. It reports two different people as roughly eight to ten times more similar than the better encoders do on the same pair [4]. That is what a compressed embedding space looks like from outside: everybody sits near everybody, and any threshold you fit inherits the crowding [8]. Neutral speech already shows it, at 0.031 EER where several ReDimNet checkpoints reach 0.000 on the identical trials [9]. The weakness is already there in neutral speech; shouting and whispering do not create it, they just expose it [9]. The margins are wide, running from +0.087 over ECAPA-TDNN at the near end to +0.185 over the lowest-error encoder at the far end, both intervals clear of zero [7]. Subtract the far margin from the worst error rate and you land on 0.048, within rounding of the best score on the list, which is a decent check that the paired design is doing what it claims [5]. Nothing on the model card predicts the ordering. Spearman rho between parameter count and expressive EER is +0.135 across the fourteen; the 20.8M-parameter model ranks 13th and a 4.81M model ranks 1st [11], so about 4.3 times the parameters bought a worse result [6]. Published VoxCeleb1-O position does not carry over either, because that leaderboard is scored on calm read speech [12]. A leaderboard on calm read speech is an excellent instrument for calm read speech. The caveat that matters downstream is about CAMPPlus. A widely used open-source TTS system conditions on it, which the dev.to writeup verifies from source code and then declines to turn into a claim about that system's audio; the encoders were measured in isolation on human recordings, and whether the conditioning choice degrades synthesis is an experiment the author says he has not run [10]. For the spread to be your spread, your pipeline has to look like this one: enrol on neutral speech, test the same person under phonation change, score raw cosine, and treat the speaker as the independent unit on both sides of a trial [5][18]. That last constraint drives the statistics. The resample is a vertex bootstrap: draw eight speakers with replacement, take the induced sub-multigraph over the 56 ordered impostor edges, weight each edge by the product of multiplicities [18]. It is also where the bug lived. The EER estimator's argmin returned the first minimiser rather than the balanced operating point, invisible at unit weights and common once a bootstrap hands it integer speaker multiplicities [15]. Fixing it moved one comparison across the Holm threshold, 28 to 29 of 91, and 91 is every pair of the fourteen [15][2]. Two of the same author's earlier claims did not survive that discipline. An 11.7 dB spectral effect was asserted with an instrument whose F0 artifact budget was 12.09 dB, and sweeping F0 with the envelope held fixed moved the descriptor further than the claimed effect [16]. A separate "shouting does not transfer" result was judged on RMS after every clip had been peak normalised, so the figure was crest factor [17]. The winner went the same way: it survives as an argmin, chosen in 89.4% of speaker resamples and leading all eight leave-one-speaker-out refits, now worded as lowest observed error [13]. One check that could have sunk the whole measurement came back quiet. Spectral denoising moved EER by at most 0.039 across 10,765 twinned trials, and whispered speech, where stripping aspiration noise was the stated worry, was among the least affected [14]. The author's reading is that the spread is larger than most architectural differences people argue about, and that for expressive audio the encoder is the decision [19].
Ranked by verification strength, evidence, and original report placement.
The trial list is a directed graph on 8 speakers, with genuine trials as self-loops and impostor trials as the 56 ordered edges; the independent unit is the speaker on both sides, so the resample is a vertex bootstrap that draws 8 vertices with replacement, takes the induced sub-multigraph, and weights each impostor edge by the product of multiplicities.
Fourteen speaker encoders were run over an identical, frozen list of 11,935 trials; the trial list was frozen in writing before a single model loaded, and every encoder scored the identical list, so every comparison is paired.
Equal error rate spans 0.047 to 0.233 across the panel, with the same trials and same audio and nothing varying but the encoder.
29 of 91 pairwise comparisons survive Holm-Bonferroni correction.
Eight speakers each recorded the same 1,360 sentences in six phonation states (neutral, happy, angry, scared, shouting, whisper), so lexical content is fixed while phonation varies; other emotional-speech corpora use different sentences for different emotions, letting a model learn angry vocabulary instead of angry delivery.
Enrolment is neutral speech, and the test is whether the encoder still recognises the person while they shout or whisper.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Eleven agent sessions on one machine settled CPU contention by writing to each other1 distinct publisher
build
Uniform INT4 beat NVFP4 in 85 of 90 real gradient tests, and the rotation barely moved them1 distinct publisher
build
Harness choice moved token use 83-fold with the model held constant1 distinct publisher
build
The repo's own control run deleted the 5-10x WASM claim from vizcrush's launch copy1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Meticulous, and corroborated by nobody
Every figure in this story — the 11,935 trials, the 0.047-to-0.233 range, the 29 surviving pairs, the 20,000-draw bootstrap — comes from one dev.to post by one author, and the internal discipline is better than most peer-reviewed work: the trial list was written down before any model loaded, families were declared before testing, and the numbers survive an arithmetic cross-check (0.233 minus the widest margin of 0.185 lands on 0.048, within rounding of the stated floor). What is absent is anything outside the author's own hands: no trial list, no code, no checkpoint identifiers, no second run. And the encoder that loses all thirteen of its comparisons is never named — the reader is left to infer it from a CAMPPlus sentence two paragraphs later.
Nothing yet shows anyone acting on it
We can see the work; we cannot see uptake. This reporting contains no instance of another team re-scoring the frozen list, no enrolment policy changed, no voice-authentication threshold moved. The only datapoint pointing at real deployments is second-hand and thin: an unnamed open-source speech system said to condition on CAMPPlus, reported from source code, with the author refusing to draw any conclusion about its output. That is too little to score.
Hedged harder than it needed to be
The rhetoric runs behind the result rather than ahead of it. A writer chasing attention had two easy escalations available — name CAMPPlus as the encoder that lost every comparison, and let readers connect it to the text-to-speech system it conditions — and this author declines both, plus retracts a spectral effect, a shouting claim and his own choice of winner. The offsetting overreach is framing: 'five-fold' is true arithmetic on eight speakers recorded in one studio where enrolment and test share a session, and our own dek's line about voice-auth thresholds sitting downstream is an extrapolation the measurement does not carry.
No sponsor, no referee either
This is self-published under a studio handle on dev.to, which means nobody paid for the conclusion and nobody checked it. The tells run against self-interest: a tie-break fix is disclosed precisely because it helped the author's own paper, a winner claim is retracted, and the CAMPPlus-degrades-synthesis inference — the most shareable thing in the piece — is explicitly left unmade. What remains is the softer incentive of building a reputation on instrument-validation rigour, and the fact that the labs behind the encoders finishing last were given no chance to answer.
Trust the ordering, not the rates
The paired design over identical trials and a speaker-level bootstrap make the relative ranking credible, and the author is candid that only 29 of 91 pairs are actually separable — so most encoder swaps here are indistinguishable, including seven of the thirteen involving his best model. The absolute numbers deserve much less weight: eight speakers, one studio, one recording chain, enrolment and test in the same session, which he flags himself as making error rates optimistic and non-comparable to VoxCeleb. Until someone else scores the same list, treat the direction as informative and the magnitudes as local.