Build1 distinct publisher3 min readPublished
The three-model set topped out at a US plus China pairing, so Mistral Small 3.2 went in late; the resulting China plus EU pair posted the run's best score and its worst capitulation rate, for about $0.15 in stranded reviews.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The unit of coverage in a harness like this is the pairing, and pairings do not grow the way the model list does. Going from three models to four takes the distinct cross-model pairs from three to six [1]. What the run actually needed was one specific cell in that matrix, and the field test behind AdversarialDebate v0.2.2 [16] had no way to fill it without a fourth lab in a fourth region.
A three-PR run went first, to confirm the pipeline worked [3]. It did, but a smoke test only proves that much: it exercises the code paths for whichever pairings you configured, and each of those pairings returns a number that reads like a finding regardless. The author, writing on dev.to, says that stopping at the small corpus would have produced a clean story and the wrong one [19].
The waste is worth doing by hand. By the time Mistral joined, the first three models had each reviewed about 148 PRs, while Mistral finished about 70 to 73 [12]. Mistral pairs were then limited to the 70 PRs all four models shared [13]. That leaves roughly 78 passes per model outside the intersection, about 234 across three models, which brackets the 228 stranded passes reported [2]. At the stated $0.15, each stranded single-pass review cost around $0.0007 [3]. Cheapest design review I have been shown a receipt for.
Treat the winning pair's average as a claim about a corpus rather than a ranking of models. For it to mean that DeepSeek plus Mistral beats GPT plus Gemini, every pair has to be scored on the same 70 shared PRs, under the same rubric, with the difficulty spread of those 70 resembling the rest of the 148. The post states the intersection restriction for the Mistral pairs [13]. If the older pairs kept their full history in the comparison, then a 70-PR sample is being read against a 148-PR one.
The behavioural numbers need the same handling. About 34 concessions per shared PR [4] is a lot of ground for one pair to give in a review. The cascade count over those same 70 PRs works out to 63% [5], close enough to the reported in-pair capitulation rate [9] that the denominator is probably debates or PRs rather than turns. Pin that down before comparing the figure to any other system, because a per-turn rate and a per-debate rate of the same value describe very different failures.
What I would change is the order, not the decision. Enumerate the corners the thesis needs first, homogeneous through strong diversity [7], then pick models to hit them, then freeze the corpus. The author's own summary is that three models proved the system ran and four made pairing strategy learnable [18], and the wider spectrum also promoted the same-model control from curiosity to evidence that weak diversity is its own failure mode [17]. That is a statement about design coverage. Design coverage is cheaper to fix on a whiteboard than at pass 148.
Ranked by verification strength, evidence, and original report placement.
The AdversarialDebate field test began with three models: GPT-4o-mini, Gemini 2.5 Flash, and DeepSeek-V3.
The three-model set gave three useful pairings: GPT + Gemini, Gemini + DeepSeek, and GPT + GPT as a homogeneous control, described as three labs, two regions, one same-model control.
The author ran a small corpus of just 3 PRs first, to validate the pipeline.
With those three models the farthest useful pairing available was US + China, so the test could suggest whether diversity helped but could not show whether maximum diversity behaved differently from moderate diversity.
The author added Mistral Small 3.2, which created three new pairings: GPT + Mistral, Gemini + Mistral, and DeepSeek + Mistral.
DeepSeek + Mistral created the strongest diversity pairing in the run: China + EU.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 30, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Shanghai AI Lab's 397B science agent shipped in July; the paper explaining it landed August 131 distinct publisher
build
A 170-goal agent field test costs $0.49. Proving it actually passed costs more.1 distinct publisher
build
Armenian ASR leaderboard: closed models take the top eight, then lose the domains that matter1 distinct publisher
invest
Debate wins the agent bake-off, then loses to one model on the same budget1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One run, one hand, no artifacts
Every figure that matters — 0.982, the 97% verdict rate, 2,352 concessions, 44 cascades, 228 stranded passes — traces to a single developer's write-up of his own run, with no logs, no scoring definition and no rerun. What the numbers do have going for them is that they check out against each other: the 148-versus-70 coverage gap implies about 234 unusable passes, which brackets the 228 claimed, and 44 cascades over 70 shared reviews is 63% against a stated 65%. Arithmetic that closes is not the same as a result anyone else has seen.
A tag and its author
Uptake amounts to a v0.2.2 release line dated the day before publication and one field test the author ran himself. No other team's pipeline, no reviewer using the debate output on real pull requests, no downloads or forks are described. The 70 shared pull requests are the entirety of the observed usage.
Modest narrator, oversized rule
The framing is unusually self-critical: the waste comes before the win, the best pair is called the least safe, and the author admits his model set was not planned tightly enough. What overreaches is the generality of the takeaway. 'Maximum diversity can create capitulation cascades' rests on exactly one China + EU pair, and 'weak diversity can be worse than no diversity' on exactly one same-model control, each measured over 70 reviews. A rule about pairing strategy is being drawn from one cell of the matrix apiece.
Own project, receipts included
This is a maker writing up his own tool the day after tagging a release, and that shapes which mid-run mistake becomes a lesson rather than a defect. Pulling the other way, he prices his own error at fifteen cents, publishes the fifty-three-cent total for the run, and states plainly that the model set had not been planned tightly enough — the sort of detail a purely promotional post leaves on the floor. No vendor money, sponsorship or third-party stake appears anywhere in the account.
Sound argument, unchecked data
Split the story in two and confidence splits with it. The design reasoning — three models could reach no farther than US + China, a fourth opened China + EU, six pairs where there were three — is verifiable by inspection and holds up. The measurements attached to it come from one author, one run, one seventy-review corpus, with nothing published for anyone to audit. Trust the lesson about experiment design well ahead of any number in it.