Build1 distinct publisher3 min readPublished
Agreement between LLM providers looks like a safety net for self-healing tests. But the measured deleted-element arm shows a panel that is confidently unanimous and wrong in the same breath, no matter how many providers you add.
The Engineer · Build desk
invest
Touchmark opens a forwards market for tokens because finance cannot forecast them1 distinct publisher
product
Nvidia's $6bn Poolside licence is the third run of the same play2 distinct publishers
product
The criminal AI market is a reseller business, and Grok's abuse desk is the chokepoint1 distinct publisher
build
Qwen 3.8 27B ships thinking at maximum, and one setting stands between you and 22,000 tokens1 distinct publisher
Compiled by The EngineerSomething wrong?How this is made
A forced choice explains the shape of the failure. In a removal-tier scenario the whole target subtree is gone from the mutated tree [10], so the candidate list the panel sees contains only survivors. Each model is asked which candidate is the successor. The nearest sibling of the hole is the best available answer to that question, and every model reasons over the same tree, so the errors correlate. That agreement looks like corroboration, but it is really several systems sharing the same bias. In seven cases the models came from three separately sourced families and still landed on the same non-existent element [4]. Buying a fourth provider buys another vote in the same room.
Combine the two arms and the panel's overall precision falls out: 52 plus 34 is 86 unanimous verdicts inside 133 usable scenarios [17], of which 52 were correct, so unanimity was right 60.5 percent of the time [18]. That figure does not transfer, and the sampling says why. Token budget targeted 25 compound-drift and 42 removed-element scenarios [13], so deletions are enormously over-represented relative to a real release. On production traffic the same panel would score much higher without gaining any ability to detect the case that matters, because the deletion arm stays at zero [3].
The experiment existed because the deterministic scorer cannot draw the line either. Decoy scores for deleted elements span 0.665 to 0.955, and true compound drift spans 0.749 to 0.874 [8]. The second interval sits entirely inside the first [20], so any cutoff that admits genuine drift also admits decoys, and the author reports that runner-up margin, cluster density and control-type filters do not separate them [8].
That leaves the commit gate carrying the load. In Automation Sandbox the LLM is an opt-in fallback behind an independent-agreement quorum, and a heal is committed only after the retried action actually succeeds [7]. Read in light of the numbers, the quorum is a cost filter rather than a correctness filter: it reduces how often you spend a retry on an element that is not there, and the retry is what stops the false heal from passing green over a real regression [6][15].
The read deserves some caveats. The measurement supports its author's own architecture and was first published on the project site, with a Benchmark and Calibration page named as the source of truth for every number [16]; the protocol at least records every provider's raw vote at temperature 0 and counts a scenario only when two or more providers answered [12]. The identifiers are also deliberately stripped: every candidate AutomationId in a mutated tree is rewritten to `ablation-` plus SHA-256 hex, which the author calls a lower bound because a descriptive `btnSaveDocument` would give a model more to work with [11]. For the moved-and-relabelled tiers, that caveat cuts in the study's favour. For the removal tier the caveat has nowhere to land, since a descriptive identifier for a control that has been deleted is not in the tree to be matched at all.
What would have to be true for the 52 out of 52 to transfer to your suite: your refactors would need to be dominated by renames, relabels and layout moves, which is the easy 90 percent the structural scorer already handles [14]. The 34 out of 34 is the part that transfers unchanged, because it does not depend on your naming conventions.
Ranked by verification strength, evidence, and original report placement.
On elements that had been deleted, unanimous agreement was wrong 34 out of 34 times, every time a confident pick of the wrong neighbour.
Four live multi-provider runs, dated 2026-08-16 to 2026-08-18, produced 133 usable scenarios.
On elements that still existed (moved or relabelled), unanimous provider agreement was correct 52 out of 52 times.
In 7 of those 34 cases, three independently sourced model families (Cloudflare/Qwen, Mistral, OpenRouter/gpt-oss) agreed on the same non-existent element at once.
Widening the provider pool from 3 to 7 did not reduce the failure rate: agreement on deleted elements ran 25%, 45%, 58%, 40% with no downward trend, and was 100% wrong throughout.
The author states the consensus check is real protection, but that it comes from providers disagreeing with each other rather than from any model recognising that the element is gone.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 30, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Unusually disciplined, entirely unreplicated
The methodology is stronger than the sample. Denominators are kept separate by mutation type, all candidate IDs are hashed so a model cannot spot the odd one out, votes are logged at temperature 0, and the author publicly withdraws his own earlier 33 percent figure for pooling the two arms. Against that: 34 deleted-element verdicts on one live WPF capture, a second application whose LLM arm is never reported, and a benchmark guide named as the source of truth that nobody outside the project has opened.
One codebase, two desktop apps, no outside users
Everything observable is inside the author's own project: the gate it ships, the 0.50 default its scorer runs at, and mutations of HandBrake and ShareX trees. No other team, suite or vendor is shown adopting the quorum pattern or the finding, and the providers named appear only as endpoints under test.
Sound result stretched into a general law
The framing outruns the sample by a modest margin. '34 out of 34' and 'no matter how many providers you add' come from four runs topping out at seven providers on one WPF application, and the deleted-element subset was deliberately capped at 42 scenarios for token cost. What keeps the overshoot small is that the author argues against the flattering version of his own numbers, refusing the pooled accuracy figure and conceding the hashed IDs make the scores a floor rather than a verdict on production naming.
A negative result that vindicates the author's architecture
The experiment fails in exactly the direction that justifies how Automation Sandbox is already built — LLM as opt-in fallback, quorum gate, heal committed only after the retry succeeds. That is not disqualifying, and the author flags the arrangement up front by noting the piece first ran on the project site and pointing at his own benchmark guide as the authority. But the person who ran the rig, chose the mutation tiers, set the sampling caps and wrote the conclusion is the person whose design the conclusion endorses.
Trust the mechanism, hold the magnitude loosely
Two things here are likely to survive scrutiny: that model panels have no way to represent absence, and that provider diversity is not statistical independence when three families converge on the same phantom control. The precise numbers are another matter — one author, one rig, 34 headline cases, no replication, and the underlying data page outside our view. Treat the direction as informative and the percentages as provisional.