Build1 publisher3 min readPublished
Adversarial Verification Killed the Two Premises 60% of This Plan Needed
A dev.to run extracted 124 atomic claims from the author's own content strategy and sent 25 to verifiers told to default to refuted. Twelve died, among them both premises the plan was built on.
The Engineer · Build desk

What happened
- A dev.to post documents a research run over the author's own content strategy: 108 agents, 25 sources fetched, 124 atomic claims extracted, 25 adversarially verified, 13 confirmed and 12 refuted.
- About 60 percent of the author's planned work depended on two specific claims, and both of them landed in the refuted set.
- Stage three sent those 25 claims to independent verifiers whose instruction was to refute, with refuted as the default verdict whenever a verifier could not settle the question.
- Refuted claims were deleted from the plan; none were kept as risks to monitor, and the strategy was rebuilt only from the claims that survived.
- Verification stopped when the budget ran out, leaving 99 of the extracted claims recorded as leads rather than facts.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost At roughly four agent invocations per verdict, coverage is priced per claim. Anyone copying this decides in advance which claims are worth a panel and accepts that the rest stay unchecked.
- decision Defaulting to refuted swaps one error for another: false confirmations become rare and true claims that cannot be settled inside a fetch budget get killed. A team adopting the harness has to say which error it prefers.
- constraint Because only single-fetch assertions pass extraction, the harness can discipline the factual base of a plan while the positioning judgements sitting on top of it go untouched.
- precedent Recording 3-0 against 2-1 alongside each finding sets an expectation that machine review ships with vote provenance, so a reader can tell a settled claim from one that squeaked through.
Three verifiers and a two-thirds rule mean two votes settle a claim. A 2-1 refute kills it, a 2-1 confirm keeps it, and the split stays in the notes so a narrow survivor can be reopened [6][25]. A verifier that cannot settle the question votes refuted, and that default is what changes the outcomes [4]. "The fix isn't a better prompt. It's making refutation somebody's job," wrote the author, who posts on dev.to as shanni [8].
The filter that makes any of this checkable sits one stage earlier. Extraction threw out "our terminology positioning is sound" because nothing can verify it. "English Wikipedia's Four Pillars article defines Day Master in prose" passed because one fetch settles it [10]. Only claims of the second kind reach a verifier. The judgement calls in a plan sit outside the harness entirely.
The first premise died on exactly that kind of fetch. The English Wikipedia article, created in 2006, binds the head term in its lead sentence and carries no cleanup banners. It runs 4,200 words with 20 footnotes and a 10-row table of the taxonomy the planned pages were going to explain [12][13]. What closed it came from the MediaWiki revision API: a last edit of 2026-05-28 and an expansion of roughly 8,000 bytes in March [14]. "I wasn't proposing to fill a vacancy. I was proposing to out-write a twenty-year-old article that gets attention every month," the author wrote [28]. Published analysis of about 76.7 million AI Overviews plus roughly 957,000 ChatGPT and 953,500 Perplexity prompts found that those systems over-index on encyclopedic and UGC sources [15].
The second premise was that structured data would win citations, and JSON-LD work had been budgeted across the site [16]. Against it the post cites a matched difference-in-differences study. It covers 1,885 pages that added JSON-LD over eight months, each matched to three control URLs on other domains at similar pre-period citation levels, with a 30-day pre/post window and four statistical approaches including event-study weekly plots [17]. Three controls for each of 1,885 pages is 5,655, and the post puts the pool at roughly 4,000; controls shared between treated pages would reconcile the two figures [23]. The post reports the design but not the measured effect, and it calls the widely quoted correlation that AI-cited pages are about three times more likely to carry JSON-LD confounded [18][27]. For that to bear on your pages, your pre-period citation levels would have to sit in the range the study matched on [17].
A 48% refutation rate is a fact about one person's plan [21]. The run also spent 108 agents to produce 25 verdicts, about four agent runs per verdict once fetching and extraction are counted [24]. The author wrote: "If you use an LLM to check your own thinking and it keeps agreeing with you, the problem is almost certainly your harness, not your thinking" [9]. That diagnosis is an argument. The run did not put the same claims through a verifier without the refute-by-default instruction, so framing and capability stay tangled in what this run measured [26]. The rebuilt plan uses 13 of the 124 extracted claims, about one in ten [22].
What to watch
- Whether a second pass verifies any of the 99 claims the budget left as leads, and what it does to the 13 survivors.
- Whether anyone runs the same extracted claims through a verifier without the refute-by-default instruction. This run never ran that comparison.
- Whether the JSON-LD difference-in-differences study publishes effect sizes and its matched control list.