Build1 distinct publisher3 min readUpdated
An agent swarm's Arbiter refused nearly everything and finished PII governance at zero. Its four-case adversarial suite handed that behaviour partial credit.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Twenty-six proposals survived the Arbiter, an approval rate of about 23 percent [1], and none of them can have been a PII classification, because PII governance closed the sweep at 0.0 percent [3][2]. The flagship capability was the one that produced nothing, and the output read as care: a report of confident explanations for why each repair was unsafe [3].
The scoring is the instructive part. Back out the arithmetic and the old suite held one approval case against three rejection cases [3], which is why a reviewer with a stuck "no" looked healthy. The author's own diagnosis generalises: when negative cases dominate a benchmark, a component can look accurate by always choosing the conservative label, and the aggregate hides whether the errors are false approvals or false rejections [9].
Now run the degenerate strategies against the replacement suite. Reject everything and you score 4 of 9, about 44 percent; approve everything and you score 5 of 9, about 56 percent [4]. Balance moved the null strategies down, but not to zero, and the approve-everything reviewer now outscores the refuser. What closes the hole is the harness counting false approval, false rejection and error separately instead of compressing opposite failures into one figure [14].
The mechanism behind the live failure was an unwritten distinction. The original prompt named schema and recorded lineage as established evidence and listed what must not be invented: ownership, row counts, refresh cadence, business meaning, downstream consumers [4]. It never drew the line between interpreting that evidence and asserting a new fact about the world [5]. A stronger model covered the gap with common sense; a weaker one obeyed literally and demanded outside corroboration that a column called cust_first_name contained first names [6]. Nothing else changed but the model family, and the ambiguity had been there all along [7].
The taxonomy in the rewrite is sensible enough: clear names can be interpreted, region and account_number on a warehouse table stay ambiguous, and ownership, cadence, row counts and trustworthiness still need evidence [10][c10b]. The part worth copying is that the new instructions price both directions, a fabrication reaching people who rely on the catalog against a sensitive column left ungoverned, and stop describing rejection as the safe default [11].
Which leaves the production gap. The harness now has a floor on acceptance. The sweep did not: 0.0 percent coverage was a line in a report rather than something that stopped the run [3]. A pipeline policed only for unsupported claims is fully satisfied by a reviewer that approves nothing, and from the outside that state is indistinguishable from a clean catalog with no defects left to fix. The threshold that would have caught this is the lower one.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In one live sweep of the ARGUS data-catalog governance swarm, the Arbiter reviewer rejected 86 of 112 proposed metadata repairs.
Rejected repairs included classifications that cust_first_name is PII, billing_zipcode is PII, and shipping_address_line1 is PII.
PII governance finished the sweep at 0.0 percent, with the report full of confident explanations about why each repair was unsafe.
The original Arbiter prompt said schema and recorded lineage were established evidence and warned against inventing ownership, row counts, refresh cadence, business meaning, and downstream consumers.
The original prompt did not explain the difference between interpreting evidence and claiming a new fact about the world.
A strong model filled the gap with common sense, while a weaker model followed the instructions literally and demanded outside corroboration that a column named cust_first_name contained a first name.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific but wholly self-reported and unreproducible
Every figure — 86 of 112, 0.0 percent PII governance, 3 of 4 on the old suite, 9 of 9 on the new one — comes from one first-person dev.to post with no repository, logs, prompt text, dataset or third-party replication. The internal reasoning is checkable (the label splits of the four- and nine-case suites make the degenerate-policy scores arithmetically necessary), which lifts this above pure assertion, but nothing is independently verifiable and the failing model family is never named.
Author's own project only
Observed use is confined to one personal agent swarm: a single production sweep plus a nine-case harness exercised against two hosted models (gpt-4o-mini via GitHub Models, gemma-4-26b via OpenRouter). No other team, product, user count, download, deployment or external uptake of the pattern is disclosed in the supplied material, and no post-fix production sweep is reported.
Broadly aligned, with mild over-generalisation
The framing is unusually disciplined for a promotional-context post: the author reports his own failure, refuses the 'use a stronger model' explanation, and keeps the numbers modest and checkable. The small positive gap comes from generalising a single-project anecdote into 'a common problem in safety-oriented systems' and from letting a 9-of-9 harness pass stand in for evidence that the production capability now works, since no post-fix sweep result is given.
Contest submission and personal-project visibility
The post opens by declaring itself a submission for DEV's Summer Bug Smash: Smash Stories powered by Sentry, so there is a contest and audience-visibility incentive to frame a debugging narrative attractively, plus reputational promotion of the author's ARGUS project. Offsetting this, the disclosed content is self-critical and no product, pricing or commercial relationship is being sold; the supplied material shows no vendor sponsorship of the technical conclusions.
Moderate on the reasoning, low on the measurements
Confidence is split. The methodological core — that a reject-heavy suite gives partial credit to a reject-everything reviewer, and that separating false approvals from false rejections makes the failure direction visible — is transparently derivable from the stated label counts and needs no trust. The empirical claims about the sweep, the model rotation and the post-fix harness depend entirely on one unverified self-report from a contest submission, with a single publisher and no corroboration available in this cluster.
build
A GPU SQL Engine Lost to One CPU Thread Because a Dispatcher Constant Was 128x Too Small1 distinct publisher
build
Sentry's defaults shipped a lifter's shoulder injury while the scrubbing policy passed its tests1 distinct publisher
build
Four indexes, none of them covering: the 78-second page and the one index that fixed it1 distinct publisher
build
One status code, two opposite remedies: the 429 that cost ARGUS seven minutes1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 23, 2026