Build1 publisherNot yet confirmed elsewhere3 min readPublished
A reviewer that rejected 86 of 112 repairs still scored 75 percent on its own safety test
An agent swarm's Arbiter refused nearly everything and finished PII governance at zero. Its four-case adversarial suite handed that behaviour partial credit.
The Engineer · Build desk

What happened
- In a live sweep, the Arbiter that guards writes to the ARGUS data catalog rejected 86 of 112 proposed metadata repairs.
- Among the refusals were PII classifications for cust_first_name, billing_zipcode and shipping_address_line1.
- The prompt had passed under a stronger model and broke only when the swarm rotated to another model family.
- The adversarial check was rebuilt as nine balanced cases, five that must be approved and four that must be rejected.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint An aggregate accuracy figure cannot serve as the gate on a reviewer like this, because the same number certifies careful judgement and total silence.
- cost Over-rejection bills the catalog and everyone reading it, and the charge never shows up in the reviewer's own output, which looks conscientious either way.
- decision Anyone swapping model families behind a fixed prompt is choosing where the prompt's unstated judgement calls surface: in a suite, or in production traffic.
- precedent Treating blanket approval of PII proposals as a test failure alongside blanket refusal sets the bar the next reviewer prompt has to clear.
Twenty-six proposals survived the Arbiter, an approval rate of about 23 percent [16], and none of them can have been a PII classification, because PII governance closed the sweep at 0.0 percent [3][17]. The flagship capability was the one that produced nothing, and the output read as care: a report of confident explanations for why each repair was unsafe [3].
The scoring is the instructive part. Back out the arithmetic and the old suite held one approval case against three rejection cases [18], which is why a reviewer with a stuck "no" looked healthy. The author's own diagnosis generalises: when negative cases dominate a benchmark, a component can look accurate by always choosing the conservative label, and the aggregate hides whether the errors are false approvals or false rejections [9].
Now run the degenerate strategies against the replacement suite. Reject everything and you score 4 of 9, about 44 percent; approve everything and you score 5 of 9, about 56 percent [19]. Balance moved the null strategies down, but not to zero, and the approve-everything reviewer now outscores the refuser. What closes the hole is the harness counting false approval, false rejection and error separately instead of compressing opposite failures into one figure [15].
The mechanism behind the live failure was an unwritten distinction. The original prompt named schema and recorded lineage as established evidence and listed what must not be invented: ownership, row counts, refresh cadence, business meaning, downstream consumers [4]. It never drew the line between interpreting that evidence and asserting a new fact about the world [5]. A stronger model covered the gap with common sense; a weaker one obeyed literally and demanded outside corroboration that a column called cust_first_name contained first names [6]. Nothing else changed but the model family, and the ambiguity had been there all along [7].
The taxonomy in the rewrite is sensible enough: clear names can be interpreted, region and account_number on a warehouse table stay ambiguous, and ownership, cadence, row counts and trustworthiness still need evidence [10][c10b]. The part worth copying is that the new instructions price both directions, a fabrication reaching people who rely on the catalog against a sensitive column left ungoverned, and stop describing rejection as the safe default [12].
Which leaves the production gap. The harness now has a floor on acceptance. The sweep did not: 0.0 percent coverage was a line in a report rather than something that stopped the run [3]. A pipeline policed only for unsupported claims is fully satisfied by a reviewer that approves nothing, and from the outside that state is indistinguishable from a clean catalog with no defects left to fix. The threshold that would have caught this is the lower one.
What to watch
- Whether the nine-case suite holds on the weaker model family, and what its false-rejection count looks like there.
- Whether a minimum approval rate is enforced on live sweeps as well as in the test harness.
- How the new evidence categories behave on genuinely ambiguous names, such as account_number on a table that does describe people.