Skip to content

Build1 publisher3 min readPublished

Do not let the model that wrote the diff approve it: the case for a cross-vendor review gate

A dev.to tip argues agent pipelines need an adversarial verifier from a different model family plus randomized human audits, because same-model review exploits a documented self-preference bias.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A dev.to post titled "AI Coding Tip 032 - Build a Dark Factory Pipeline" recommends running an AI coding pipeline like a dark factory: automated, sampled, and policed by an adversarial model.
  • The described anti-pattern is letting one model write a pull request and then handing the same model, or a suspiciously agreeable fresh instance of it, the job of approving its own work.
  • Models like their own output when asked to judge it, a documented self-preference bias that makes same-model review about as objective as a suspect grading their own trial.
  • 100 percent human review does not scale once agents merge dozens of pull requests a day.
  • Fake full coverage looks safe on a dashboard the same way vanity coverage does: fine on a slide deck and useless in production.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A tip post published on dev.to, "AI Coding Tip 032 - Build a Dark Factory Pipeline", argues that an agent-heavy coding pipeline should be automated, sampled, and policed by an adversarial model rather than trusted or fully hand-inspected [1]. It matters for a narrower reason than the factory metaphor suggests: the default shortcut, letting one model open a pull request and then handing the same model or a fresh instance of it the approval, runs straight into a documented self-preference bias in which models favour their own output when asked to judge it [2][3].

Both obvious alternatives fail in a way you can predict. Full human review stops scaling the moment your agents merge dozens of pull requests a day [4]. Self-approval produces the other failure, coverage that looks complete on a dashboard and is worthless in production, the same trick vanity test coverage plays [5].

The arrangement the post proposes is specific enough to argue with. One model writes the spec, a builder model writes the code from it [6]. A separate model from a different vendor or family acts as the verifier, with the single job of finding fault in the builder's output [7]. Anything the verifier flags is rejected and returned to the builder to fix before a human sees it [8]. Merged pull requests are then sampled at a defined rate, the way quality control samples a manufacturing lot [9], with the sampled subset chosen at random rather than picking the easiest or newest diffs [10]. Every human override of the verifier is logged and fed back as a correction signal [11], and any merge that skips both the adversarial gate and the sampling gate is blocked, with no manual bypass [12].

The consequence for staffing is arithmetic: daily human reviews equal the sampling rate times merged pull requests, so the rate is the only lever you actually set, and the post never names a number [1]. The consequence for procurement is that a cross-vendor verifier makes a second provider a hard dependency of your merge path, with its own quotas and outages [2].

Two things in the material are weaker than the recommendation built on them. The claim that sampling catches the same class of defects a full review would, at a fraction of the human hours, is asserted by analogy to acceptance sampling, where a small randomized batch decides whether the whole lot passes [13][14]; no measured catch rate for sampled versus full code review appears in the source [3]. And the passage citing studies that quantify self-preference bias breaks off before any figure [c14b]. The qualitative version stands on its own: research on adversarial code review holds that the maker should not grade the checker, because the assumptions that produced the defect also produce the blind spot in the review [15].

The factory material is a useful corrective to the branding. A lights-out plant runs a line with no on-site workers because human presence adds variability [16]; Foxconn replaced over 60,000 workers at its Kunshan plant with robots running around the clock [17]; and even the most automated line still keeps someone walking the floor [18]. The data centre detail is the same shape: hypoxic fire suppression holds oxygen near 14 to 15 percent, low enough that nothing burns and still breathable for a technician [19], a tuned guardrail rather than an empty room.

Watch the override log [11]. It is the only instrument in this design that tells you whether the verifier and the sample agree with your humans, which is the one thing self-approval can never report [4].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories