Skip to content

Build1 publisher3 min readPublished

The check that never fires: why every agent-built detector needs a negative control

A grep probe written to find blind detectors was itself blind, and returned a clean, plausible number. Silence is both the success state and the total-failure state.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying The check that never fires: why every agent-built detector needs a negative control
Generated illustration

What happened

  • The author asked his coding agent to count how many of his tooling scripts carry a self-test; the agent grepped and answered 12 of 13.
  • The answer was wrong: one script labels its control in uppercase (Ukrainian for 'negative control') while the probe's regex was lowercase with no -i flag. The real answer was 13 of 13.
  • A probe written to find blind detectors was itself blind.
  • The probe misclassified 1 of 13 scripts, about 7.7 percent of the corpus it was auditing.
  • The probe returned a clean, specific, entirely plausible number and nothing in its output hinted that it had missed anything; the author caught it only because the total felt one short and he opened the file by hand.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Writing on dev.to, Volodymyr Kubiria asked his coding agent to count how many of his tooling scripts carried a self-test; it grepped and answered 12 of 13 [1]. The answer was wrong, because one script labelled its control in uppercase while the probe's regex was lowercase with no `-i` flag, and the true count was 13 of 13 [2].

That is a 1-in-13 miscount, roughly 7.7 percent of the corpus, produced by a tool whose entire job was to find blind detectors [1][3]. According to Kubiria, nothing in the output hinted that anything had been missed; the number was clean, specific and plausible, and he caught it only because the total felt one short and he opened the file by hand [4].

The structural problem is cheap supply. Pre-commit hooks, custom lint rules and audit scripts cost one sentence to request, so you request them constantly [5]; Kubiria reports 29 in a single project, running on every commit and almost always printing nothing, which is what you want them to print [6]. And a detector that found nothing and a detector that cannot see produce byte-identical output [7]. The trust curve runs backwards: the longer a check stays quiet the more you rely on it, and a broken detector is silent more reliably than a working one [8].

Test suites have a partial defence. Mutation testing perturbs production code, re-runs the suite and reports any mutant that survived, and it exists because you can reach 100 percent line coverage with tests that assert nothing [9]. But it points at suites, in CI, over application code, and nobody mutation-tests the 200-line script an agent wrote on Tuesday to check something about the docs [10]. The guardrail layer around LLM coding, the pre-call and post-call interceptors, is built to stop the model doing something dangerous, not to prove that a bespoke checker can still see [11]. So the fastest-growing category of quality machinery in the repo is the one with no soundness check at all [12].

Kubiria's fix borrows from lab practice. A negative control is the assay run with everything except the sample: no signal is expected, and a signal means the run is contaminated and its results void; a positive control is a known-present sample that must produce a signal, and if it does not, the instrument is dead and every clean reading it gave you today means nothing [13]. Translated: give every detector a `--self-test` flag with a positive case it must catch and a vacuum case, invented but realistic, it must stay silent on [14]. Run the controls first, and if any control fails the tool prints that it is unsound instead of printing a verdict [15].

His published example, from a tool auditing whether every project rule has an enforcement mechanism, checks that it can still read its inputs at all, that a known rule slug resolves fully, that an absent slug resolves to nothing, and that an invented filename is not accepted as a witness [16]. The output is not a claim about the codebase; it is a claim that this run's verdict is worth reading [17].

Two things to watch. Kubiria notes that controls themselves go blind, though the supplied text is cut off before he explains how he handles it [18]. And in your own repo, the number worth tracking is not how many detectors you have but how many have ever fired: a check that has never once produced output is an untested code path with a green badge on it.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories