Megapixel99's assay scan pairs duplicate functions whose outputs match on one fixed input ladder, a check it can apply to roughly a tenth of functions. A match means only that no probe split the pair, so the tool fails the run and hands it to a person.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap−10
- Incentives40
- Confidence40
zerocase fails a CI run when its JUnit, TAP, LCOV, Cobertura or ESLint report shows zero executed items, even if the command exited 0. It counts only what ran, so fifty skipped tests fail the gate and fifty failed tests are left to the exit code.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence45
Mutation testing needs a function to mutate, so the YAML and Terraform that gate production go untested. canfail edits an anchored string in those files, runs your check and reports whether it failed for the reason you declared.
Reality
- Evidence46
- Adoption9
- Hype gap−12
- Incentives28
- Confidence54
Filament Studio's developer stopped trusting his 1,800 passing tests in September and installed the package into a real Laravel app. The defects he logged share one shape: success reported, nothing done.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap−15
- Incentives45
- Confidence58
NFKD maps U+FF0C onto an ordinary comma, so the name 北京,上海 sanitizes down to a single comma and never trips a check for the empty string. The transaction imports with the right amount and a blank payee.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+6
- Incentives25
- Confidence66
An exit status of 0 covers both a suite that passed and a suite that never ran, so ten new verification tools each commit a control test that fails if the tool's own reason for existing has gone away.
Reality
- Evidence55
- Adoption12
- Hype gap+12
- Incentives65
- Confidence50
A conformance suite hashed a repository, killed mutmut, cosmic-ray, Stryker and PIT at the instant a source file went dirty, then hashed again. One of the four came back mutated, and a plain SIGTERM was enough to do it.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−8
- Incentives55
- Confidence55
In tomlkit, the official TOML conformance corpus builds every expected datetime value by calling the parser under test, so parser and expectation drift together. Three hand-written tests caught the break.
Reality
- Evidence62
- Adoption24
- Hype gap+8
- Incentives62
- Confidence56
A dev.to post proposes scoring how much of an agent-written test came from the patch that shipped with it. The literal half of that score depends on ast.parse succeeding, and the invocation it documents feeds the checker a diff.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+12
- Incentives65
- Confidence72
A developer ran three change tasks in one Python and TypeScript product against two code-index tools, and counted a run only after the focused test passed and then failed again with the defect put back on purpose.
Reality
- Evidence46
- Adoption12
- Hype gap−15
- Incentives32
- Confidence41
Salesforce Core failed globally on the second morning of Dreamforce 2026, and the standard Apex loop keeps its compiler, test runner and debugger inside the org. The only account of shipping through it comes from a toolmaker.
Reality
- Evidence44
- Adoption10
- Hype gap+36
- Incentives86
- Confidence34
Prose in AGENTS.md did not stop it, so the project now enumerates agent skills from disk and runs a script that exits non-zero on any copy outside the canonical root, including an empty leftover directory.
Reality
- Evidence42
- Adoption12
- Hype gap+12
- Incentives30
- Confidence46
A dev.to how-to proposes four CI gates for LLM-written tests: execution, coverage, mutation, drift. Read the sample configs line by line and two of them would let through a test that asserts nothing.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+40
- Incentives50
- Confidence55
Counting how many wrong implementations a test suite rejects is only as honest as the list of wrong implementations, and on a small Python order filter mutmut generated five candidates without the one that hid the regression.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap−12
- Incentives22
- Confidence62
A rule can evaluate cleanly, appear in the rule list, report healthy, and still be incapable of going red. One team found four of those in its own stack, and the PromQL one turned on a missing keyword.
Reality
- Evidence72
- Adoption18
- Hype gap−8
- Incentives25
- Confidence62
A dev.to workflow makes every recorded golden prove it can fail before a refactor starts. The gate script that enforces it drops any of its three hardcoded patterns the target file does not contain.
Reality
- Evidence62
- Adoption12
- Hype gap+18
- Incentives72
- Confidence58
A dev.to post reports ten bugs in one measurement harness, every one of them flattering and none found by reading the code. The three checks that eventually caught them cost a few lines each to write.
Reality
- Evidence45
- Adoption12
- Hype gap−10
- Incentives25
- Confidence48
DHSeaDev's write-up on the browser card game Prismwar shows a replay check passing ten of ten matches because the recorder and the playback player both skipped the AI's defence step, so the round trip compared a bug against a copy of itself.
Reality
- Evidence45
- Adoption12
- Hype gap+10
- Incentives55
- Confidence45
Eleven passes and one failure on a footer change turned out to predate the diff. The test built tomorrow in UTC while the function counts days in America/Los_Angeles, so it goes red from midnight UTC to about 07:00.
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap+5
- Incentives25
- Confidence65
A reader pointed out that flipping < to > leaves source size and whole-second mtime unchanged, so Python can import the original bytecode and log the mutant as surviving. The unparseable-garbage canary changes the file size, invalidates the cache and reports success.
Reality
- Evidence62
- Adoption22
- Hype gap−14
- Incentives38
- Confidence58
Earlier coverage
- Five pull requests reported findings filed into an empty issue tracker
Build · September 11, 2026 · 1 publisher
- Nine test cases pin an error taxonomy before an agent rewrites one except clause
Build · September 11, 2026 · 1 publisher
- Nine bugs in a self-built mutation-testing harness, all favoring the builder's story
Build · September 10, 2026 · 1 publisher
- Mutation evidence cleared a wheel guard that FishNet's teardown had already disabled
Build · September 7, 2026 · 1 publisher
- Seed known-bad diffs into the review queue to find the volume where review stops working
Build · September 2, 2026 · 1 publisher
- Mutation scoring drops a 47% coverage test to a 9.5% kill rate
Build · September 1, 2026 · 1 publisher
- Mutation testing pinned the rounding logic, but a separate 3p VAT bug slipped past 202 green tests
Build · August 30, 2026 · 1 publisher
- Mutation-testing an agent-patch gate scores it at 74% recall on injected defects
Build · August 28, 2026 · 1 publisher
- Deleting guard lines one at a time found 40 of 61 unmeasured by any test
Build · August 27, 2026 · 1 publisher
- Green tests proved the converter agreed with itself, not with GnuCash
Build · August 24, 2026 · 1 publisher
- 3,845 tests, 94.22% coverage, and nine things the suite could not see
Build · August 24, 2026 · 1 publisher
- The eval suite was green while the agent lied about a declined charge
Build · August 24, 2026 · 1 publisher
- Coverage at 80% was a price on human attention, and CI is the wrong place to charge it
Build · August 22, 2026 · 1 publisher
- Coverage Is A Line Counter, So A Coverage Gate Buys You Line-Counting Tests
Build · August 21, 2026 · 1 publisher
- An AI test suite hit 94% coverage and missed the one branch that mattered
Build · August 20, 2026 · 1 publisher
- The check that never fires: why every agent-built detector needs a negative control
Build · August 16, 2026 · 1 publisher