Open-source tool runtape traced an agent's unrequested invoice forward to one sentence in a tool result, 10 of 10 reruns with it against 0 of 10 without. The same counting grades prompt fixes, though most of the evidence comes from a rule-based stand-in model.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence50
PostgreSQL scans a 20-row test table sequentially even when a perfect index exists, so a test that fails on Seq Scan cannot find a missing index. Re-running EXPLAIN with enable_seqscan off catches it at any row count, since only a filter no index can serve keeps its scan.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence58
GPT-5.5, Gemini 3.7 Flash and Grok 4.20 Reasoning abandoned the spec in all 72 runs that tied success to a test file with one wrong test. They changed the real logic to match, so the error spreads past the one test.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives25
- Confidence40
zerocase fails a CI run when its JUnit, TAP, LCOV, Cobertura or ESLint report shows zero executed items, even if the command exited 0. It counts only what ran, so fifty skipped tests fail the gate and fifty failed tests are left to the exit code.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence45
Three of the four layers in a dev.to testing pyramid for MCP servers are plain pytest checks on schemas, error envelopes and session expiry. Only the fourth puts a model in the loop, and the available text breaks off before describing it.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence50
A dev.to scorecard for agent-authored patches digests every file under testdata/ and replays recorded property seeds before pytest runs. The published fragments enforce less than the prose around them describes.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+30
- Incentives20
- Confidence70
A dev.to design essay argues CI was sized for human review speed and proposes tiering checks by cost and failure probability. Its figures are examples, so the case rests on the failure modes it names.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+30
- Incentives25
- Confidence45
A dev.to proposal freezes fixture bytes, runner config and two test floors at the parent SHA, then denies the coding agent write access to the file holding them. Its content hashes hold up better than the path denylist it prints.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives20
- Confidence55
A dev.to write-up sorts every test path into four classes and hashes the scoring ones on main, so a patch that retouches a fixture fails with a different exit code than a patch that breaks a property.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+18
- Incentives18
- Confidence45
A dev.to engineer moves the admit/deny verdict back into an in-process token bucket and wraps it so the five verdict words cannot reach a model call, arguing that a free inference tier still charges the defender under a flood.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives15
- Confidence60
A dev.to FAQ proposes failing any agent session whose diff touches a test path without the ticket ID in a signed allowlist. Both scripts in the post are labeled unexecuted, so the false-positive count is yours to find.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+14
- Incentives20
- Confidence50
An exit status of 0 covers both a suite that passed and a suite that never ran, so ten new verification tools each commit a control test that fails if the tool's own reason for existing has gone away.
Reality
- Evidence55
- Adoption12
- Hype gap+12
- Incentives65
- Confidence50
Python's logging module keeps an RLock behind every handler. One shared setup_logging() therefore deadlocks a forked pool on one machine and silently drops worker records on another, and a single-OS review only ever sees one of them.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives20
- Confidence55
A dev.to worked example keeps HTTP clients and model SDKs out of the lease renewal path. Its Postgres fence rises on every renewal, so a write that outlives one five-second interval is rolled back by its own lease.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives25
- Confidence60
A dev.to harness stamps connect, TTFB, body, patch and test time on every generate call. In live mode the first two clocks come from one expression, so curl still names the handshake, and the published numbers are synthetic.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+55
- Incentives70
- Confidence68
A dev.to case study builds the timezone calculator and its pytest suite first, then lets a model write the mail worker that may only call it, with 8 March 2026 pinned as the fixture no regeneration may delete.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+12
- Incentives20
- Confidence60
A dev.to proposal freezes the budget in clock.env and hashes src and tests before any prompt is sent. Its sample predicate then runs a test file from outside the spike's own allowed-paths list.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+15
- Incentives20
- Confidence55
A queue worker built its remaining-seconds budget from time.time(), and after the host woke from sleep it logged remaining=-1842.7. Two days of debugging went to sockets, Docker and NTP before the clocks got compared.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap0
- Incentives60
- Confidence55
A post on dev.to publishes accept_run.py, a gate that starts the test command itself and writes the command, directory, exit code and output hashes to JSON, then refuses any repo whose git tree is dirty.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+24
- Incentives58
- Confidence58
A dev.to write-up swaps the single assertion for per-field hit rates with 95 percent intervals over repeated trials. On its own example table, the pair it calls a real regression has intervals that overlap by 0.3 points.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives75
- Confidence60
Earlier coverage
- Running this contamination gate on git diffs zeroes out its literal check
Build · September 17, 2026 · 1 publisher
- This merge gate runs decode(encode(x)) before it trusts a green pytest run
Build · September 16, 2026 · 1 publisher
- A workflow_run listener reads the shutdown-signal line before it retries a dead CI job
Build · September 16, 2026 · 1 publisher
- Censoring infra failures shrinks the denominator behind a coding-agent pass rate
Build · September 15, 2026 · 1 publisher
- PYTHONUTF8=1 in a .zshrc hid the encoding bug from every local test run
Build · September 15, 2026 · 1 publisher
- Committed byte fixtures catch the ensure_ascii flip a dict-equality test lets through
Build · September 15, 2026 · 1 publisher
- Classifying test files by glob hands the merge bit to whoever owns tests/oracles
Build · September 15, 2026 · 1 publisher
- A must-reject lockfile blocks the merge when a frozen bad input starts returning 200
Build · September 14, 2026 · 1 publisher
- The mutation gate skips every pattern it cannot find in the target file
Build · September 14, 2026 · 1 publisher
- No source edits until a flake repeats twice, tested across a 16-case environment matrix
Build · September 14, 2026 · 1 publisher
- A generic loop-cost harness for any OpenAI-compatible endpoint asks the model once per repeat
Build · September 14, 2026 · 1 publisher
- Trusting messages[0] lets yesterday's email verify today's signup
Build · September 13, 2026 · 1 publisher
- Nine test cases pin an error taxonomy before an agent rewrites one except clause
Build · September 11, 2026 · 1 publisher
- Proposed merge gate fails an agent patch when a committed seed stops drawing
Build · September 10, 2026 · 1 publisher
- An except Exception clause lets the recorder script record its own AssertionError as the timeout exception
Build · September 10, 2026 · 1 publisher
- Nine bugs in a self-built mutation-testing harness, all favoring the builder's story
Build · September 10, 2026 · 1 publisher
- A fixture-digest lockfile moves the merge gate out of CI and into commit identity
Build · September 10, 2026 · 1 publisher
- The archrule glob makes your package tree the architecture spec
Build · September 10, 2026 · 1 publisher
- A trailing dollar sign disarms the fixture check in this merge-promotion hook
Build · September 9, 2026 · 1 publisher
- A bare ToolNode.invoke() has needed a Runtime object for eleven months
Build · September 8, 2026 · 1 publisher
- An assertion budget scores the test AST before CI installs anything
Build · September 8, 2026 · 1 publisher
- A decorative checkmark crashes pytest on a runner that boots with LANG=C
Build · September 8, 2026 · 1 publisher
- Freezing an expected-result file moves the pass condition outside the agent's writable surface
Build · September 7, 2026 · 1 publisher
- SWE-Gate flunks 221 of 644 test-passing agent patches on rules mined from PR comments
Build · September 5, 2026 · 1 publisher
- Pytest catches the pipeline bugs that ship as wrong numbers
Build · September 2, 2026 · 1 publisher
- Mutation scoring drops a 47% coverage test to a 9.5% kill rate
Build · September 1, 2026 · 1 publisher
- Classify the CI failure before you rewrite the agent's patch
Build · August 29, 2026 · 1 publisher
- Django's boring core is paying for the churn in everything above it
Build · August 28, 2026 · 1 publisher
- Mutation-testing an agent-patch gate scores it at 74% recall on injected defects
Build · August 28, 2026 · 1 publisher
- A model documenting a retry wrapper hands you tenacity's parameters
Build · August 27, 2026 · 1 publisher
- Five rewrites later, the LLM is out of the test loop and into the selectors
Build · August 25, 2026 · 1 publisher
- A green test suite that proved nothing: aiortc's mangled cert slot versus libp2p certhash pinning
Build · August 20, 2026 · 1 publisher
- The 21-cent model bake-off that inverted when the judge got audited
Build · August 20, 2026 · 1 publisher
- An AI test suite hit 94% coverage and missed the one branch that mattered
Build · August 20, 2026 · 1 publisher
- Your Databricks Pipeline Is A Demo Until Promotion Only Runs One Way
Build · August 20, 2026 · 1 publisher
- 255 tool schemas, 91K tokens: pricing the two MCP costs nobody budgets
Build · August 19, 2026 · 1 publisher