Across twelve ORBITAL tasks, eight failed a browser check at least once, and the verifier's own triage marked six of them automation bugs. Those six were cleared by override, two of them with per-case approval.
Reality
- Evidence45
- Adoption12
- Hype gap−10
- Incentives52
- Confidence45
A post on dev.to publishes accept_run.py, a gate that starts the test command itself and writes the command, directory, exit code and output hashes to JSON, then refuses any repo whose git tree is dirty.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+24
- Incentives58
- Confidence58
Migrating a legacy suite with a coding agent turns on proof, so this harness runs the migrated Playwright test itself, checks the required assertions, hashes the report, and only then lets the work commit.
Reality
- Evidence32
- Adoption10
- Hype gap+25
- Incentives35
- Confidence45
Wildcard DNS answers names nobody configured, which makes a resolving hostname weak evidence. So the control experiment moved into the Actor's output contract: three random probes per run, four evidence classes, 98 of 129 rows flagged.
Reality
- Evidence58
- Adoption20
- Hype gap−8
- Incentives70
- Confidence57
Rubber Duck runs a complementary model over what the primary coding agent produced, and a new Agent Host keeps the session alive outside the window that started it. Both changes concede an agent cannot check itself.
Reality
- Evidence45
- Adoption25
- Hype gap+15
- Incentives65
- Confidence55
An eldercare agent published at a live URL puts its two hard rules in Python instead of prompts, and the demo tests them by calling the send function directly, which is the only version of that claim a reader can check.
Reality
- Evidence32
- Adoption8
- Hype gap+14
- Incentives68
- Confidence46
Across roughly 130 agents and more than 72 sessions, one report passed seven integrity gates that never executed. The operator's conclusion is that completion has to be read off disk by something the agent cannot author.
Reality
- Evidence30
- Adoption14
- Hype gap+12
- Incentives32
- Confidence44
A new tool diffs a Git checkpoint instead of trusting a coding agent's summary, with no LLM in the analysis path. The baseline is the clever part; the risk weights are the part still unpublished.
Reality
- Evidence38
- Adoption9
- Hype gap+14
- Incentives62
- Confidence42
OpenWorkProof v0.5 refuses to form a high-risk decision unless two independently keyed verifiers each execute the work and agree field by field. Its own adoption evidence is still empty.
Reality
- Evidence34
- Adoption3
- Hype gap+12
- Incentives72
- Confidence38
An agent reported a bulk insert complete when zero rows had landed. The fix is not a better guardrail on the text but a mandatory re-read of the system of record before completion can be claimed.
Reality
- Evidence34
- Adoption11
- Hype gap+24
- Incentives62
- Confidence33