NVIDIA released SoL-Pi, an MIT-licensed Pi extension with four token-saving harness mechanisms that an optimizer agent found. A dev.to review puts the saving near one third on long sessions, bought with a few lost solves.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence45
The same model scored twice under two scaffolds. A dev.to post uses that gap to argue the dividing line in AI coding is whether the model can run your repo's own commands and read the failure.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+18
- Incentives45
- Confidence50
Huntress ran seven Rails tasks past Claude Fable 5.1 and watched it copy hand-rolled code out of its own repository. Changing the harness, in three cheap steps, took API recall from 48 percent to 100.
Reality
- Evidence58
- Adoption25
- Hype gap+15
- Incentives40
- Confidence55
Researchers held a coding harness's execution loop fixed and varied planning, tools and context management across four models, and found that what each component is worth depends on how strong the model already is.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence58
A developer kept the model and the repository fixed and changed only the agent framework. Six days of Codex produced no completed run, while goose's own engine drove the same model to working code.
Reality
- Evidence30
- Adoption12
- Hype gap+30
- Incentives55
- Confidence40
Coding-agent failures cluster in the plumbing between knowing what to change and changing it. One harness builder measured what that costs and published the failure rates for Grok 4 and GLM-4.7.
Publishers:substack.com
Reality
- Evidence36
- Adoption22
- Hype gap+28
- Incentives78
- Confidence42
A Labs engineer says prompt engineering and harness tuning both topped out on multi-hour builds, and the lever that moved them was a separate grading agent working from written criteria. The costs are named; the gain is not measured.
Reality
- Evidence42
- Adoption20
- Hype gap+18
- Incentives72
- Confidence55
Anthropic's engineering blog says the context-reset logic it wrote to stop Sonnet 4.5 quitting early was unnecessary on Opus 4.5. Its answer is to sell hosted interfaces, not harness code.
Reality
- Evidence42
- Adoption30
- Hype gap+12
- Incentives78
- Confidence55
Anthropic's harness for agents that work across days is mostly ordinary repository hygiene, which is also the best reason to think it will still be worth something after the next model release.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence65
A June finding said local models cannot iterate on code. It said so about chat UIs, and it named the fix. July supplied that fix and measured it. The follow-up test result is the number worth reading.
Reality
- Evidence46
- Adoption12
- Hype gap+12
- Incentives32
- Confidence44
Aikido spent 11.7 billion tokens rediscovering 32 fresh CVEs with ten models, three attempts each. The number that should move a scanning budget is the marginal cost of the second and third pass.
Reality
- Evidence44
- Adoption18
- Hype gap+32
- Incentives76
- Confidence52
A developer grepped pi for the five features it refuses and got two hits, both false. The audit that survives is one line of default tool names, not the size of the repo.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+12
- Incentives50
- Confidence42
A practitioner argues the 80% coverage norm was a budget for review time, not a standard, and that any check living only in the pipeline reports to you rather than constraining the agent.
Reality
- Evidence20
- Adoption14
- Hype gap+28
- Incentives
- Insufficient
- Confidence34
A five-person Nvidia team published the architecture behind its AVO agent alongside a perfect public-set score. The model did not change. The system around it did.
Perspective Coverage
4 publishers
- Builder
- Builder 50%
- Operator
- Operator 29%
- Investor
- Investor 21%
Reality
- Evidence58
- Adoption20
- Hype gap+34
- Incentives82
- Confidence72
A dev.to series argues "prompt engineering" is too narrow for agents: the package sent to the model is reassembled on every call, which makes it a runtime concern, not a file.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives65
- Confidence33