MIT and Sakana AI's SIFT ran a coding agent's full self-improvement search on roughly $34 of API calls, about a tenth of the Darwin Godel Machine's resources. The saving comes from a language-model judge screening patches, so it holds only while that judge picks correctly.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
Removing the qwen3-coder:30b planner from the olo app stack lost no working build across 88 requests, its developer reports. The planner is the stack's only neural part, and on this corpus it showed no measurable gain, in figures from a private repository.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap−10
- Incentives20
- Confidence40
The coding-trained model came out ahead over thirty runs, but four of the five fixtures tied, which leaves the whole margin in a single Docker log task run at default temperature between a 30b model and an 8b one.
Reality
- Evidence58
- Adoption10
- Hype gap−15
- Incentives25
- Confidence55
A dev.to post hands four models a file whose comment contradicts its code, then asks a clean session to fix the inconsistency. It is a well-built way to expose the failure, and no counts are published yet.
Reality
- Evidence38
- Adoption12
- Hype gap+10
- Incentives22
- Confidence45
After weeks of reading 100 percent on its own curriculum, one team rebuilt its coding-agent bench around unseen prompts and guard redirects per run, and found two harness contaminations on the way.
Reality
- Evidence36
- Adoption11
- Hype gap−9
- Incentives58
- Confidence42
A June finding said local models cannot iterate on code. It said so about chat UIs, and it named the fix. July supplied that fix and measured it. The follow-up test result is the number worth reading.
Reality
- Evidence46
- Adoption12
- Hype gap+12
- Incentives32
- Confidence44
A dev.to post scores ten coding models on five real tasks and divides by price. The method is cheap to copy; the vendor plumbing it recommends deserves more scrutiny than the arithmetic.
Reality
- Evidence20
- Adoption12
- Hype gap+45
- Incentives70
- Confidence55