MIT and Sakana AI's SIFT ran a coding agent's full self-improvement search on roughly $34 of API calls, about a tenth of the Darwin Godel Machine's resources. The saving comes from a language-model judge screening patches, so it holds only while that judge picks correctly.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
build3 publishersConfirmed Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
Antigravity's free tier includes Gemini 3.1 Pro and Claude Opus 4.6 with unlimited tab completions. XDA reports Opus burns credits about four times faster than Gemini, and Google moved the free tier onto credits in March 2026.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence40
build1 publisherOne report The merged 27B finished 37 of 100 internal coding tasks against 34 for the fast checkpoint and 39 for the reasoning one, and it did that while emitting 71% fewer output tokens than the 39, with no post-training in the recipe.
Reality
- Evidence55
- Adoption30
- Hype gap+18
- Incentives78
- Confidence58
build1 publisherOne report Alibaba ran the reviewer internally for two years across tens of thousands of developers before releasing it. The write-up is specific about how files get selected and comments get positioned, and it stops short of publishing the benchmark scores.
Reality
- Evidence32
- Adoption30
- Hype gap+55
- Incentives68
- Confidence38
build1 publisherOne report The Kotlin Benchmark grades agents on 105 verified repository tasks. Its token column shows setups that solve within a few tasks of each other burning between 66,000 and 777,000 tokens per fix.
Reality
- Evidence62
- Adoption30
- Hype gap+22
- Incentives60
- Confidence58
build1 publisherOne report SWE-Bench ProMax puts frontier coding agents on 170 curated refactoring commits. The number that should move procurement is a different one: nearly 60% of unsolved SWE-bench Verified tasks have flawed tests.
Reality
- Evidence38
- Adoption14
- Hype gap+12
- Incentives62
- Confidence42
A June 13 Commerce Department order cut non-US users off from two Anthropic models overnight, and OpenAI followed with limits of its own. Model access is now a sovereign risk, not a vendor term.
Reality
- Evidence22
- Adoption20
- Hype gap+45
- Incentives68
- Confidence30