KAIST and Seoul National University's AgSpec nearly doubles accepted draft length for coding agents by indexing files in the diff and JSON forms agents emit. It lives entirely in the retrieval index and leaves model weights untouched.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Google Cloud AI Research has released RRSI, an Apache 2.0 tool whose self-rewriting agent harness lifted Terminal-Bench 2.1 scores from 74.2% to 80.2%. Any team can use it commercially, though on tasks the agent never trained against the reported gain falls to 4.7 points.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence40
Fireworks' Ember-1 used 23% fewer reasoning tokens than Kimi K3 in The New Stack's tests, yet Kimi on the cheapest host would cost $1.96 to Ember's $2.48. Ember beats Fireworks' own Kimi rate and loses at the cheapest, so buyers have to price the host before the model.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 30%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption30
- Hype gap+25
- Incentives70
- Confidence58
GPT-5.6 Luna says "frog" 70-95% of the time when asked for an amphibian after capability benchmarks, against 12-38% after real use, a LessWrong post reports. Anyone with black-box access can run the check, though its authors cannot yet say whether it detects evaluation awareness or lexical cues.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
Broken answer keys and graders that punish correct tool calls drove the verdicts. Epoch AI says it stops each review once it has enough evidence, so the published defect counts are floors.
Reality
- Evidence65
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence62
A study of eight frontier models on SWE-bench Verified puts agentic coding at roughly 1,000 times the token cost of code chat, dominated by input, with the models' own pre-run estimates correlating no better than 0.39.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence55
A dev.to post credits Anthropic with running the same model and the same prompt under two harnesses, 20 minutes and $9 for a broken result against six hours and $200 for a working one. The hourly spend barely moved.
Reality
- Evidence25
- Adoption
- Insufficient
- Hype gap+30
- Incentives45
- Confidence40
GLM-5.3-Flash leads agentic terminal work, DeepSeek V4 Flash is billed as the cheapest per token, and a 2.52B MiniCPM5-2B runs locally under Apache 2.0. The comparison flags most of those numbers as vendor-reported.
Reality
- Evidence34
- Adoption27
- Hype gap+26
- Incentives58
- Confidence41
NVIDIA's developer blog sets out the five-level rollup from step to benchmark and the two scores that read the same trace. Step-level says where the chain broke. End-to-end reads the environment and says whether the refund posted.
Reality
- Evidence55
- Adoption20
- Hype gap+12
- Incentives50
- Confidence60
OpenRouter has dropped the stealth listing. The model is Pareto 26.9, from unbiased.ai, at $2.50 per million input tokens before a 10 October launch. That is double what 26.8 charged.
Reality
- Evidence64
- Adoption42
- Hype gap+45
- Incentives68
- Confidence55
Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.
Reality
- Evidence40
- Adoption20
- Hype gap+15
- Incentives55
- Confidence45
Researchers held a coding harness's execution loop fixed and varied planning, tools and context management across four models, and found that what each component is worth depends on how strong the model already is.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence58
A new evaluation gives models explicit rules about what may appear in their reasoning and scores whether they comply while still solving the problem. Compliance rose with model size and fell with more RL training.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−5
- Incentives35
- Confidence60
A position paper argues that CLT-based intervals dramatically understate uncertainty below a few hundred datapoints. The specialized benchmarks frontier teams build are already smaller than that before anyone slices them by task.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence58
Microsoft reports Qwen3.5-9B going from 41.8% to 56.4% on SWE-bench Verified using 6K examples. The load-bearing change is who owns the agent loop, and what the service boundary costs.
Reality
- Evidence32
- Adoption18
- Hype gap+28
- Incentives66
- Confidence44
Token traffic, survey reach and production-model ledgers rank different vendors because they count different things. The autonomy figures say the hard part is still unbought.
Reality
- Evidence58
- Adoption64
- Hype gap+32
- Incentives68
- Confidence52
A killed side project produced the useful result: 11 of 31 pilot pairs, $5.60, and zero Java fixes from the cheap model. A savings figure without a pass rate is not a number.
Reality
- Evidence46
- Adoption14
- Hype gap+42
- Incentives24
- Confidence55
One model was verified at 30.16 in the official harness and reported at 100.00 in NVIDIA's. Microsoft's Agent Lightning now trains the harness into the weights.
Reality
- Evidence54
- Adoption58
- Hype gap+28
- Incentives66
- Confidence48
SemiAnalysis says a $200 plan can yield $8,000 of Anthropic API-equivalent usage and $14,000 of OpenAI's. Even the best published caching fix leaves the gap several times wide.
Reality
- Evidence32
- Adoption44
- Hype gap+34
- Incentives58
- Confidence38
SWE-Bench ProMax puts frontier coding agents on 170 curated refactoring commits. The number that should move procurement is a different one: nearly 60% of unsolved SWE-bench Verified tasks have flawed tests.
Reality
- Evidence38
- Adoption14
- Hype gap+12
- Incentives62
- Confidence42
Earlier coverage
- Ornith-1.5 moves the RL loop upstream, and the hard job becomes reward design
Build · August 19, 2026 · 2 publishers
- Ornith-1.0's benchmarks are fine. Ollama can't parse its tool calls.
Build · August 18, 2026 · 1 publisher
- The AI bill nobody reconciles: cost per finished task, not per million tokens
Leadership · August 18, 2026 · 1 publisher
- Kraken's parent now runs a security model that Washington can switch off
Invest · August 17, 2026 · 1 publisher
- NIST says AI benchmarks are now an attack surface, not just a measuring stick
Science · August 16, 2026 · 1 publisher
- NIST's own logs show agents looking up the answers, making public benchmark scores soft evidence
Science · August 16, 2026 · 1 publisher