Anthropic's Claude Sonnet 5.5 beats Opus 5.5 at coding for half the per-token price, on its own tests and on Artificial Analysis's. At max effort it writes 60% more tokens per task, so moving coding work down a tier saves nearer a fifth than a half.
Perspective Coverage
4 publishers
- Builder
- Builder 39%
- Operator
- Operator 36%
- Investor
- Investor 25%
Reality
- Evidence60
- Adoption30
- Hype gap+25
- Incentives65
- Confidence58
OpenAI's own benchmarks put GPT-6 Sol and Luna at half the list price and a fraction of a rival's cost per finished task. For defenders, the thing getting cheaper is autonomous tool calls into business systems.
Reality
- Evidence32
- Adoption38
- Hype gap+34
- Incentives80
- Confidence52
RoboCurve scored 19 successes in 20 physical attempts on a pair of I2RT YAM arms, at roughly 2.1k output tokens a run. The humanoid pickup described in the same account happened in a simulator.
Reality
- Evidence20
- Adoption8
- Hype gap+55
- Incentives55
- Confidence28
Artificial Analysis scores the new Xiaomi model first among open weights and twelve points behind Claude Opus 5.5, and the cheaper Flash tier is the one an operator should put in front of a real queue.
Reality
- Evidence55
- Adoption30
- Hype gap+15
- Incentives72
- Confidence55
Bespoke Nimble's LoRA fine-tune of Qwen3.5-9B scored 90% against Jev's 93% on an eval it curated itself, and five more replications landed at scales between 421M and 35B parameters. No standard benchmark exists for the category yet.
Reality
- Evidence32
- Adoption46
- Hype gap+42
- Incentives72
- Confidence38
Artificial Analysis now reports how often a provider or model refuses coding-benchmark work on safety grounds. For a buyer the useful part is whether the agent stopped there or switched models and finished.
Reality
- Evidence55
- Adoption25
- Hype gap+10
- Incentives55
- Confidence48
Vals keeps its test materials private and charges the model developers it scores, and Andreessen Horowitz has now put $40 million behind that arrangement. Buyers reading the numbers cannot run the test themselves.
Reality
- Evidence33
- Adoption35
- Hype gap+34
- Incentives82
- Confidence46
In a paper posted on 13 September, nine models across the Claude, GPT and Gemini families edited already-optimal EffiBench solutions in every trial. One added sentence in the prompt recovered 20 refusals.
Reality
- Evidence45
- Adoption18
- Hype gap+15
- Incentives35
- Confidence55
Give Atlas a few images and it generates views a camera never shot, then exports depth, point clouds, and 3D splats you can move through. The invented geometry looks as measurable as the observed, a gap that matters more to a robot than a filmmaker.
Reality
- Evidence36
- Adoption10
- Hype gap+24
- Incentives70
- Confidence44
NVIDIA's submission runs the same Qwen3.6-27B as the llama.cpp reference on the same Jetson board and finishes 6.4x sooner. Most of the gap comes from prompt tokens the runtime never has to prefill.
Reality
- Evidence58
- Adoption18
- Hype gap+20
- Incentives85
- Confidence62
A dev.to post gives every failed coding-agent run one of four labels and scores skill only over the runs where the harness stayed healthy. SSH drops and disk-full errors get published as their own rates.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives15
- Confidence55
Korea's Dokpamo program gives the benchmarking firm's index 25 of the 100 points that decide which teams advance. In the results published on August 27, the team that led that component finished last of the four.
Reality
- Evidence30
- Adoption55
- Hype gap+25
- Incentives55
- Confidence38
A dev.to walkthrough proposes MemoryBench: four tasks, ten metrics and a million-fact store, so buyers can measure recall, latency and write cost themselves before signing.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence35
Stanford's AI Index puts a 142-fold parameter cut and a more than 280-fold price cut behind one fixed MMLU threshold. That narrows where building your own still pays, and it lands in a state-law count that doubled in a year.
Publishers:hai.stanford.edu
Reality
- Evidence60
- Adoption68
- Hype gap+10
- Incentives40
- Confidence57
OpenAI's invite-only Ultrafast tier runs the same GPT-5.6 Sol up to 14 times quicker, while Google halves Gemini Flash pricing until December 31. Latency is now its own budget line.
Perspective Coverage
3 publishers
- Builder
- Builder 35%
- Operator
- Operator 33%
- Investor
- Investor 32%
Reality
- Evidence54
- Adoption42
- Hype gap+27
- Incentives74
- Confidence60