Enterprises buy the cheapest model that clears their bar. On Ramp's July billing data, that leaves Anthropic's flagship with about an eighth of its maker's platform spend.
Reality
- Evidence55
- Adoption25
- Hype gap+30
- Incentives40
- Confidence55
Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.
Reality
- Evidence40
- Adoption20
- Hype gap+15
- Incentives55
- Confidence45
Alibaba's Qwen3.8-Flash-Next preview activates 6B of its 125B parameters per token. Per-token compute drops to under a quarter of the dense 27B's, and about 125GB of weights still has to stay on device.
Reality
- Evidence34
- Adoption18
- Hype gap+15
- Incentives60
- Confidence42
Oriol Vinyals has left Google DeepMind to build Discovery Loop with Jeff Dean, Sanjay Ghemawat and Quoc Le, on the argument that AI already writes the code and runs the experiments while idea generation and judging results lag.
Reality
- Evidence45
- Adoption15
- Hype gap+20
- Incentives75
- Confidence45
Alibaba's largest open-weight release fits on a single eight-GPU node only because a community four-bit build takes it down to 1.2 TB. AWS publishes the vLLM config for it and skips the price.
Reality
- Evidence52
- Adoption22
- Hype gap+32
- Incentives78
- Confidence48
Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.
Perspective Coverage
5 publishers
- Builder
- Builder 52%
- Operator
- Operator 28%
- Investor
- Investor 20%
Reality
- Evidence58
- Adoption52
- Hype gap+32
- Incentives76
- Confidence71
An arXiv tracing study of Claude Code agents on Gemma and Qwen measured prefix-cache hit rates between 84.6 and 99.5 percent, which moves the serving bottleneck to how long you can keep KV blocks resident between tool calls.
Reality
- Evidence58
- Adoption25
- Hype gap+12
- Incentives
- Insufficient
- Confidence55
Token traffic, survey reach and production-model ledgers rank different vendors because they count different things. The autonomy figures say the hard part is still unbought.
Reality
- Evidence58
- Adoption64
- Hype gap+32
- Incentives68
- Confidence52
Alibaba's Apache-2.0 Qwen3.8-27B fits in about 17GB and matched near-frontier scores, per Artificial Analysis. It also burned 3.7x the median output tokens getting there.
Reality
- Evidence62
- Adoption64
- Hype gap+18
- Incentives60
- Confidence55
MAI-Thinking-1 is in public preview in Microsoft Foundry at $2 per million input tokens. For C# shops the consequence is an interface swap, not a new runtime to operate.
Reality
- Evidence34
- Adoption20
- Hype gap+32
- Incentives46
- Confidence41
Qwen 3.8 27B ran on a MacBook Pro from a 17GB GGUF and spent 21 minutes on one SVG. Licensing and access stopped being the blocker; latency and KV cache budgeting became the job.
Reality
- Evidence48
- Adoption58
- Hype gap+18
- Incentives62
- Confidence52