GLM-5.3-Flash leads agentic terminal work, DeepSeek V4 Flash is billed as the cheapest per token, and a 2.52B MiniCPM5-2B runs locally under Apache 2.0. The comparison flags most of those numbers as vendor-reported.
Reality
- Evidence34
- Adoption27
- Hype gap+26
- Incentives58
- Confidence41
The merged 27B finished 37 of 100 internal coding tasks against 34 for the fast checkpoint and 39 for the reasoning one, and it did that while emitting 71% fewer output tokens than the 39, with no post-training in the recipe.
Reality
- Evidence55
- Adoption30
- Hype gap+18
- Incentives78
- Confidence58
PrismML's ternary build of Qwen3.8 27B keeps 98.2 percent of the full-precision benchmark average on both of the company's inconsistent scorecards, and the loss it does take is concentrated in knowledge and reasoning.
Reality
- Evidence38
- Adoption20
- Hype gap+32
- Incentives72
- Confidence55
OmnisBench's author rebuilt his split from date-stamped LiveCodeBench problems. The cheap tier fell from 90% to 60%, and a 4,096-token cap had been marking the frontier model absent.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives78
- Confidence48
Alibaba's Apache-2.0 Qwen3.8-27B fits in about 17GB and matched near-frontier scores, per Artificial Analysis. It also burned 3.7x the median output tokens getting there.
Reality
- Evidence62
- Adoption64
- Hype gap+18
- Incentives60
- Confidence55