Ollama since 0.34.4 lets Gemma 4 skip the requested JSON schema, returning bare text with HTTP 200 in 8 of 24 test calls. Until the open fix ships in a release, structured output on local thinking models needs a shape check in the client.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives20
- Confidence58
Grok 4.20 followed an injected wrong step about five times as often with reasoning mode on, in a 15-question Kaggle benchmark entry. On the Anthropic test it uses, that makes the trace a better record of how the model answered and a weaker guard against a bad step.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives20
- Confidence35
IBM's 3B, 8B and 30B dense models all get a thinking switch and native tool calling, but only the two larger ones get agentic RL, and the tuning mixture leans hard on software engineering.
Perspective Coverage
3 publishers
- Builder
- Builder 63%
- Operator
- Operator 28%
- Investor
- Investor 9%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence72
Ben Thompson now says he does not think AI is a bubble, and the evidence he offers is a capability jump that showed up in Claude Code in December, weeks after Anthropic shipped the Opus 4.5 weights.
Reality
- Evidence38
- Adoption35
- Hype gap+32
- Incentives45
- Confidence40
DigitalOcean says step-by-step reasoning is billed as output and invisible by design. The 90% figure it cites comes from a paper that estimates hidden token counts, so moving it onto your own invoice takes a matching task mix.
Reality
- Evidence32
- Adoption26
- Hype gap+38
- Incentives88
- Confidence42
In IEEE Spectrum's account of the 2026 compute market, reasoning and agentic workloads have pushed serving toward memory-heavy silicon, and Amazon now runs a single inference job across two vendors' chips.
Reality
- Evidence47
- Adoption55
- Hype gap+24
- Incentives66
- Confidence46
A new arXiv benchmark runs seven extended-thinking models under an explicit order to conceal and under an offhand contextual detail. A routine anti-bias system prompt pushes implicit detection as low as 5%.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence55
A new evaluation gives models explicit rules about what may appear in their reasoning and scores whether they comply while still solving the problem. Compliance rose with model size and fell with more RL training.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−5
- Incentives35
- Confidence60
Eleven tasks with pre-computed answer keys, three runs each, seven effort settings. Everything from low upward scored 33 of 33, so the only thing the top rung buys is the number on the launch page.
Reality
- Evidence62
- Adoption45
- Hype gap+15
- Incentives55
- Confidence58
A position paper argues that reasoning traces carry safety signal precisely because reinforcement learning treats them as latents rather than outputs, and it asks frontier developers to weigh training decisions against that.
Reality
- Evidence56
- Adoption
- Insufficient
- Hype gap−12
- Incentives55
- Confidence60
MAI-Thinking-1 is in public preview in Microsoft Foundry at $2 per million input tokens. For C# shops the consequence is an interface swap, not a new runtime to operate.
Reality
- Evidence34
- Adoption20
- Hype gap+32
- Incentives46
- Confidence41