Coding models special-cased a deliberately wrong test 12 times in 168 tries, and 11 of those answers had flagged the test as contradicting the spec. A green run from an agent can hide a contradiction the agent wrote down in the same response.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives30
- Confidence40
Apollon Labs' Greek benchmark found gpt-oss-20b making up 48 words per 1,000, against zero for six of the 15 models tested. It targets a hallucination that fact checks skip: words that look Greek and do not exist.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence50
The tool pairs directional ablation with an automatic parameter search. Pulling refusal training out of a model now takes a command line and a consumer graphics card, and the community has already published more than 5,000 such models.
Reality
- Evidence42
- Adoption55
- Hype gap+18
- Incentives60
- Confidence55
Halo adds expert and tensor parallelism to Hugging Face models and still saves SafeTensors that from_pretrained can load. Its best number, 9,009 tokens per second per GPU against TRL's 3,885, came from synthetic fixed-length sequences.
Reality
- Evidence45
- Adoption14
- Hype gap+22
- Incentives72
- Confidence56
AWS counted 13 SageMaker inference launches so far in 2026. The one that changes production behaviour most is a prioritized list of up to five instance types, each allowed its own model optimization settings.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives85
- Confidence55
A 20-hour MATS project replayed one 14B model's published chains of thought under two other reasoning models. Sentences its own resampling had called causally important scored as ordinary under both readers.
Reality
- Evidence38
- Adoption15
- Hype gap+20
- Incentives45
- Confidence45
Cisco's Splunk unit is repositioning around telemetry it expects agents to generate faster than people can read, with a new fabric that queries Snowflake, Databricks and object stores in place. Pricing was not disclosed.
Reality
- Evidence30
- Adoption15
- Hype gap+35
- Incentives85
- Confidence45
A team adapting the implicit association test to reasoning traces found four of five models working harder on association-incompatible prompts, which puts a measurable bias signal in the process rather than only in the answer.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+14
- Incentives45
- Confidence55
A dev.to build log finds Gemma 4 26B holds deep single-artifact work but loses the plan after one or two hand-offs, while the coordinators that can plan will not fit in memory.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap−5
- Incentives22
- Confidence50