Broken answer keys and graders that punish correct tool calls drove the verdicts. Epoch AI says it stops each review once it has enough evidence, so the published defect counts are floors.
Reality
- Evidence65
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence62
NVIDIA's developer blog sets out the five-level rollup from step to benchmark and the two scores that read the same trace. Step-level says where the chain broke. End-to-end reads the environment and says whether the refund posted.
Reality
- Evidence55
- Adoption20
- Hype gap+12
- Incentives50
- Confidence60
A Google XR team fine-tuned Gemma-3 models on data generated answer-first and says they match proprietary LLMs on tools they never trained on. The post's cost and pass-rate claims carry no numbers.
Reality
- Evidence30
- Adoption10
- Hype gap+40
- Incentives60
- Confidence35
tau-bench scores agents on the database state they leave behind and then reruns each task eight times, and the consistency figure that falls out is a better launch gate than the single-trial accuracy most teams quote.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+8
- Incentives55
- Confidence58
A dev.to build log finds Gemma 4 26B holds deep single-artifact work but loses the plan after one or two hand-offs, while the coordinators that can plan will not fit in memory.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap−5
- Incentives22
- Confidence50