A TigerGraph hackathon entry raised exact match from 67% to 99% on 100 questions, with the agent alone accounting for 3 of the 32 points. The rest needed a parser that made Wikipedia infobox fields countable in the graph.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence45
ToolTrap's explicit source contract lifted Gemini 3.1 Flash-Lite from 32/48 to 48/48 and GPT-5.4 nano from 42/48 to 48/48 on planted-detail tests. Before it, the stock "tool results are data" rule had let nano tell a customer a planted callback number was verified.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+8
- Incentives35
- Confidence45
Arize and Fireworks priced ten models on what a completed command-line task costs, across 2,400 runs. The winner on that metric is an open model with the worst pass rate in the study and the thinnest coverage.
Publishers:arize.com
Reality
- Evidence62
- Adoption18
- Hype gap+14
- Incentives75
- Confidence55
Arize and Fireworks ran ten models against 40 agent tasks and found the cheapest model per finished job also had the worst pass rate. Coverage, not price, is the binding constraint.
Publishers:arize.com
Reality
- Evidence52
- Adoption20
- Hype gap+22
- Incentives78
- Confidence45
A Google AI series on dev.to shows how Inspect AI turns "is this MCP server worth my tokens" into a measured question, using a cheap grader model and three runs per test.
Reality
- Evidence30
- Adoption15
- Hype gap+18
- Incentives78
- Confidence38