KaliBench, 8,504 query-command pairs across 1,642 Kali tools, found no open-weight model among 24 configurations topped 42% exact-command accuracy. Scores rise when the model is handed the tool name, but production analysts describe only intent.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
A developer's Claude agent crashed on 388 of 1,200 runs because its loop answered only the first of several parallel tool calls. Deleting the extra calls stopped the errors but cut its eval score from 46 to 38 out of 50.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+8
- Incentives
- Insufficient
- Confidence50
ToolTrap's explicit source contract lifted Gemini 3.1 Flash-Lite from 32/48 to 48/48 and GPT-5.4 nano from 42/48 to 48/48 on planted-detail tests. Before it, the stock "tool results are data" rule had let nano tell a customer a planted callback number was verified.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+8
- Incentives35
- Confidence45
Omni Calculator says any number a customer acts on should come from a deterministic tool, after its benchmark scored AI models at 48.4% to 70.4% on math. Its own Toronto BMW example went wrong at the input, so a calculation engine protects operators only when the right data reaches it.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives80
- Confidence40
A dev.to walkthrough of MCP internals locates the integration saving in deployment coupling and leaves the residual risk with the host process that validates each tool call and raises the consent prompt.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+22
- Incentives25
- Confidence42
The same platform ships regex and BM25 ranking in front of its tool catalog, while its memory store offers only depth, limit, page, path_prefix and view, so ranking memories is work the developer supplies.
Reality
- Evidence66
- Adoption
- Insufficient
- Hype gap−5
- Incentives70
- Confidence62
A Google XR team fine-tuned Gemma-3 models on data generated answer-first and says they match proprietary LLMs on tools they never trained on. The post's cost and pass-rate claims carry no numbers.
Reality
- Evidence30
- Adoption10
- Hype gap+40
- Incentives60
- Confidence35
Two Bern researchers solved a rotation CAPTCHA in 0.006 seconds using circle detection from the 1970s, then fed that answer to frontier models as a tool result and watched one of them argue with it.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+18
- Incentives30
- Confidence45
A published production loop for customer service agents shows where the cost really sits: not the model, but the classify-execute-confirm middle where a write hits a payment processor.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+28
- Incentives72
- Confidence44