Pi shipped MCP in v0.99.0 using a sandbox that keeps tool schemas out of context, where three servers took 143,000 of Perplexity's 200,000 tokens. Other harnesses can copy the design if they are willing to host an interpreter that runs model-written code.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence35
Polyglot's developer ran six coding agents 30 times on each of seven local models, and three never made a tool call on models that write calls as text. The author's own error bars say 30 runs can sort agents into tiers but cannot rank two agents inside one.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives70
- Confidence55
layaAgent's 421M-parameter encoder picks tool arguments from extracted spans and gets 71% right on its author's test split. Unsure steps go to a local 1.5B LLM or the user, and any tool that changes data waits for approval whichever system chose it.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives40
- Confidence40
Codex batched two view_image calls and put a resize notice between their outputs. OpenAI's Responses API pairs items by call ID wherever they sit. DeepSeek's endpoint wants them adjacent, and returns the same 400 on every later message.
Reality
- Evidence62
- Adoption22
- Hype gap−8
- Incentives32
- Confidence58
A consultancy's account of production MCP work on Django client apps calls the protocol a transport layer. Its sample server calls django.setup(), imports the Order model, and dispatches tool calls on a string.
Reality
- Evidence60
- Adoption25
- Hype gap+12
- Incentives55
- Confidence55
The MIT-licensed TypeScript framework at v0.16 exposes seven named run phases with hooks on each side, and its case for harness over model rests on one incident-triage run that a 4B local model and Claude both finished.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+28
- Incentives78
- Confidence44
A dev.to design post puts a per-task scope check under every tool invocation and caps the agent's cloud credentials at fifteen minutes. The restrictions work by taking capability away from the model.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+30
- Incentives20
- Confidence45
Two models read the same tool description and disagreed about one field name. The fix moved the shape into the JSON Schema for the nine pattern kinds that account for 85 percent of emissions, and left the rest loose.
Reality
- Evidence45
- Adoption20
- Hype gap−10
- Incentives25
- Confidence55
A dev.to design post starts from a publish_page call that succeeded while its response timed out. Its answer is a layer of application code that owns authorization, idempotency and verification before anything reaches the external system.
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+12
- Incentives18
- Confidence44
The Go port of Gemma's chat template filters schema keywords at all four call sites where the Jinja macro filters at one, so the parameter name reaches the model with no definition and stays on the required list.
Reality
- Evidence78
- Adoption20
- Hype gap−10
- Incentives25
- Confidence70
A dev.to design splits the agent loop into observe, propose, verify and commit phases. The Zod schema check is mandatory, the domain-invariant guard is optional per tool, and a retry counter on state is what ends a runaway run.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+35
- Incentives25
- Confidence55
A no-code support agent called an endpoint it invented, was refused twice, and still told the customer it could see her charges. Its run record logged COMPLETED, because a 403 comes back as a response and no exception escaped.
Reality
- Evidence58
- Adoption10
- Hype gap−6
- Incentives68
- Confidence45
More than half the traffic to one developer's API documentation now comes from AI agents. The changes that followed landed in the summary and description fields, where an agent does operation selection and error recovery.
Reality
- Evidence34
- Adoption15
- Hype gap+10
- Incentives45
- Confidence45
Copilot Extensions, Claude tools, LangChain tools and Spring AI functions share one three-part anatomy. A dev.to walkthrough shows that the schema ports for free and the registration does not.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+30
- Incentives30
- Confidence32
The tools specification dated 2026-07-28 lets a server advertise a cache lifetime for tools/list, which turns every tool name into a contract clients can keep holding well after you deploy a change to it.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+15
- Incentives55
- Confidence42
AgentArch sweeps orchestration, ReAct versus function calling, memory scope and a thinking tool across 18 setups on frontier models. Because the best cell moves with the model, the grid is what you reuse and the ceiling is what you budget for.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+8
- Incentives55
- Confidence50
A Gemini function call arrives as a name and structured arguments, and nothing in the SDK decides whether the caller is permitted or whether the booking already went through. That gate is yours to build.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+9
- Incentives28
- Confidence44
A dev.to walkthrough on sandboxing LLM tool calls names old web-app failure modes and old web-app controls. The delta is when you apply them: at registration, not after the first exfiltration.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+28
- Incentives35
- Confidence46
A developer's LLM issued the same refund three times after a timeout. The defect was in the wiring, not the model, and the fixes are schema validation and scoped toolsets.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+12
- Incentives21
- Confidence44