Penn State researchers moved agent attack-success rates by up to 13.21 points by renaming tools, with the attack, task and policy held fixed. The size of the shift varied by model and benchmark, so a robustness score only describes the tool names and schemas it was measured on.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence45
Meta's Prompt Guard 2 caught 6 of 629 buried attacks at a 0.5 cutoff and 621 at 0.003, in a benchmark posted on dev.to. Health checks pass the same at both settings, so only attacks sent through the deployed threshold show which one a team runs.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence40
One developer scored ten open-source prompt-injection detectors against 629 AgentDojo attacks and found the best caught about half. His own gate at the tool boundary then held the legitimate payment as well as the injected one.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence45
The AgentDojo framework pairs 97 tool-using tasks with 629 security test cases in one stateful environment, and current LLMs solve fewer than 66 percent of those tasks with no attacker in play. That baseline bounds what its defense numbers can mean.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−10
- Incentives60
- Confidence52
The paper's enforcement point reads a whole session's provenance before committing any action, so a read-email, call-payroll, send-externally sequence gets judged as a flow. Adoption means instrumenting every execution path.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence35
The Agentic AI Detection and Response system runs in production at Uber. Three of its five components are now open source, including a 300-task, 133-MCP-server benchmark.
Reality
- Evidence58
- Adoption34
- Hype gap+22
- Incentives66
- Confidence52
A framework running at Uber for over ten months argues endpoint tooling is structurally blind to agent reasoning. On the authors' own benchmark it still misses a third of attacks.
Reality
- Evidence46
- Adoption54
- Hype gap+24
- Incentives68
- Confidence44