Scores in the buried-injections benchmark show a well-known prompt-injection classifier catching 6 of 629 hidden attacks at its 0.5 default and 621 at 0.003. A reviewer who re-ran the saved scores found its pooled cross-domain rates hide how detectors do on their weakest suite.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence55
Dreadnode benchmarked eight judges as pre-execution gates for offensive agents. The top performer reached human range on 4,897 real tool calls, which still leaves a tenth of out-of-scope actions arriving at whatever sits on the other end.
Publishers:dreadnode.io
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+12
- Incentives68
- Confidence44
Aikido spent 11.7 billion tokens rediscovering 32 fresh CVEs with ten models, three attempts each. The number that should move a scanning budget is the marginal cost of the second and third pass.
Reality
- Evidence44
- Adoption18
- Hype gap+32
- Incentives76
- Confidence52
A framework running at Uber for over ten months argues endpoint tooling is structurally blind to agent reasoning. On the authors' own benchmark it still misses a third of attacks.
Reality
- Evidence46
- Adoption54
- Hype gap+24
- Incentives68
- Confidence44
CrowdStrike argues AI security evaluations have converged on vulnerability discovery because it is easy to score, while most breaches still start with stolen credentials, phishing and trusted access.
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+14
- Incentives78
- Confidence34