A dev.to post gives every failed coding-agent run one of four labels and scores skill only over the runs where the harness stayed healthy. SSH drops and disk-full errors get published as their own rates.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives15
- Confidence55
Datadog Security Labs put one document-portal prompt through three coding agents in both modes and audited all six builds. An insecure direct object reference appeared in each, and one build per cell leaves mode effects entangled with noise.
Publishers:securitylabs.datadoghq.com
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives68
- Confidence45
A replay that held the agent's actions fixed produced different labels before and after delayed operations resolved, and one late write moved the following run's score.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap−8
- Incentives34
- Confidence50
A self-published log of 157 agent runs says planning depth beats model size. Model size was never one of the dimensions the runs varied.
Reality
- Evidence14
- Adoption8
- Hype gap+66
- Incentives58
- Confidence63
One good demo is not evidence that a new prompt, skill or rules file helped. A control arm, thirty tasks and a held-constant model will tell you, within limits worth knowing before you build it.
Reality
- Evidence18
- Adoption
- Insufficient
- Hype gap+24
- Incentives28
- Confidence42
A preprint from ServiceNow AI Research and university co-authors reports that a plain web agent with a longer horizon matches or beats AWM, ASI and ReasoningBank, often on fewer tokens.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives45
- Confidence42
A proxy on the MCP stdio pipe caught one client killing most of its own tool calls before they reached the server. On the wire it looked exactly like a weak model.
Reality
- Evidence68
- Adoption30
- Hype gap−18
- Incentives28
- Confidence62
One research agent, built three times on three hyperscaler frameworks with the same prompt and the same search tool, diverged in nine ways. None of them were protocol problems.
Reality
- Evidence62
- Adoption34
- Hype gap+4
- Incentives33
- Confidence55
A thirty-run experiment finds agents followed long, buried rules anyway. The two failures came where the file contradicted the repository, not where it was long or buried.
Reality
- Evidence52
- Adoption18
- Hype gap+12
- Incentives32
- Confidence55
Princeton and the UK AI Security Institute gave a frontier agent six days and $3,000 to answer real unpublished research questions. The original authors reviewed the output and rejected both papers.
Reality
- Evidence55
- Adoption15
- Hype gap+15
- Incentives60
- Confidence45