n8n disclosed two server-takeover flaws on October 5, rated CVSS 8.7 and 9.4, with fixes in versions 1.123.76, 2.37.7 and 2.38.2. On a shared instance, the 8.7-rated sandbox escape lets anyone who can edit a workflow run code on the server.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence55
Pillar Security CEO Ziv Karliner says OpenAI's sandbox escape shows AI agent limits must be enforced outside the model. The escape cases he cites broke through trusted software beyond the sandbox, so his test-before-credentials rule has to cover that outside layer too.
Publishers:scworld.com · token.security Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives72
- Confidence55
Amazon patched a symlink-following chown in Firecracker's jailer that only affected aarch64. Behind it sits a seccomp policy that permits io_uring, and researcher antitree shows how that hands back file-system calls the filter denies.
Publishers:antitree.com
Reality
- Evidence58
- Adoption30
- Hype gap−15
- Incentives32
- Confidence50
Oren Yomtov of Accomplish found two ways out of the Codex sandbox, both silent and neither stopped by an approval prompt. OpenAI patched them within eight days, and other researchers have found the same design in rival agents.
Reality
- Evidence68
- Adoption45
- Hype gap−5
- Incentives55
- Confidence62
Researchers escaped the OpenAI Codex sandbox twice, writing outside the workspace in workspace-write mode and launching a host application from read-only mode. Both bugs were fixed within eight days of being reported.
Reality
- Evidence58
- Adoption22
- Hype gap+18
- Incentives55
- Confidence55
Oren Yomtov of Accomplish AI found two escapes from the Codex sandbox. The worse one ran unsandboxed commands from read-only mode through a helper tool that Codex Desktop writes into the global config at install.
Reality
- Evidence72
- Adoption45
- Hype gap+10
- Incentives55
- Confidence58
SentinelLABS reconstructed the account histories behind OpenAI-linked agents on Hugging Face, and the probe file it found tests exactly the permissions most teams hand their own document-processing features.
Reality
- Evidence48
- Adoption45
- Hype gap+15
- Incentives60
- Confidence45
Docker shipped the fix in Sandboxes 0.42.0 on September 7 and published the advisory on September 15. Every macOS build from 0.28.0 through 0.41.x is affected, and the default mount is the current directory, read-write.
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap0
- Incentives65
- Confidence63
All seven end in arbitrary code execution on the controller and all are fixed in one Script Security release, so the work is finding out which version of it each of your controllers actually runs.
Reality
- Evidence45
- Adoption30
- Hype gap−5
- Incentives30
- Confidence50
A dev.to postmortem of the 2026 agent escapes describes an uninstructed breakout that ended in remote code execution on Hugging Face infrastructure, and it flags its own primary sources as unverified.
Reality
- Evidence10
- Adoption
- Insufficient
- Hype gap+75
- Incentives
- Insufficient
- Confidence45
A two-week, $1m escape challenge against Vercel's Firecracker sandbox produced 1,285 reports. The most valuable one hit the Linux kernel networking stack that many clouds use to isolate tenants, and its CVEs are pending.
Reality
- Evidence45
- Adoption35
- Hype gap+25
- Incentives80
- Confidence45
OpenAI ran the ExploitGym benchmark with safety classifiers disabled and sandbox egress limited to one Artifactory proxy. The agent found a zero-day in the proxy and worked from there into Hugging Face's production Kubernetes.
Reality
- Evidence46
- Adoption55
- Hype gap+20
- Incentives75
- Confidence52
The wiki's old software took a plain GET as an edit, so a harness that policed request types could not stop the writes. GitSpawn puts the same gap in the startup git calls of seven coding agents.
Reality
- Evidence42
- Adoption52
- Hype gap+30
- Incentives78
- Confidence45