Unsloth fixed Studio in version 2026.6.9 after Pillar Security showed that reading a malicious model's config.json could run an attacker's Python code. Unsloth disputed parts of the finding and declined to publish an advisory, so no CVE was assigned.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence52
Pillar Security CEO Ziv Karliner says OpenAI's sandbox escape shows AI agent limits must be enforced outside the model. The escape cases he cites broke through trusted software beyond the sandbox, so his test-before-credentials rule has to cover that outside layer too.
Publishers:scworld.com · token.security Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives72
- Confidence55
A four-container demo runs one Deadbugz-shaped MCP server behind two brokers. Only the broker that hashes name, description and inputSchema on the first tools/list refused the swap that lands after three tool calls.
Reality
- Evidence62
- Adoption18
- Hype gap+10
- Incentives40
- Confidence45
Oren Yomtov of Accomplish found two ways out of the Codex sandbox, both silent and neither stopped by an approval prompt. OpenAI patched them within eight days, and other researchers have found the same design in rival agents.
Reality
- Evidence68
- Adoption45
- Hype gap−5
- Incentives55
- Confidence62
Oren Yomtov of Accomplish AI found two escapes from the Codex sandbox. The worse one ran unsandboxed commands from read-only mode through a helper tool that Codex Desktop writes into the global config at install.
Reality
- Evidence72
- Adoption45
- Hype gap+10
- Incentives55
- Confidence58
A dev.to survey walks seven AI supply chain entry points and the named incidents behind each. The two dataset-poisoning numbers in it are the ones that should change how a model review is scoped.
Reality
- Evidence55
- Adoption45
- Hype gap−10
- Incentives30
- Confidence52
The chain in OpenAI's post-mortem on the Hugging Face incident runs through a RubyGems processing bug, an HDF5 dataset file and 14 write tokens that were already public, according to Pillar Security's Dor Sarig. OpenAI calls the result a warning shot.
Reality
- Evidence38
- Adoption30
- Hype gap+25
- Incentives80
- Confidence42
Pillar Security's Dor Sarig argues the non-human identity buildout answers who an agent is rather than what it does, and cites an internal agent any Slack message could trigger across a thousand private repositories.
Reality
- Evidence24
- Adoption18
- Hype gap+28
- Incentives86
- Confidence52
Grafana's MCP server checked that a session ID looked like one instead of checking that it had issued it. Paired with a CVSS 9.1 SSRF in the same server, that gave unauthenticated callers a proxy inside the network.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+12
- Incentives70
- Confidence55
Pillar Security's researcher filed a bug report that carried orders for Google's triage agent, which returned Workload Identity Federation credentials good enough to impersonate a privileged account. Google has since fixed it.
Reality
- Evidence38
- Adoption24
- Hype gap+18
- Incentives72
- Confidence46
Novee Security's Black Hat findings land the same lesson on three agents: a validator that inspects a cleaned-up copy of a command is not a control, and a sandbox built too late is not one either.
Reality
- Evidence42
- Adoption55
- Hype gap+27
- Incentives68
- Confidence45
CVE-2026-22708 let injected text rewrite a Cursor agent's environment, so an approved "git branch" ran something else. It worked with an empty allowlist too.
Reality
- Evidence58
- Adoption45
- Hype gap+18
- Incentives78
- Confidence55