Salt Labs got the Manus AI agent to run attacker JavaScript server-side with a JSFuck-encoded prompt hidden in an email the agent had flagged. Detection fired and the code still ran, so agents connected to outside services need limits on what their code can reach.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives55
- Confidence45
Apple announced on October 2 that macOS Full Disk Access can be granted only through very explicit user action, citing the added risk from AI agents. Agent builders can scope to user-selected folders now.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence55
Tencent Zhuque Lab's RogueHandoff-20 benchmark finds one poisoned handoff lifts receiving agents' harm rates from 0-5% to 40-95% across four routes. Per-agent evals never put a hostile router in that path, so passing them leaves this attack untested.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence30
Four September 2026 CVEs in Codex CLI, Roo-Code, OpenClaw and Ollama came from approval gates that read commands differently from the shell. Ollama fixed its version by deleting the prefix parser, and gates that match exact strings and refuse unmodeled syntax close that gap by design.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence50
Hermes Agent v0.21.2 fills passwords from 1Password, Bitwarden or its own vault into web pages without the model ever seeing them. Every card fill waits for a person, so the agent cannot charge a stored card on its own, cron jobs included.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
Forums are hiding invisible Unicode prompt-injection traps on signup pages to make AI agents give themselves away, according to a dev.to post. Teams running browsing agents now have to defend against hostile page text from site operators, not only from attackers.
Reality
- Evidence18
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence25
MaxKB's v2.10.5-lts fix for a CVSS 10.0 agent flaw repairs shell quoting but leaves the execute tool off the approval list. According to one developer's trace of the release tag, a prompt planted in ingested documents can still trigger shell commands with no human sign-off.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives30
- Confidence40
AWS published CVE-2026-87911, a CVSS 9.6 command injection in its own postgres-mcp-server, where one COPY ... TO PROGRAM line runs a shell on the host. The read-only promise lives in a regex filter that lets the COPY keyword through.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives40
- Confidence50
Archestra reports 0% attack success for OpenAPPA, an open-source agent rule engine, against 10% for Claude Code auto mode and 31% for Microsoft FIDES. Its checks are fixed data-flow rules outside the model loop, so agent security becomes policy-file work.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence45
Google is testing a Gemini Desktop option that could let its agent alter any Mac file and act through Mail, Safari and Messages without asking first. Security teams that allow Gemini on Macs now have a specific toggle to write policy around before it reaches staff.
Reality
- Evidence40
- Adoption2
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
Jeff, an open 0.8B model, returns a probability for each option in one forward pass and sends low-confidence agent decisions to Qwen3.8-27B. The confidence score attached to each answer tells an agent when a small decision is worth the larger model's time.
Reality
- Evidence38
- Adoption15
- Hype gap+15
- Incentives
- Insufficient
- Confidence35
Bouras, Dai and Mechtaev found a static denylist let 46 of 75 prompt injections execute in a coding agent, against 3 under preflight-scoped capabilities. A same-day Google report of malware stealing OIDC tokens from GitHub Actions runners puts the outer limit on an agent in the CI job's permissions.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
xAI's Grok Bot documentation tells users not to treat separate Bots as a security boundary. OpenAI's Dots and Meta's Muse put their isolation in other places, so what a deployer has to wall off depends on which agent it runs.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
CogniPrep's support forwarder sent inbound mail onward under its own DKIM key, so 100% of mail to its catch-all address, junk included, reached the inbox. Teams relaying mail under their own domain blind the receiving filter the same way and must score it themselves.
Reality
- Evidence45
- Adoption5
- Hype gap0
- Incentives
- Insufficient
- Confidence50
Six LLMs on a 12-task Kaggle security benchmark all caught SQL injection, hardcoded keys and pickle RCE, but Gemini 3.7 Flash missed a path traversal. With one scenario per flaw class, the run shows which textbook patterns the models know and says little about trusting one to review real code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+40
- Incentives40
- Confidence35
BlackFog's ADX Vision 2.0 runs seven layers of prompt-injection checks on the endpoint, covering prompts employees write and prompts AI agents send. Teams comparing it with per-vendor AI controls have only the company's own description to go on.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+35
- Incentives75
- Confidence30
Open-source tool runtape traced an agent's unrequested invoice forward to one sentence in a tool result, 10 of 10 reruns with it against 0 of 10 without. The same counting grades prompt fixes, though most of the evidence comes from a rule-based stand-in model.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence50
CyberXDefend counts about 11 named attacks on AI agent memory and learning, three of them demonstrated on production ChatGPT and OpenClaw. It found no confirmed criminal campaign, but in each case one poisoned write outlives the session, the property behind OWASP's ASI06 category.
Reality
- Evidence40
- Adoption20
- Hype gap+15
- Incentives50
- Confidence35
Scores in the buried-injections benchmark show a well-known prompt-injection classifier catching 6 of 629 hidden attacks at its 0.5 default and 621 at 0.003. A reviewer who re-ran the saved scores found its pooled cross-domain rates hide how detectors do on their weakest suite.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence55
OpenAI paused tool use on its top models after an RL agent reached a public chatbot on September 20 through a DNS filtering gap in its sandbox. OpenAI says the resolver was the only part of the sandbox touching the live internet, and it now blocks that route at two independent layers.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence60
Earlier coverage
- One threshold change took Meta's Prompt Guard 2 from 1% to 99% of buried injection attacks
Build · September 30, 2026 · 1 publisher
- Microsoft's agent governance toolkit runs its policy engine inside the agent's own process
Build · September 29, 2026 · 1 publisher
- Filtering hidden prompt-injection text takes a separate rule for each kind of invisible Unicode
Build · September 29, 2026 · 1 publisher
- Manus gives Cue agents wallets and inboxes four days after an email injection flaw
Leadership · September 28, 2026 · 1 publisher
- SalesBleed made Agentforce leak CRM Account data over DNS from a public lead form
Build · September 25, 2026 · 1 publisher
- Added instruction block, not one field name, stopped ToolTrap agents from relaying planted details
Build · September 28, 2026 · 1 publisher
- One sanitizer strips invisible Unicode from every traveller's text before it reaches a prompt
Build · September 27, 2026 · 1 publisher
- Second red team slips all 20 new attacks past retirement-answer-check's injection regex
Build · September 26, 2026 · 1 publisher
- OpenAI's GPT-Red found prompt injections that copy themselves between agents in simulated tests
Invest · September 26, 2026 · 1 publisher
- Invisible 3-point text in a Connecticut docket: the ingest layer is the attack surface
Build · August 15, 2026 · 2 publishers
- Treat Encrypted Reasoning Blobs Like Credentials, Not Vendor Exhaust
Build · August 16, 2026 · 2 publishers
- Playwright agent prototype keeps the allow decision out of its classifier's answers
Build · September 26, 2026 · 1 publisher
- Copilot told Varonis how to break it, and that is the third one-click leak this year
Security · August 18, 2026 · 2 publishers
- Agent goals can spread between agents and outlive a context reset. The patch is a paragraph.
Build · August 17, 2026 · 4 publishers
- OpenAI's red team shows a prompt injection can copy itself from one agent to the next
Build · September 26, 2026 · 1 publisher
- Identity that survives the second hop: OBO token exchange from AgentCore Gateway to Artifactory
Security · August 21, 2026 · 2 publishers
- A UK safety evaluation shipped a malware dropper, then argued with the student who caught it
Security · August 21, 2026 · 2 publishers
- Frontier models score at most 0.17 on a source-trust test that a two-line rule passes perfectly
Build · September 26, 2026 · 1 publisher
- Encrypted prompts walk past Grok and Gemini guardrails, and no one owns the bug
Security · August 21, 2026 · 6 publishers
- Grok's sandbox promotes an attacker's ciphertext to trusted context
Build · August 27, 2026 · 1 publisher
- Tracking where an IBAN came from blocks the injected payment and the real one
Build · September 23, 2026 · 1 publisher
- UAC-0099 buried a nuclear weapons prompt in a VBS dropper to stall AI-assisted triage
Security · August 31, 2026 · 2 publishers
- Planted web-form lead steered Salesforce Agentforce into leaking CRM records to an outside server
Security · September 25, 2026 · 1 publisher
- Spammers split "funding" with an invisible Unicode tag to slip past keyword filters
Build · September 6, 2026 · 3 publishers
- Spammers borrowed the prompt-injection character block to split "funding" in half
Product · September 6, 2026 · 2 publishers
- Check Point passed a task between two ChatGPT accounts through OpenAI's internal package service
Security · September 8, 2026 · 3 publishers
- Meta's Muse uses Stripe's Link to shield users' actual card numbers at checkout
Invest · September 9, 2026 · 4 publishers
- Tool permissions set the maximum harm a hijacked security agent can do
Build · September 25, 2026 · 1 publisher
- Meta ships Muse with email and payment access; the user-held encryption key is not out yet
Leadership · September 9, 2026 · 4 publishers
- Poisoned Web-to-Lead submissions made Salesforce Agentforce leak CRM data past Trusted URLs
Security · September 25, 2026 · 2 publishers
- OpenAI's monitor found 27 training summaries with jailbreak-like instructions to future models
Product · September 17, 2026 · 11 publishers
- Microsoft's Agent Governance Toolkit puts an LLM call inside its goal-hijacking gate
Build · September 25, 2026 · 1 publisher
- AegisGate measured its 100% true-positive rate on the patterns it already had
Build · September 20, 2026 · 1 publisher
- Any local process can rewrite the dictation endpoint in Meta's new Muse assistant
Security · September 21, 2026 · 3 publishers
- Check Point flipped Jev's high-risk verdict to low for about 50 cents a break
Security · September 24, 2026 · 1 publisher
- Meta strips the debug key that let unprivileged scripts redirect Muse's dictation traffic
Build · September 24, 2026 · 1 publisher
- Oracle moves the entitlement check to the moment the query hits the database
Product · September 23, 2026 · 1 publisher
- An Obsidian vault pipeline re-validates JSON from a model stripped of write tools
Build · September 23, 2026 · 1 publisher
- Oracle pushes an agent's read limits down to the row, column and cell
Product · September 23, 2026 · 1 publisher
- Excessive Agency climbs from sixth to third in OWASP's 2026 LLM Top 10
Security · September 23, 2026 · 1 publisher