Skip to content

Topic

Prompt injection

A security flaw where untrusted text in an AI model's context is treated as commands, letting attackers hijack its behavior.

Current stories

build1 publisher

Hidden Unicode traps on signup forms turn prompt injection against AI agents

Forums are hiding invisible Unicode prompt-injection traps on signup pages to make AI agents give themselves away, according to a dev.to post. Teams running browsing agents now have to defend against hostile page text from site operators, not only from attackers.

Publishers:dev.to

Reality

Evidence18
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence25
build1 publisher

One COPY line runs a shell despite AWS's read-only Postgres MCP filter

AWS published CVE-2026-87911, a CVSS 9.6 command injection in its own postgres-mcp-server, where one COPY ... TO PROGRAM line runs a shell on the host. The read-only promise lives in a regex filter that lets the COPY keyword through.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives40
Confidence50
build1 publisher

A static denylist stopped one more prompt injection than no protection in a coding-agent study

Bouras, Dai and Mechtaev found a static denylist let 46 of 75 prompt injections execute in a coding agent, against 3 under preflight-scoped capabilities. A same-day Google report of malware stealing OIDC tokens from GitHub Actions runners puts the outer limit on an agent in the CI job's permissions.

Publishers:dev.to

Reality

Evidence45
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence50
build1 publisher

Idiomatic os.path.join trips Gemini 3.7 Flash in a 12-task LLM security benchmark

Six LLMs on a 12-task Kaggle security benchmark all caught SQL injection, hardcoded keys and pickle RCE, but Gemini 3.7 Flash missed a path traversal. With one scenario per flaw class, the run shows which textbook patterns the models know and says little about trusting one to review real code.

Publishers:dev.to

Reality

Evidence30
Adoption
Insufficient
Hype gap+40
Incentives40
Confidence35

Earlier coverage

  1. One threshold change took Meta's Prompt Guard 2 from 1% to 99% of buried injection attacks

    Build · September 30, 2026 · 1 publisher

  2. Microsoft's agent governance toolkit runs its policy engine inside the agent's own process

    Build · September 29, 2026 · 1 publisher

  3. Filtering hidden prompt-injection text takes a separate rule for each kind of invisible Unicode

    Build · September 29, 2026 · 1 publisher

  4. Manus gives Cue agents wallets and inboxes four days after an email injection flaw

    Leadership · September 28, 2026 · 1 publisher

  5. SalesBleed made Agentforce leak CRM Account data over DNS from a public lead form

    Build · September 25, 2026 · 1 publisher

  6. Added instruction block, not one field name, stopped ToolTrap agents from relaying planted details

    Build · September 28, 2026 · 1 publisher

  7. One sanitizer strips invisible Unicode from every traveller's text before it reaches a prompt

    Build · September 27, 2026 · 1 publisher

  8. Second red team slips all 20 new attacks past retirement-answer-check's injection regex

    Build · September 26, 2026 · 1 publisher

  9. OpenAI's GPT-Red found prompt injections that copy themselves between agents in simulated tests

    Invest · September 26, 2026 · 1 publisher

  10. Invisible 3-point text in a Connecticut docket: the ingest layer is the attack surface

    Build · August 15, 2026 · 2 publishers

  11. Treat Encrypted Reasoning Blobs Like Credentials, Not Vendor Exhaust

    Build · August 16, 2026 · 2 publishers

  12. Playwright agent prototype keeps the allow decision out of its classifier's answers

    Build · September 26, 2026 · 1 publisher

  13. Copilot told Varonis how to break it, and that is the third one-click leak this year

    Security · August 18, 2026 · 2 publishers

  14. Agent goals can spread between agents and outlive a context reset. The patch is a paragraph.

    Build · August 17, 2026 · 4 publishers

  15. OpenAI's red team shows a prompt injection can copy itself from one agent to the next

    Build · September 26, 2026 · 1 publisher

  16. Identity that survives the second hop: OBO token exchange from AgentCore Gateway to Artifactory

    Security · August 21, 2026 · 2 publishers

  17. A UK safety evaluation shipped a malware dropper, then argued with the student who caught it

    Security · August 21, 2026 · 2 publishers

  18. Frontier models score at most 0.17 on a source-trust test that a two-line rule passes perfectly

    Build · September 26, 2026 · 1 publisher

  19. Encrypted prompts walk past Grok and Gemini guardrails, and no one owns the bug

    Security · August 21, 2026 · 6 publishers

  20. Grok's sandbox promotes an attacker's ciphertext to trusted context

    Build · August 27, 2026 · 1 publisher

  21. Tracking where an IBAN came from blocks the injected payment and the real one

    Build · September 23, 2026 · 1 publisher

  22. UAC-0099 buried a nuclear weapons prompt in a VBS dropper to stall AI-assisted triage

    Security · August 31, 2026 · 2 publishers

  23. Planted web-form lead steered Salesforce Agentforce into leaking CRM records to an outside server

    Security · September 25, 2026 · 1 publisher

  24. Spammers split "funding" with an invisible Unicode tag to slip past keyword filters

    Build · September 6, 2026 · 3 publishers

  25. Spammers borrowed the prompt-injection character block to split "funding" in half

    Product · September 6, 2026 · 2 publishers

  26. Check Point passed a task between two ChatGPT accounts through OpenAI's internal package service

    Security · September 8, 2026 · 3 publishers

  27. Meta's Muse uses Stripe's Link to shield users' actual card numbers at checkout

    Invest · September 9, 2026 · 4 publishers

  28. Tool permissions set the maximum harm a hijacked security agent can do

    Build · September 25, 2026 · 1 publisher

  29. Meta ships Muse with email and payment access; the user-held encryption key is not out yet

    Leadership · September 9, 2026 · 4 publishers

  30. Poisoned Web-to-Lead submissions made Salesforce Agentforce leak CRM data past Trusted URLs

    Security · September 25, 2026 · 2 publishers

  31. OpenAI's monitor found 27 training summaries with jailbreak-like instructions to future models

    Product · September 17, 2026 · 11 publishers

  32. Microsoft's Agent Governance Toolkit puts an LLM call inside its goal-hijacking gate

    Build · September 25, 2026 · 1 publisher

  33. AegisGate measured its 100% true-positive rate on the patterns it already had

    Build · September 20, 2026 · 1 publisher

  34. Any local process can rewrite the dictation endpoint in Meta's new Muse assistant

    Security · September 21, 2026 · 3 publishers

  35. Check Point flipped Jev's high-risk verdict to low for about 50 cents a break

    Security · September 24, 2026 · 1 publisher

  36. Meta strips the debug key that let unprivileged scripts redirect Muse's dictation traffic

    Build · September 24, 2026 · 1 publisher

  37. Oracle moves the entitlement check to the moment the query hits the database

    Product · September 23, 2026 · 1 publisher

  38. An Obsidian vault pipeline re-validates JSON from a model stripped of write tools

    Build · September 23, 2026 · 1 publisher

  39. Oracle pushes an agent's read limits down to the row, column and cell

    Product · September 23, 2026 · 1 publisher

  40. Excessive Agency climbs from sixth to third in OWASP's 2026 LLM Top 10

    Security · September 23, 2026 · 1 publisher