Build1 publisher3 min readPublished
Three of 11 catalogued attacks on agent memory and learning reached production systems
CyberXDefend counts about 11 named attacks on AI agent memory and learning, three of them demonstrated on production ChatGPT and OpenClaw. It found no confirmed criminal campaign, but in each case one poisoned write outlives the session, the property behind OWASP's ASI06 category.
The Engineer · Build desk

What happened
- In 2024 Johann Rehberger got hidden instructions in documents and web pages saved as ChatGPT user preferences, which then sent conversation data to an attacker's server across sessions.
- Radware's ZombieAgent proof of concept, in January 2026, chained ChatGPT connectors and memory so indirect prompt injection persisted across sessions and spread through email attachments.
- Researchers sent OpenClaw payloads through a real Gmail integration in July 2026, getting past spam filtering in more than half of attempts, and coverage reported no substantive fix yet.
- MemMorph, from May 2026, steered which tools an agent picks with up to 85.9% success using only three planted records disguised as technical facts and policies.
- P-Trojan, presented at AAAI 2026, kept a backdoor alive through continual fine-tuning with over 99% persistence on Qwen2.5 and LLaMA3 while clean-task accuracy held.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint OpenAI's patch closed the exfiltration path but left memory-manipulating injection open, so anything already written to an agent's memory has to be handled as untrusted input for as long as it is stored.
- exposure Agents that read email, web pages or repository files and save what they learn can be poisoned by whoever writes that content, without any access to the memory store.
- cost Investigating delayed poisoning means keeping a dated history of memory writes long enough to find the one that went bad weeks before the agent misbehaved.
- decision Anyone building a learning-through-use loop has to decide how outcome signals get authenticated, because forged outcomes can train the decision head directly.
All of these attacks depend on state that survives the session. A stateless model forgets an attack when the session ends. A model that remembers or learns does not [1]. Anything that lands in memory comes back in later sessions as context the agent treats as its own. The CyberXDefend post calls the result "one bad write, exploited forever" [15].
According to the post, persistence is why OWASP separates Memory and Context Poisoning, ASI06 in its Top 10 for Agentic Applications, from ordinary prompt injection [2]. The author wrote that he had not opened the OWASP primary and worked from a secondary summary published by Akto [3].
In the memory cases, the payload arrives through content the agent reads. Rehberger's instructions sat in documents and web pages [4]. MemoryGraft used harmless-looking content such as a README. Weeks later the agent retrieved the poisoned "successful experience" and copied it [7]. eTAMP, in April 2026, showed cross-session, cross-site compromise with no direct access to the memory at all [8].
OpenAI's fix covered the exit. It patched the exfiltration path in Rehberger's attack, and the post says OpenAI acknowledged that prompt injection that manipulates memory storage is still an open problem [5]. Radware's ZombieAgent went back through memory against the same product in January 2026 [6].
The record shows these attacks working on production systems; it does not show criminals using them. Of about 11 named attacks or studies, three ran against production systems and the rest are research demos [13], roughly eight [1]. The author found no confirmed criminal campaign in the wild. He argued that with attacks built to stay dormant for weeks, absence of evidence is not reassuring [14]. An IEEE Access study supports the dormancy point. Across 2,614 multi-step attack trajectories, some poisoning stayed indistinguishable from normal behavior until much later interactions [11].
The research figures describe the researchers' setups. MemMorph's 85.9% is a ceiling from a research demo, reached with three planted records posing as technical facts and policies [10][13]. I'd expect it to transfer to an agent that feeds unvetted retrieved records straight into tool selection. It should transfer less well where each record carries provenance the agent checks before acting. P-Trojan's persistence figure was measured on Qwen2.5 and LLaMA3 [12]. A team fine-tuning a different base model would need to reproduce it before planning around it.
The post closes with a design review of Ontogen, a project aimed at skill learning through use [18]. Its feedback(decision_id, outcome) call is an attack surface. An attacker who can fake outcomes can train the decision head, which is the MemMorph pattern written into slow traces instead of text [16]. The author credits the design with per-user regions, age-indexed rollback and a surprise gate [17]. I think age-indexed rollback matters most of the three. Poisoning that stays quiet for weeks [11] can only be undone by returning state to a point before the bad write. That requires knowing when each write happened.
What to watch
- A substantive OpenClaw fix for the Gmail-delivered payloads disclosed in mid-July 2026, or a statement that none is coming.
- The first confirmed criminal campaign using memory poisoning; the post's tally has none, and payloads designed to stay dormant would make one slow to surface.
- OWASP's own ASI06 text, since the post's account of how it differs from prompt injection comes from a secondary summary.