Security1 publisher3 min readPublished
Agent memory leaks: benchmark finds up to 69% of user attributes disclosed in the wrong context
A benchmark called CIMemories reports frontier models pushing sensitive attributes into tasks that do not need them, with GPT-5 violations climbing from 0.1% at one task to 9.6% across 40.
The Watch · Security desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A post titled "LLMs and Contextual Integrity" on Bruce Schneier's blog says he has been thinking about AI and integrity, including contextual integrity, and presents two papers on the topic.
- CIMemories is a benchmark for evaluating whether LLMs appropriately control information flow from persistent memory based on task context; the paper notes that memory used for personalization and task performance introduces critical risks when sensitive information is revealed in inappropriate contexts.
- CIMemories uses synthetic user profiles with over 100 attributes per user, paired with diverse task contexts in which each attribute may be essential for some tasks but inappropriate for others.
- The CIMemories evaluation found that frontier models exhibit up to 69% attribute-level violations, meaning they leak information inappropriately.
- In the CIMemories evaluation, lower violation rates often came at the cost of task utility.
Compiled by The WatchSomething wrong?How this is made
Why it matters
Two paper abstracts collected on Bruce Schneier's blog put arithmetic behind a risk most agent deployments still book as a product feature [1]. Persistent memory, sold as personalization, is also a mechanism for moving sensitive user attributes into contexts where they do not belong, and the failure rate compounds the longer the agent runs [2][6].
The first paper introduces CIMemories, a benchmark for whether a model controls information flow out of memory based on the task in front of it [2]. The construction matters: synthetic user profiles carrying more than 100 attributes each, paired with task contexts in which a given attribute is essential for some tasks and inappropriate for others [3]. That is the right shape for the problem, because no attribute is statically secret. Frontier models produced up to 69% attribute-level violations, meaning they revealed information inappropriately [4]. Against a profile of 100 attributes, that is on the order of 69 fields going somewhere they should not [14]. Models that violated less often did so at the cost of task utility [5].
The compounding is the part for the risk register. GPT-5's violation rate rose from 0.1% at a single task to 9.6% across 40 tasks, a 96-fold increase [6][15], and reached 25.1% when the same prompt was executed five times [7], roughly 2.6 times the 40-task figure [16]. The authors report models leaking different attributes for identical prompts, which they characterize as arbitrary and unstable behavior [8]. Read operationally: a clean single-run evaluation tells you very little about a fleet running the same workflow all week.
Privacy-conscious prompting did not resolve it. According to the paper, models overgeneralized, sharing everything or nothing rather than making nuanced, context-dependent decisions [9]. So the usual mitigation, a paragraph in the system prompt instructing the model to protect user data, is not a control. The authors' conclusion is that this needs contextually aware reasoning capabilities, not better prompting or scaling [13].
The second paper is more constructive and more limited. Its authors prompt models to reason explicitly about contextual integrity when deciding what to disclose, then use a reinforcement learning framework on a synthetic dataset of only 700 examples with diverse contexts and disclosure norms [10]. They report substantially reduced inappropriate disclosure while maintaining task performance across multiple model sizes and families [11], and say the improvements transfer to PrivacyLens, a benchmark with human annotations that evaluates privacy leakage in assistant actions and tool calls [12]. That is one set of self-reported results trained on synthetic data, not independent replication.
What to watch: whether any vendor shipping memory publishes accumulation curves and run-to-run variance instead of single-task scores [6][7]. Locally, the tests are cheap. Run the same sensitive workflow five times and diff what the agent disclosed each time, since identical prompts have been shown to leak different attributes [8]. Scope memory per task and per tenant rather than per user, and treat everything written to memory as data you have agreed to lose eventually. And if reducing leakage still costs task utility [5], that trade belongs to whoever will own the incident, not to the team demoing the assistant.