Security1 distinct publisher3 min readUpdated
A benchmark called CIMemories reports frontier models pushing sensitive attributes into tasks that do not need them, with GPT-5 violations climbing from 0.1% at one task to 9.6% across 40.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
Two paper abstracts collected on Bruce Schneier's blog put arithmetic behind a risk most agent deployments still book as a product feature [1]. Persistent memory, sold as personalization, is also a mechanism for moving sensitive user attributes into contexts where they do not belong, and the failure rate compounds the longer the agent runs [2][6].
The first paper introduces CIMemories, a benchmark for whether a model controls information flow out of memory based on the task in front of it [2]. The construction matters: synthetic user profiles carrying more than 100 attributes each, paired with task contexts in which a given attribute is essential for some tasks and inappropriate for others [3]. That is the right shape for the problem, because no attribute is statically secret. Frontier models produced up to 69% attribute-level violations, meaning they revealed information inappropriately [4]. Against a profile of 100 attributes, that is on the order of 69 fields going somewhere they should not [14]. Models that violated less often did so at the cost of task utility [5].
The compounding is the part for the risk register. GPT-5's violation rate rose from 0.1% at a single task to 9.6% across 40 tasks, a 96-fold increase [6][15], and reached 25.1% when the same prompt was executed five times [7], roughly 2.6 times the 40-task figure [16]. The authors report models leaking different attributes for identical prompts, which they characterize as arbitrary and unstable behavior [8]. Read operationally: a clean single-run evaluation tells you very little about a fleet running the same workflow all week.
Privacy-conscious prompting did not resolve it. According to the paper, models overgeneralized, sharing everything or nothing rather than making nuanced, context-dependent decisions [9]. So the usual mitigation, a paragraph in the system prompt instructing the model to protect user data, is not a control. The authors' conclusion is that this needs contextually aware reasoning capabilities, not better prompting or scaling [13].
The second paper is more constructive and more limited. Its authors prompt models to reason explicitly about contextual integrity when deciding what to disclose, then use a reinforcement learning framework on a synthetic dataset of only 700 examples with diverse contexts and disclosure norms [10]. They report substantially reduced inappropriate disclosure while maintaining task performance across multiple model sizes and families [11], and say the improvements transfer to PrivacyLens, a benchmark with human annotations that evaluates privacy leakage in assistant actions and tool calls [12]. That is one set of self-reported results trained on synthetic data, not independent replication.
What to watch: whether any vendor shipping memory publishes accumulation curves and run-to-run variance instead of single-task scores [6][7]. Locally, the tests are cheap. Run the same sensitive workflow five times and diff what the agent disclosed each time, since identical prompts have been shown to leak different attributes [8]. Scope memory per task and per tenant rather than per user, and treat everything written to memory as data you have agreed to lose eventually. And if reducing leakage still costs task utility [5], that trade belongs to whoever will own the incident, not to the team demoing the assistant.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The CIMemories evaluation found that frontier models exhibit up to 69% attribute-level violations, meaning they leak information inappropriately.
In the CIMemories evaluation, lower violation rates often came at the cost of task utility.
A post titled "LLMs and Contextual Integrity" on Bruce Schneier's blog says he has been thinking about AI and integrity, including contextual integrity, and presents two papers on the topic.
CIMemories is a benchmark for evaluating whether LLMs appropriately control information flow from persistent memory based on task context; the paper notes that memory used for personalization and task performance introduces critical risks when sensitive information is revealed in inappropriate contexts.
CIMemories uses synthetic user profiles with over 100 attributes per user, paired with diverse task contexts in which each attribute may be essential for some tasks but inappropriate for others.
As usage increases from 1 to 40 tasks, GPT-5's violations rise from 0.1% to 9.6%.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Quantified but single-source and unverifiable
The cluster carries specific, falsifiable measurements (up to 69% attribute-level violations; GPT-5 0.1% to 9.6% to 25.1%) and a described benchmark methodology, which is more than assertion. But all of it reaches us through one aggregator post that reproduces abstracts only - no paper titles, authors, venues, links, model lists, or per-model tables - and there is no second publisher or replication. Evidence quality is therefore mid-band: precise numbers, thin provenance.
No deployment or usage evidence supplied
The supplied material contains only benchmark publication signals. Nothing in the source reports products shipping memory features, users affected, vendor remediation, incidents, or uptake of either the CIMemories benchmark or the CI-reasoning/RL method by labs or evaluation suites. Benchmark existence is not adoption, so this dimension cannot be scored without inventing facts.
Slightly overstated by worst-case framing
The framing leads with the maximum figure - up to 69% of attributes disclosed - which the abstract supports but which sits far above the same paper's typical measurements (GPT-5 at 0.1% in the single-task condition, 9.6% across 40 tasks). Derived restatements that convert 69% into an absolute count of leaked attributes push further than the source's 'over 100 attributes' floor allows. The underlying accumulation and instability findings are real and understated by nobody, so the gap is modest rather than severe; the source itself adds no promotional language.
Author self-reporting, no disclosed affiliations
Both quantitative results are reported by the teams that built the artifact being evaluated: the CIMemories group frames its own benchmark as revealing fundamental limitations, and the second team reports that its own RL method substantially reduces disclosure and transfers to PrivacyLens. That is a standard, moderate research incentive to favor one's contribution. Offsetting factors: the intermediary is a commentary blog with no commercial stake evident in the text, and no funding, employer, or vendor relationships are disclosed either way, so the read stays mid-band rather than high.
Moderate-low: precise numbers, one unverifiable source
Confidence is limited by cluster structure rather than by internal contradiction. A single publisher, abstracts reproduced without identifiers or links, no independent replication, and no adoption data mean the direction of the finding (memory-backed agents leak context-inappropriate attributes, and leakage compounds with usage and repetition) is more trustworthy than any specific figure. The mitigation half of the story is weakest, being entirely unquantified in the supplied text.
invest
Google says frontier models already know the facts they get wrong. That is a budget decision.1 distinct publisher
security
The nationalization argument is really a vendor-continuity memo1 distinct publisher
science
NIST's own logs show agents looking up the answers, making public benchmark scores soft evidence1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026