Skip to content

Build1 publisher3 min readPublished

Session isolation is not data isolation, and write-side memory defenses cannot see the difference

A read-only sensor for cross-tenant memory leaks ships with no tagged release, because its own backend cannot enumerate what the retriever exposed. That limitation is the story.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • CrossSessionMemoryGuard is a read-only sensor that detects unauthorized cross-session memory flow using three signals (provenance, content, and a write/read graph) over the chunks the engine exposes to a principal, never blocks anything, and documents its own limitations with versioned raw evidence in the repo.
  • Persistent agent memory is today defended mainly on the write side, against poisoning and manipulation.
  • The prior question that almost nobody monitors is whether this data should go out towards this principal.
  • The post cites recent work, identified as arXiv 2607.23444, 'Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents', as demonstrating a persistent memory extraction attack in practice against agents that operate isolated by session.
  • Session isolation is not the same as data isolation: if the retrieval engine crosses tenants (a broken filter, an aggressive consolidation, a relabeling), one session can silently read what another wrote.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer has published CrossSessionMemoryGuard, a read-only sensor that watches the chunks an agent's memory engine hands to one principal and asks whether another principal wrote them [s1c1]. It matters because persistent agent memory is currently defended almost entirely on the write side, against poisoning and manipulation, while the prior question -- should this data leave towards this principal -- goes largely unmonitored [s1c2][s1c3].

The post cites recent work it identifies as arXiv 2607.23444, "Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents", as a practical demonstration against agents that are isolated per session [s1c4]. The distinction the project draws is the useful one: session isolation is not data isolation, and if the retrieval engine crosses tenants through a broken filter, an aggressive consolidation, or a relabeling, one session can silently read what another wrote [s1c5]. Nothing in an integrity check on the write path is looking at that.

As evidence that the read side is uncovered, the author reports that a GitHub repository search using the five exact terms describing the problem returns zero results, while control searches return 4,632 repos for "agent persistent memory" and 435 for topic:llm-memory [s1c6][s1c7]. The raw JSON from those calls is versioned in the repo as github_gap_2026-08-17.txt rather than asserted in prose [s1c8]. Adjacent projects the author names -- OWASP Agent Memory Guard, memlineage, dent8 -- are described as covering write-side integrity and poisoning [s1c9].

The design is deliberately unambitious in scope: it observes, compares and alerts, never blocks, modifies or participates in authorization, sits outside the critical path, fails open, and carries a kill-switch via CSMG_DISABLED=1 [s1c10][s1c11]. Events carry a SHA-256 hash plus a minimal span, not the content [s1c12]. Attribution of who wrote what is taken from the schema layer, never from the observed view, because a retriever that crosses tenants must not be trusted to relabel ownership; the author says that contamination produced mass false positives in the benchmark, logged as KI-8 [s1c13].

The honest part is why there is no tagged release yet [s1c14]. The Engram adapter reads the schema layer rather than the real search path because the author verified that engram search cannot list everything visible: an empty query errors with "search query is required", wildcards and FTS escapes return zero results, and there is a hard cap of 20 results per query that persists at --limit 1000 and --limit 5000, against a default of 10 [s1c15][s1c16]. Coverage therefore becomes a function of the query set: 15 broad queries covered 87.9% of real observations, 80 of 91, 23 queries reached 96.7%, and 100% arrived only with queries tailored to the rows that were missing [s1c17]. At 15 queries, 11 of 91 observations were invisible [s1d1]. An independent audit reproduced the cap with the engine's official binary and 30 test observations: 30 stored, 20 returned, two thirds of what was there [s1c18][s1d2].

That is the structural point for anyone building read-side monitoring. If your ground truth for what the engine exposed is itself a set of queries, a leak nobody thought to query for does not exist. The author also flags that mem0, langmem, zep and letta implement the real read but not schema_chunks(), so a broken tenant filter in those adapters would contaminate attribution the way bug t1 did in the SQLite and JSONL paths before it was fixed [s1c19], and that composite collusion cases -- fragments of a secret each below the similarity threshold, scenario t4 -- are not detected in the MVP and are reported as declared limitation AC7 rather than a pass [s1c20].

Watch whether external adapters get a schema layer of their own, and whether memory engines start shipping an enumeration path that a monitor can trust.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories