Build1 distinct publisher3 min readUpdated
A preprint pairs a 'do you know this preference' test with a 'now act on it' test on the same item, and reports a large gap between the two. Health and therapy preferences fare worst.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The useful move in this design is reading the same stored fact twice. The Know test asks the agent to state a specific preference; the Act test drops it into a scenario where that preference should change the answer [6]. Scoring both on the same item produces four outcomes instead of two, and a retrieval benchmark can only separate remembered from not-remembered, which means the cell where the agent recites the constraint and then violates it is scored as a pass [17].
The paper's illustration is a user who reveals a peanut allergy while asking for help polishing an email. The memory module extracts it correctly. In a later session the agent recommends pad thai with crushed peanuts [14]. That failure is invisible to a recall metric, because recall is exactly the part that worked.
This matters more than a lab curiosity because of how preferences actually arrive. The authors cite a large-scale analysis of ChatGPT usage finding that most users still treat the model as a tool for refining emails, translating messages, or debugging code, so preferences surface implicitly inside task content rather than as declarations [13]. The benchmark embeds its 1,000 preferences at three levels of expression strength for that reason [3]. Weakly expressed preferences are the normal case in production, not the hard tail.
Memory architectures do help. The reported result is that they narrow the Know-Act gap, but that utilization stays weakest on health and therapy preferences [5], which is the inverse of where you would want the residual error. Part of the mechanism is that a memory layer is not a pure addition: distilling history into condensed context introduces its own failure points, semantic mismatch at retrieval among them [11]. And the field already knows models skip relevant information sitting fully in context [10]; adding a retrieval stage in front of that does not remove it.
Two caveats on how far this travels. What we have is one anonymised preprint with code and data behind an anonymous repository [16], and no per-system breakdown, so it does not license a ranking of ChatGPT's persistent memory against Claude's, or Mem0 against Letta/MemGPT, Zep, or HippoRAG, all of which the paper names as the systems being equipped with this capability [12]. Sixteen systems spread over five memory architectures averages 3.2 systems per architecture [18], which is thin for any claim about a family rather than a build.
The second caveat is the authors' own, and it cuts both ways: a good behavioural response may reflect coincidence rather than memory [15]. Neither half of the pair is a scorecard alone. That is the argument for pairing, and it is also why the existing evidence base is weaker than its numbers look. The benchmarks these systems post strong results on test factual recall, multi-hop reasoning, and long-range understanding [9], while personalization and agentic-memory work has measured recall or behaviour but not both on the same stored item [7], mostly in static long-context setups rather than the incremental pipelines a long-running agent actually uses [8]. So the closest thing to evidence that a memory layer changes what an agent does, rather than what it can repeat, is a test design that until now nobody was running.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
While memory architectures reduce the Know-Act gap, utilization remains especially weak for health and therapy-related preferences, where failures to act carry the greatest real-world stakes, according to the paper.
The paper introduces a decoupled evaluation paradigm that administers paired Know and Act tests to the same user preference.
The authors run large-scale experiments across 16 systems and five memory architectures, evaluating 1,000 preferences embedded at three levels of expression strength.
Results show a large gap between Know and Act outcomes: agents often pass the recall test for a user preference but fail to reflect that same preference in the paired behavioural scenario.
The Know test directly asks whether the agent can recall a specific preference; the Act test presents a natural scenario in which the preference should matter.
The authors note that testing only behaviour leaves open whether a good response reflects genuine memory or coincidence, and testing only recall leaves open whether the information influenced behaviour.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but single-source and unreplicated
The methodology is specified at unusual granularity for a cluster this small - paired Know/Act tests on the same preference, 16 systems across five architecture families, 50 personas, 1,000 preferences, three expression levels, an incremental memory-construction protocol, and a named per-system utilization figure. That specificity, plus a stated code-and-data release, lifts evidence above assertion level. It is capped by structural limits: one publisher, an un-peer-reviewed preprint, all numbers self-reported by the authors, code behind an anonymous review repository, and supplied text truncated before the per-preference-type results that underpin the health/therapy conclusion.
Artifact published, no third-party uptake observed
Observable adoption in the supplied material is limited to the authors' own publication: a preprint plus a stated anonymous code-and-data repository, and one internally run benchmark sweep. There are no citations, forks, downstream evaluations, vendor responses, or production deployments of the KnowAct harness in the cluster. The evaluated memory products (ChatGPT and Claude memory, Mem0, Letta/MemGPT, Zep, HippoRAG) are named as test subjects, which evidences their existence as a landscape, not adoption of this work.
Mildly overstated relative to verification
The paper's own language is restrained and scoped - it reports a gap, credits memory architectures with narrowing it, and flags the health/therapy weakness as the residual problem - and the cluster framing tracks those findings rather than inflating them. The small positive gap reflects that every claim is self-reported in an unreplicated preprint, that the sharpest assertion (health and therapy worst) is not backed by visible per-category numbers in the supplied text, and that per-architecture conclusions rest on roughly three systems each. Nothing in the cluster contradicts the paper, so the gap stays near alignment.
Standard academic novelty incentive, no disclosed vendor stake
The visible incentive is scholarly: the paper positions itself against prior work that tests recall or behaviour but not both, and against static long-context evaluation, so demonstrating a large and previously unmeasured gap is what establishes novelty. The anonymous 4open.science link indicates a blind review submission, consistent with an academic venue rather than a commercial launch. No funding source, vendor sponsorship, or commercial affiliation is disclosed in the supplied text, and the paper both credits existing memory architectures with large gains and criticises them, which tempers a pure adversarial-framing incentive.
Moderate-low: plausible and specific, but unverified
Confidence is held down by the cluster's structure rather than by any internal contradiction. A single arXiv preprint, all figures self-reported, no peer review, no independent replication, code behind anonymisation, and a truncated results section mean the direction of the finding is credible - it is consistent with the well-documented knowledge-utilization problem the paper cites - while the magnitudes should be treated as provisional. The design detail and paired-test logic are strong enough to act on as a testing hypothesis, not as a settled measurement.
build
Session isolation is not data isolation, and write-side memory defenses cannot see the difference1 distinct publisher
build
The one signal agent memory learns from is the one production traffic almost never sends1 distinct publisher
science
GJ 523b gives 'Mega-Earth' a number: 23 Earth masses inside 2.5 Earth radii1 distinct publisher
build
Multi-agent LLM gains largely vanish once the thinking-token budget is held constant1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026