Build1 distinct publisher2 min readUpdated
A new benchmark penalises agents for acting on facts the user already revoked. The dedicated memory layers barely beat the plain models on it.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A retrieval score cannot tell you which of two contradictory memories an agent actually used. If a user says one thing in March and the opposite in June, a system that stores both and pulls both back is credited with recall, and there is no column for which version it then acted on. FAMA is built to charge for exactly that: it penalises reliance on obsolete or invalidated memory instead of rewarding the presence of a hit [4]. The paper's account of why this went unmeasured is the more useful part. Retrieval-centric evaluation implicitly assumes stored information stays valid forever, while real interaction is non-stationary, with user facts updated, corrected or withdrawn over time [13].
The depth figures reward a little arithmetic. In LoCoMo, 94% of evaluation questions need grounding evidence from no more than two previous sessions [7]; the authors report the same pattern for 85% of LongMemEval questions [8]. That leaves at most 6% and 15% respectively reaching further back than two sessions [16][17]. Whatever "long-term" means in those suites, it is mostly one hop.
On mutation the ceiling is lower still. LongMemEval includes knowledge-update operations but caps them at two sessions before evaluation [10], and PersonaMem handles updates across no more than three [11], so the deepest revision chain any of these benchmarks asks a model to reconcile is three sessions long [18]. A production assistant that a user has corrected four times across five months is outside the tested envelope entirely.
The word carrying the weight in the results is "frequent". The abstract and introduction describe frequent reuse of invalid memories without putting a number on it [19], so nobody can yet rank the six systems by how often they act on a dead fact, which is the number a buyer would want. What makes the finding checkable rather than assertion is the build: the authors say data quality was controlled with automated memory-grounding checks plus human evaluation [3], and that code and data are published [15].
If it survives replication, the engineering consequence is not a larger index. It is a write path that can mark a fact dead, with supersession and a timestamp on the moment the user changed their mind. Statelessness was the problem the memory layer was sold against, since the cache holding short-term context is discarded when an interaction ends [14]. Retention without invalidation is a different fault with a similar symptom: an assistant answering confidently from information its user already retracted.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Memora is a long-term memory benchmark spanning weeks to months long user conversations, evaluating three memory-grounded tasks: remembering, reasoning and recommending.
To ensure data quality, the authors employ automated memory-grounding checks and human evaluation.
The paper introduces Forgetting-Aware Memory Accuracy (FAMA), a metric that penalizes reliance on obsolete or invalidated memory when evaluating long-term memory.
Evaluations of four LLMs and six memory agents reveal frequent reuse of invalid memories and failures to reconcile evolving memories.
Memory agents offer marginal improvements, exposing shortcomings in long-term memory for personalized agents.
In LoCoMo, 94% of the evaluation questions require grounding evidence from no more than two previous sessions.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single self-published preprint with quantified critique but no released scores
The cluster contains exactly one source, the authors' own arXiv preprint, with no peer review or independent replication. Its critique of prior benchmarks is specific and checkable (94% / 85% session-depth figures, Table 1 consolidation average, two- and three-session mutation caps), and code and data are claimed public, which raises verifiability. But the headline result -- six memory agents failing on forgetting -- arrives without per-system numbers or system names in the supplied text, so the central finding cannot be inspected from this material.
First-party release only
Observed adoption is limited to the authors' own publication: a preprint plus a stated public repository, and one internal evaluation run over four LLMs and six memory agents. The supplied material shows no third-party use of Memora or FAMA, no leaderboard, no vendor response, and no production deployment, so adoption is real but confined to the originating team.
Framing runs slightly ahead of disclosed numbers
The cluster framing -- six memory agents fail, dedicated memory layers barely beat plain models -- tracks the abstract's own wording, and the benchmark-design critique is well quantified. The overstatement is modest and definitional: 'fail' and 'marginal' are the authors' qualitative characterisations, delivered without failure rates, per-system scores or system names, and scored by a metric the same authors introduced. That leaves the strongest conclusions resting on unpublished-in-this-source numbers.
Authors grade rivals with the benchmark and metric they authored
The paper both defines the evaluation (Memora and FAMA) and reports that four LLMs and six competing memory agents fail on it, with the code hosted under an industry organisation's GitHub account (geniesinc). That is a structural interest in the finding that existing memory layers underperform. Mitigating factors are the public code-and-data release and the disclosed data-quality process, which expose the construction to outside checking; no funding, vendor relationship or competitive positioning is disclosed in the supplied text.
Design claims solid, headline verdict thinly documented
Confidence is moderate-low. The benchmark-design and prior-benchmark-depth claims are stated precisely by the source and are checkable, so those hold up. The consequential claim -- that six memory agents fail and add little over base models -- is single-sourced, author-scored, unreplicated and unquantified in the supplied excerpt, which is truncated mid-sentence before results. No second publisher is present to corroborate or contest.
build
AgentCL: if the task stream is not controlled, agent memory gains prove nothing1 distinct publisher
build
The failure mode your dashboard cannot see: 43% of agent errors are wrong-shaped output at HTTP 2001 distinct publisher
build
Solar Pro 4 turns model routing into a procurement decision, not a research one1 distinct publisher
invest
Google's 180-config sweep: extra agents cut sequential-task scores by 39 to 70%1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026