Published · yesterdayProduct2 min read
Memory is graded on two sessions: 94% of LoCoMo questions need no more than that
The figure comes from Memora's audit of existing long-term memory benchmarks, and it explains why stored preferences can both fail to bind and steer an agent into a single lane.
Not a builder's beat, but builders have a standing stake in it.See today for builders

What happened
- The Memora paper reports that in LoCoMo, 94% of the evaluation questions require grounding evidence from no more than two previous sessions.
- The same authors observe the same pattern for 85% of the evaluation questions in LongMemEval.
- Memora's Table 1 shows that average memory consolidation across existing benchmarks is approximately one session.
- Memora is a long-term memory benchmark spanning weeks to months long user conversations, evaluating three memory-grounded tasks: remembering, reasoning and recommending.
- The paper introduces Forgetting-Aware Memory Accuracy (FAMA), a metric that penalizes reliance on obsolete or invalidated memory when evaluating long-term memory.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
Stored text is the whole mechanism, and it does not distinguish between a hard constraint and a passing mood. In the Ben's Bites account of building a personal agent, Codex advised folding build preferences into the main instruction file, which the author reads as a product of the coding variant's own system prompt telling it that it is a coding agent [14]. The same property produced the behaviour he objected to: preferences he had asked the agent to log came back at him as reasons to stay in the lane he had already described, when what he wanted was a new direction [12]. The Know/Act authors note that most users never state preferences outright, so what a memory module captures is often whatever was implied in a request to polish an email [11].
That is why the benchmark figure matters more than it looks. Memora itself covers conversations spanning weeks to months and scores remembering, reasoning and recommending [4]; the prior benchmarks its authors tabulate concentrate 85% of LongMemEval's questions and 94% of LoCoMo's inside two sessions [2][1]. The FAMA metric exists because ten systems, four base LLMs and six memory agents, reused invalid memories and failed to reconcile revisions, with the memory agents adding only marginal improvement [6][7][5].
So the field now has a penalty for memory that is out of date, and a paired test for memory that is present but unused [5][9]. What neither scores is the case the practitioner hit: memory that is present, current, correctly applied, and used to agree with him [16]. An agent that never contradicts its own stored profile passes both.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The Memora paper reports that in LoCoMo, 94% of the evaluation questions require grounding evidence from no more than two previous sessions.
- [2]
The same authors observe the same pattern for 85% of the evaluation questions in LongMemEval.
- [3]
Memora's Table 1 shows that average memory consolidation across existing benchmarks is approximately one session.
ReportedView cited source - [4]
Memora is a long-term memory benchmark spanning weeks to months long user conversations, evaluating three memory-grounded tasks: remembering, reasoning and recommending.
ReportedView cited source - [5]
The paper introduces Forgetting-Aware Memory Accuracy (FAMA), a metric that penalizes reliance on obsolete or invalidated memory when evaluating long-term memory.
ReportedView cited source - [6]
Evaluations of four LLMs and six memory agents reveal frequent reuse of invalid memories and failures to reconcile evolving memories, with memory agents offering marginal improvements.
ReportedView cited source
Sources & coverage · 2 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- bensbites.comyesterdayHow I built this - Ben's Bites



