Build1 publisher2 min readPublished
Storing only user messages kept nearly all LongMemEval memory value at an eighth of the size
Aura Memory's developer kept almost all memory value on 120 LongMemEval questions by storing only the user's words, at an eighth of the size. Summaries in the same tests dropped facts and invented advice, so raw messages are the safer default.
The Engineer · Build desk

What happened
- Each experiment had its protocol and pass/fail thresholds written down and committed before the run, and runs that missed are reported as misses.
- Nearly all of the gap to storing everything came from one question type, 'what did the assistant tell me?', which user-only storage cannot answer by design.
- In one week of the developer's own 319 messages to coding assistants, 45% carried something durable, led by 64 facts about the setup and 34 preferences.
- With room for only 25% of LoCoMo turns, promoting often-retrieved records was the worst forgetting rule by 10 points, and the developer removed it from Aura.
- A 66k-weight network trained in seconds on a CPU from consequence labels beat time-based decay by 16 points on conversations it had not seen.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams designing agent memory can default to storing raw user messages and add assistant turns only when users will need the agent to recall its own answers.
- exposure A summarizing memory can write recommendations the assistant never made into a user's long-term record, where later sessions will retrieve them as history.
- capability A team that logs which memories changed an answer has enough data to train its own eviction model on a CPU, with no human labelling pass.
An eighth of the size is 87.5% less to store [2]. The test used the same reader and judge for every storage arm [2].
Summaries failed in two ways. "Summaries were a disaster: they keep the topic and drop the facts," Aura's developer wrote [5]. Half of them also described things the assistant "recommended", although the summarizer had been shown only the user's messages [6]. Putting them in the store alongside the raw text made retrieval worse. "Added on top of the raw words, they hurt, because they push real messages out of the top results," the developer wrote [7].
This is one developer's comparison, run while trying to get the developer's own product to decide what to keep [18]. The result describes LongMemEval's question mix. I think it carries over to a team whose users mostly ask an agent about things they said themselves. The author flags a wider limit. Benchmarks are built from questions about the past, so they cannot show whether real conversations contain anything worth keeping [8].
A week of labelled personal traffic answers that for one user. The four durable categories add up to 142 of 319 messages, or 44.5% [1]. Twelve times that week the author repeated something a model had lost to a new session or a context reset, and five of those repeats crossed projects [10]. The author concluded that memory has to be shared across projects and tools [19].
Storing raw messages still leaves the question of what to drop when space runs out. Crediting the single record whose removal broke an answer moved results by 1.5 points. The author called that noise, because a record that helped once is rarely needed again [12]. The small learned network did generalise across datasets. Trained only on LoCoMo, it scored 45% on LongMemEval, against 32% for time-based forgetting [14].
Two cheap signals failed. On user-assistant data, keeping the longest messages retained 98% assistant turns and 5% of what was actually needed [15]. Surprisal under a local language model scored an AUC of 0.55 on benchmark outcomes and 0.41 on the author's own messages [16]. Content helped. The network fed embeddings kept 45% of needed messages, against 6% for the same network fed only structure [17].
The method is what I would copy, more than any single figure. The author's working definition of value is the decision-theoretic one: "A memory is worth what the future answer gains from having it" [20].
What to watch
- A rerun of the storage comparison on a second benchmark, or with a reader and judge other than the ones used for the 120 LongMemEval questions.
- Whether the consequence-trained forgetter holds up on real-world signals the author named as candidates, such as a coding agent's exit codes.
- Whether Aura Memory ships the learned forgetter in place of time-based decay.