Skip to content

Build1 publisher3 min readPublished

The bug in agent memory is not volume, it is that everything recalled has equal authority

A dev.to post argues coding agents fail because abandoned decisions and standing conventions come back with identical confidence. The proposed fix is a schema with explicit precedence.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Teams that wire long-term memory into a coding agent hit the same wall about three weeks in: the agent remembers the abandoned migration, the replaced library and the convention one person suggested once in a thread, recalling all of it with the same flat confidence, so the team ends up debugging memory instead of code.
  • The instinct to store less or store better is the wrong axis; the problem is not how much the agent remembers but that everything it remembers has the same authority.
  • The common framing that separates the current conversation from facts kept forever is a real distinction but does not help, because it says nothing about what the agent should do when two remembered things disagree.
  • Evidence is what happened: the agent tried a fix and it failed, a user rewrote a function, a test went red. Evidence is cheap to produce, accumulates fast, and any single piece of it can be wrong or unrepresentative.
  • Policy is what should happen: use pnpm, tests mirror the src layout, never touch the legacy billing module. Policy is expensive to produce because a human usually decides it, and it should be hard to change by accident.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A dev.to post on coding-agent memory makes a claim worth taking seriously: teams that wire long-term memory into an agent hit the same wall roughly three weeks in, when the agent starts recalling the migration you abandoned, the library you replaced and the convention one person floated once in a thread, all with the same flat confidence [s1c1]. The consequence is that you end up debugging memory instead of code, and the usual responses (store less, store better) are aimed at the wrong axis, because the defect is authority, not volume [s1c1][s1c2]. The framing most systems ship with separates the current conversation from facts you keep forever. The author's objection is narrow and correct: that split says nothing about what the agent should do when two remembered things disagree [s1c3]. The proposed cut is evidence versus policy. Evidence is what happened: a fix was tried and failed, a user rewrote a function, a test went red. It is cheap to produce, it accumulates fast, and any single piece of it can be wrong or unrepresentative [s1c4]. Policy is what should happen: use pnpm, tests mirror the src layout, never touch the legacy billing module. It is expensive because a human usually decides it, and it should be hard to change by accident [s1c5]. Hold them apart and two familiar failure modes get names. An agent that promotes one observation to policy overfits to a single incident; an agent that demotes a merged architecture decision to evidence keeps relitigating it [s1c6]. Most systems put both in one undifferentiated bucket of "things we know," which is why their behaviour feels unpredictable [s1c7]. Operationally that becomes three layers. Shared project truth holds ADRs, API contracts, naming conventions and the deployment runbook, versioned and source-linked, with the same copy for every agent on the project [s1c8]. Role memory holds heuristics that belong to a job rather than a project: what a frontend reviewer checks, how the QA pass is structured, which failure modes a migration tends to hit. The author says this is the layer most systems skip, which is why teams re-teach review standards to every new session [s1c9]. Episodes hold what was attempted, what failed and what feedback followed; they grow fastest and rot fastest [s1c10]. The ranking is not importance but how easily something should change: episodes written constantly, role memory slowly, shared truth only when a human decides [s1c11]. The part that generalises beyond agents is the time model. If a codebase moved from Redux to Zustand and the store keeps one time axis, the old note is overwritten or decays, and the agent can no longer explain why a component written in March looks the way it does, because the fact that explains that code is gone [s1c12]. "Recorded at" and "was true from/until" are different questions: "use Redux" was true from January to June and recorded in February, which makes it closed rather than wrong [s1c13]. Note the gap: the record lands a month after the fact became true, so anything keyed only on recorded-at misdates it and has nowhere to put the June closure [1]. Keep closed facts with their validity windows and the agent can explain old code, flag a pattern as belonging to a superseded era, and stop rewriting history [s1c14]. Replacing rather than closing is, the author argues, the most common destructive operation in agent memory, and it stays invisible until someone asks about the past [s1c15]. None of this is new engineering. Separating when a fact was true from when the system learned it is bitemporal modelling, standardised in SQL:2011 as application-time and system-versioned period tables [s1c16]. What to watch is the promotion path, which is where these schemas usually leak.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories