Memory handoffs cut replaced agent context by 98.65% across 10,241 completed runs, according to an audit published on dev.to. For agent builders, the open risk now sits in whether the agent recalls the right stored text, since lossy summaries are no longer the failure point.
Reality
- Evidence55
- Adoption15
- Hype gap+10
- Incentives
- Insufficient
- Confidence45
Throughline's on-call agent returns a three-way coverage verdict with every memory recall, so a timed-out search cannot pass as "no prior incidents". Its author hit the same error-as-empty trap in CockroachDB's managed MCP server, where failures come back as HTTP 200.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap−10
- Incentives30
- Confidence40
Every 10 minutes an unattended job summarizes Claude Code and Codex logs into an Obsidian vault and pushes the commit, so the pipeline treats the summarizer's own JSON as text that may carry pasted keys or injected instructions.
Reality
- Evidence58
- Adoption6
- Hype gap+10
- Incentives20
- Confidence55
A runnable two-agent example moves every refusal into the catalog that vends credentials, because a key covers a storage prefix and there are no rows at the point of enforcement. The author works on that catalog.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+12
- Incentives82
- Confidence55
A dev.to writeup by nasiko_labs describes a semantic memory tier whose first version inserted a row per session, so one agent could read "cautious" and "aggressive" at once. The rewrite keys each belief by fact type and archives the superseded version.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence45
A Strands hook recorder writes each agent step into Neo4j through Neo4j Labs' agent-memory SDK. The audit query that finds every decision touching a bad source depends on a dict lookup keyed by tool name.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+20
- Incentives65
- Confidence55
Both terminal agents recalled planted facts inside a single project across three two-session tests on identical Node repos. Only Grok Build carried a stated convention into an unrelated repo, at about a third of the reported cost.
Reality
- Evidence58
- Adoption25
- Hype gap+14
- Incentives40
- Confidence62
Dhravya Shah spent three days probing the personal agent from the outside and reports git-tracked Markdown searched by keyword, with a background pass he clocked taking 23 hours and 16 minutes to commit one preference.
Reality
- Evidence42
- Adoption24
- Hype gap+28
- Incentives76
- Confidence58
A dev.to walkthrough of Claude Code compaction says the system prompt and the last 10 to 15 turns survive while the messages between them are summarized and deleted without an error. Constraints you need later belong in a file.
Reality
- Evidence20
- Adoption
- Insufficient
- Hype gap+35
- Incentives45
- Confidence60
An arXiv paper names the failure behavioral state decay and runs a side-car memory agent next to an unmodified action agent, reporting gains of 8.3 and 6.8 points of pass@1 on two long-horizon benchmarks.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+18
- Incentives60
- Confidence57
In a dev.to account of self-reinforcing memory loops, an agent's own hedged inference is captured as a flat fact, retrieved a week later as context, and generalised until the store recommends replacing the cache layer.
Reality
- Evidence33
- Adoption
- Insufficient
- Hype gap+20
- Incentives40
- Confidence50
Ramsauer and colleagues proved the equality in 2020, three years after attention was published. It holds for a single retrieval step inside a forward pass, and the capacity figures quoted alongside it come from a different model.
Reality
- Evidence74
- Adoption
- Insufficient
- Hype gap−8
- Incentives62
- Confidence66
Mem0, Zep, Cognee, Letta, Supermemory and Mnemoverse all ship a knowledge graph, and each one's docs mean something different by a node. Read-time behaviour is the part a buyer can settle from the documentation.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap−8
- Incentives78
- Confidence52
The opencode-agent-memory plugin gives an OpenCode agent scoped Markdown blocks it rewrites through three tools, plus an opt-in journal it can search locally but never revise. The maintainer documents the failure modes himself.
Reality
- Evidence34
- Adoption14
- Hype gap−12
- Incentives38
- Confidence52
Turing Post talked through agent feedback with Salesforce's chief AI scientist at Dreamforce. A correction lands in one of four places: the context, persistent memory, the software around the model, or the weights.
Publishers:turingpost.com
Reality
- Evidence46
- Adoption20
- Hype gap+12
- Incentives62
- Confidence48
IBM Research's ALTK-Evolve distils an agent's own trajectories into scored guidelines and injects the top five at inference time, and a companion post puts a number on the reliability an average success rate hides.
Reality
- Evidence45
- Adoption20
- Hype gap+15
- Incentives85
- Confidence55
Serge Kernbach's dev.to build notes put a 9B to 27B assistant on two RTX 4070s totalling 24 GB and argue that integration beats raw model quality. The figures he publishes are power draw and PCIe bandwidth.
Reality
- Evidence32
- Adoption12
- Hype gap+30
- Incentives22
- Confidence42
Vendor pages for Claude Code, Cursor, Codex, VS Code and Windsurf, re-checked on 2026-09-12, show that all five resume a previous session. The complaint about agents forgetting points at the other three questions.
Reality
- Evidence80
- Adoption
- Insufficient
- Hype gap+10
- Incentives45
- Confidence68
A developer instrumented 89 coding sessions to test whether rules the agent had already read changed what it did. Enforcement only started working once the rule was rewritten into something a hook could see.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap−10
- Incentives30
- Confidence55
An Apache-2.0 memory layer for coding agents returns a candidate only when several signals agree, and it labels every answer STRONG, WEAK or MISS so the calling agent has to branch on confidence before reusing an old fix.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence52
Earlier coverage
- Read-only environment probing moves a 40-question agent memory benchmark by 1.2 questions
Build · September 16, 2026 · 1 publisher
- Unmarked supersession lets embedding similarity pick which decision an agent believes
Build · September 15, 2026 · 1 publisher
- myc's PreCompact hook writes the session to disk before the summary drops the reason
Build · September 15, 2026 · 1 publisher
- Running every AppWorld task five times drops a ReAct agent from 77% to 53%
Build · September 14, 2026 · 1 publisher
- A key minted in a second console account returns an empty MCP memory store without an error
Build · September 13, 2026 · 1 publisher
- "Infinite context" is not a spec: a four-task harness for testing agent memory
Build · August 16, 2026 · 1 publisher
- Numbering six retrieved chunks turned one handbook into three agreeing sources
Build · September 13, 2026 · 1 publisher
- The server injects up to five project notes before the agent takes its first turn
Build · September 11, 2026 · 1 publisher
- Cursor's plugin.json puts agent tool scope where a code reviewer can diff it
Build · September 11, 2026 · 1 publisher
- Putting the name inside a MERGE pattern forks one person into two nodes
Build · September 11, 2026 · 1 publisher
- Claude's memory list hands back 20 full items per call against a 10,000-memory store
Build · September 10, 2026 · 1 publisher
- Claude Code's auto memory relearned the same Chrome fix in three separate repositories
Build · September 10, 2026 · 1 publisher
- Adjudicating every extracted fact against its five nearest memories costs one LLM call apiece
Build · September 6, 2026 · 1 publisher
- Disproving one pointer in Lemmalog retracts every conclusion that rested on it
Build · August 28, 2026 · 1 publisher
- The cheapest retrieval win this week sat at ingest, not in the agent loop
Build · August 24, 2026 · 1 publisher
- Coding agents cost $4,125 a month because 73% of it is context you already sent
Build · August 23, 2026 · 1 publisher
- The bug is the tutorial's first line: "set up your vector database"
Build · August 23, 2026 · 1 publisher
- Agent memory rots by accumulation, and the missing primitive is a supersession key
Build · August 22, 2026 · 1 publisher
- The one signal agent memory learns from is the one production traffic almost never sends
Build · August 21, 2026 · 1 publisher
- LoreKit puts agent memory in Markdown files you can grep, not a vendor's database
Build · August 16, 2026 · 1 publisher