A runnable two-agent example moves every refusal into the catalog that vends credentials, because a key covers a storage prefix and there are no rows at the point of enforcement. The author works on that catalog.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+12
- Incentives82
- Confidence55
Mem0, Zep, Cognee, Letta, Supermemory and Mnemoverse all ship a knowledge graph, and each one's docs mean something different by a node. Read-time behaviour is the part a buyer can settle from the documentation.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap−8
- Incentives78
- Confidence52
A preprint from Ariel University reports laundering attack success of up to 68% against existing agent-memory defenses, and argues in machine-checked TLA+ that authority has to be bound to an item's origin at the moment it is written.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+35
- Incentives62
- Confidence47
Microsoft Security counted 50 memory poisoning attempts across 31 companies in 60 days, according to a security engineer's account of the report. The scanners most teams already run check inputs at arrival and never read the memory file.
Reality
- Evidence28
- Adoption20
- Hype gap+38
- Incentives82
- Confidence32
A developer who builds one of the six persistent memory systems he checked walks through how each can stop sharing while Cursor and Claude Code both report a healthy connection. The split happens at the account, and nothing logs it.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+8
- Incentives72
- Confidence52
The MCP schema defaults destructiveHint to true, but only where readOnlyHint is false. A ChatGPT app directory scanner asked for the field anyway on four read-only tools, and the tools came out of it behaving exactly as before.
Reality
- Evidence74
- Adoption22
- Hype gap−8
- Incentives40
- Confidence66
Mem0-style memory replaces embed-and-store with extraction, write-time retrieval, an ADD/UPDATE/DELETE/NOOP decision and consolidation. The similarity scores in the dev.to writeup show why no threshold can make that call for you.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+22
- Incentives30
- Confidence55
Mithil Vakde's from-scratch transformer landed one point behind TRM on the public eval. His own ablations put 20 of those 44 points on two representation choices rather than on any amount of compute.
Reality
- Evidence44
- Adoption18
- Hype gap+16
- Incentives72
- Confidence41
Jordy Zomer built a Datalog engine so an agent maintains what it currently knows instead of searching its own transcript, and his own benchmark runs put the weak link in the model that writes the facts.
Reality
- Evidence46
- Adoption12
- Hype gap−8
- Incentives28
- Confidence54
A comparison of Mem0, Zep and LangChain's memory classes puts the split at adjudication rather than storage. Mem0 pays two model calls per message to get it.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives
- Insufficient
- Confidence38
A preprint pairs a 'do you know this preference' test with a 'now act on it' test on the same item, and reports a large gap between the two. Health and therapy preferences fare worst.
Reality
- Evidence52
- Adoption14
- Hype gap+8
- Incentives55
- Confidence48
A vendor audited 359,388 edges in its own memory store and found feedback-shaped graphs only where its benchmark harness supplied the feedback. Live tenants got one reinforcement per edge.
Reality
- Evidence48
- Adoption17
- Hype gap+12
- Incentives74
- Confidence44
A read-only sensor for cross-tenant memory leaks ships with no tagged release, because its own backend cannot enumerate what the retriever exposed. That limitation is the story.
Reality
- Evidence41
- Adoption9
- Hype gap−14
- Incentives63
- Confidence40
A runnable Mem0-backed wrapper blocks an agent's repeat tool call before it burns the request. The pattern is sound; the default policy will block you forever.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+12
- Incentives35
- Confidence52
Mads Thines has shipped an open-source memory layer for coding agents that starts on local disk. The design bet is that a record of past mistakes should be inspectable and portable.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+14
- Incentives68
- Confidence40