Skip to content

benchmark

Mem2ActBench

Evaluation benchmark for memory-grounded actions by AI agents, with a taxonomy of memory-related error types.

Current clusters

build1 publisher

Crystals push agent memory into the hook that runs before each tool call

Crystals, a memory design written up on dev.to, deliver notes to an agent just before a matching tool call runs, from a hook firing about 300 times a day. Its most useful finding is a matched note that the token budget cuts before the model sees it while the logs still count a hit.

Publishers:dev.to

Reality

Evidence30
Adoption5
Hype gap+5
Incentives
Insufficient
Confidence35