Skip to content

Build1 publisher3 min readPublished

past.dev bets agent builders will pay for memory that knows when a fact went stale

past.dev has opened a memory API for AI agents and reports 85.03% on the 10-million-token split of the BEAM benchmark. Rival scores came from differently configured runs and no customer is named, so the case for a separate memory layer rests on past.dev's own figures.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying past.dev bets agent builders will pay for memory that knows when a fact went stale
Generated illustration

What happened

  • Developers send it timestamped emails, meeting transcripts, messages and documents, then ask questions through a recall request that carries the requester's identity.
  • Answers are meant to return the current fact, what it replaced and dated source material, filtered to what the requester may see, including for questions about an earlier date.
  • The API supports MCP and self-serve signup, and enterprise customers can self-host it or use dedicated regions.
  • A past.dev blog post dated September 30 already presented the API as available, six days before the October 6 PR Newswire release that formally introduced it.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction Depending on which of its own launch artifacts a buyer reads, past.dev's lead at 100,000 tokens is either 5.88 points or 15.18, so the claimed margin cannot be taken at face value.
  • cost A buyer who wants a like-for-like comparison has to pay for a rerun that scores every system with one model and one judge, on its own data.
  • capability Agents built on it can be asked what was true on a past date and what superseded it, a question plain search over messages cannot settle.
  • exposure Putting team history and access rights in one shared store means a wrong identity or a stale permission at recall time affects every agent that calls it.

Mehdi Djabri and Alex Ronse built past.dev as the memory behind their own agents for product work, engineering, email and meetings [1]. The store keeps each fact with its history, its source and its audience [3]. Working out what replaced what has to start from the timestamp attached to each item at ingest [4]. The founders list late-arriving information among the problems they ran into [14]. A meeting transcript filed under its upload time instead of its meeting time would let an old decision overwrite a newer one.

The access filter has a similar dependency. It acts on the identity the caller passes with each recall request [4] and on whatever rights the store holds for that identity. I would ask the vendor two things first: how those rights get into the store, and how quickly a revoked right reaches it. Point-in-time queries add a third: whether a question about an earlier date applies the permissions in force then or now.

The benchmark page describes a single run on September 29, 2026, over complete BEAM splits [11]. The 100,000-token split has 400 questions, the 500,000- and 1-million-token splits have 700 apiece, and the 10-million-token split has 200, for 2,000 in total [15]. The headline 10-million result comes from the smallest split [11]. The scores are not whole-question counts: 92.08% of 400 is 368.32 [16]. So the grading is finer than right or wrong. The page also says rival systems were run with different models, judges and configurations [13]. Exabase M-1's 68.0% is the next result at 10 million tokens, 17.03 points behind [12][17]. For that gap to carry over, one model and one judge would have to score every system. BEAM is a benchmark of long conversations and memory tasks published at ICLR 2026 [10]. A team's email and meetings, with an audience attached to every item, is a different workload. I would not treat the score as evidence that the permission filter works.

For a product sold on knowing which version of a fact is current, the launch materials disagree with themselves about the baseline. The release table lists the previous best at 86.2%, 80.1% and 79.1% for 100,000, 500,000 and 1 million tokens [18]. Its accompanying graphic shows 76.9% and 71.1% for the first two [19]. Measured against the table, past.dev's 92.08% at 100,000 tokens [9] leads by 5.88 points. Against the graphic, the lead is 15.18 [20]. At 500,000 tokens the lead is 9.53 or 18.53 points [21].

Runtime Wire describes the commercial bet this way: teams running agents across internal systems will want one auditable memory layer, so that each application does not have to build and maintain its own history [23]. The release says the system ran in production for large teams across very large volumes of data [22]. past.dev has not quantified deployments, data volumes, revenue or usage [7]. Its About page says early customers get help with evaluation, deployment and production testing [8].

What to watch

  • A controlled BEAM rerun of past.dev and its listed rivals under one model and one judge, or a corrected baseline that settles 86.2% against 76.9% at 100,000 tokens.
  • A named production customer with published data volumes, the first outside check on the release's claim of large-team production use.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories