Skip to content

Product1 publisher3 min readPublished

Timescale's agent-run memory sprint turned on a human checking what each metric counted

Timescale says Claude wrote nearly all the code and ran all 41 experiments that lifted its LoCoMo memory score from 0.392 to 0.666 F1 in six days. By the team's own account, the calls that decided the result came from a person checking whether each metric measured what it claimed to.

The Product Desk · Product desk

Illustration accompanying Timescale's agent-run memory sprint turned on a human checking what each metric counted

What happened

  • Of the 41 experiments, 17 were adopted and 24 reverted, over 108 eval runs and 85 commits.
  • The eval framework had been passing question categories into the prompt, giving temporal questions an "answer with a date" hint that production would never supply.
  • A recall gain from interleaving extracted facts with dialogue turns was an artifact, because the metric counted turn IDs and facts have none.
  • An ablation showed the Haiku-based fact extraction step added zero F1, and cutting it made ingestion 50% faster while improving retrieval recall.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Teams putting an agent on an experiment loop have to staff the reviewer seat with someone able to challenge the metric; in this run that review set the pace while the agent handled the code.
  • exposure Benchmark scores from agent-run loops can overstate production results whenever the harness hands the agent an input production lacks, since the agent will use it silently.
  • contradiction The post's summary credits all four decisive calls to measurement checks, but its own detail counts three, with the fourth a judgement about which hypothesis to try first.
  • cost Complex pipeline steps that nobody ablates keep costing ingestion time and adding search noise; this team found its dead step only when a human asked for the direct test.

On Timescale's harness, one eval run took five to seven minutes [6]. The team wrote that in that window a person could form a hypothesis, make one change, score it and decide to keep or revert "before you have finished reading the previous result" [6].

The machine time was modest. At five to seven minutes each, 108 eval runs come to between 9 and 12.6 hours across six days [3]. The team averaged about seven experiments a day [4]. The post warns that a loop with a subtly wrong objective produces "40 confident experiments per day, all climbing a metric that does not mean what you think it means" [17]. In my view the gap between those two figures is the human reviewer. The team put it this way: "writing code was never the constraint. The agent produced hypotheses and implementations faster than we could evaluate them" [7].

According to the post, teams tell themselves that the agent removes the tedium and drafts the next hypothesis while you get coffee [18]. The post says that part is true [18]. The agent used the leaked question category "without comment, because from inside the loop it is just an available input that improves the score" [9].

The agent was better at volume. It read hundreds of failing questions and found three bugs the scores never showed: image captions stored but never indexed for search, grep-only queries falling into a path that sorted by creation time, and a case mismatch in tree paths that left thirty searches per eval returning nothing [13]. The human's side of the split, as the post frames it, was one question asked repeatedly: "is this measuring what we think it is measuring?" [15]

The post's own accounting is slightly inconsistent. Its summary says all four decisive interventions were a human asking whether the team was measuring the right thing [4]. Its body says three of the four were [14]. The fourth was a human reordering the agent's four ranked hypotheses, which the post calls "taste rather than validity" [12]. That reordered pick, forcing grep to combine with ranked search, was the largest win of its round [12].

The figures here come from the team's own post [1]. Its raw F1 of 0.638 beats the previous published best of 0.598 by 0.040 [2][2]. The post does not explain how its headline 0.666 differs from the raw 0.638, and the raw number is the one it sets against the prior best [1][2]. Of the 41 experiments, 41% were kept [1].

For whoever is putting an agent on an experiment loop on Monday, the three validity catches sort into a 2x2 for every change the loop wants to keep. One axis is whether production would supply every input the run used. The other is whether the metric can observe what the change produced. A change that passes both can be adopted on its score. One that passes the first and fails the second, like the interleaved facts, needs someone to read the retrieved results before it ships [10]. One that fails the first, like the category hint, means the harness gets fixed before any score counts [9]. The other rule is to ablate the most expensive step first. In this run, the most complex step in ingestion was contributing zero F1 [11].

What to watch

  • An independent reproduction of the 0.638 raw F1 on LoCoMo by a team outside Timescale.
  • Whether Timescale publishes how its 0.666 figure is computed and how it differs from raw F1.
  • Whether the harness and the log of adopted and reverted experiments are released, so others can check which gains survived removal of the category hint.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories