Skip to content

Build1 publisher2 min readPublished Updated

Pi Durable checkpoints every agent step so crashed runs resume from storage

Pi Durable ports the Pi agent harness to TypeScript and checkpoints every step to one of three stores, so crashed agents resume where they stopped. Whether it can replace hand-built resume code depends on how it treats a tool call cut off mid-flight.

The Engineer · Build desk

Illustration accompanying Pi Durable checkpoints every agent step so crashed runs resume from storage

What happened

  • Pi Durable runs on any JavaScript runtime, including Node, Bun and Cloudflare, with local or remote execution and Memory, SQLite or JSONL storage backends.
  • One harness can run parallel, branching conversations, such as a main channel and side threads, without any of them blocking the others.
  • Application state such as a to-do list is kept in documents beside the transcripts, so several users or UIs can watch and steer one agent at once.
  • Pi, now part of Earendil, put both Pi 1.0 and Pi Durable on the Hacker News front page the same day.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams that maintain their own checkpoint-and-resume layer for long agent jobs now have to decide whether to retire it and make Pi Durable's step records the system of record.
  • cost Spreading one agent's workers across hosts needs a store they can all reach, and a team would have to supply that beyond the SQLite and JSONL files on the list.
  • exposure Any tool that moves money or sends messages needs idempotency or a rollback path before it runs under automatic resume, or a restart can repeat the action.

Externalizing state turns the agent loop into a log. Every step is written to a store as a checkpointed task. After a failure or restart, agents and subagents resume "from their last exact state," according to Latent Space's AINews summary of the release [2][11]. In that design the process is a replaceable worker that reads the latest record and carries on [1][2].

Which of the three storage backends a team picks decides whether that promise holds [1]. A Memory store lives inside the process it is meant to protect. I'd expect it to serve tests and local work, with real crash survival coming from SQLite or JSONL [3]. JSONL, one appended record per line, suits a log of steps [3].

The hard case is a tool with side effects. Suppose the checkpoint is written when a step completes. A process that dies after a tool has charged a card, and before the checkpoint lands, resumes from the state before the charge, and a plain resume issues the call again [2]. The summary's own example of a durable task is a multi-step checkout with rollback, packaged in an installable Extension alongside prompts, tools and hooks [5]. The authors appear to expect compensation to happen at the task level. The summary does not say whether a tool call interrupted before its checkpoint is retried, skipped or rolled back on resume [2][5].

Two more features write into the same state. Background compaction summarizes older messages to stay inside token limits without pausing the agent [6]. A summary built off the main path has to be committed against the exact transcript it replaces. Otherwise a resume could load both the summary and the messages it was meant to drop. Hot-swapping lets tool and extension code change while the agent runs, and the next tool call picks up the new version [8]. Add checkpoints, and a run that crashed on one version of a tool can resume on the next. This is pleasant until a tool's argument schema changes between checkpoint and resume, and new code has to accept arguments that old code recorded [8].

In my view the store is the right home for a harness's state once a run lasts long enough to span a deploy or a host restart. A process that owns the run loses it on either, and Pi Durable moves every stateful component out of the process [1][2].

What to watch

  • Docs or source that state what happens on resume to a tool call that was running when the process died: retry, skip, or rollback.
  • Which backend a Cloudflare deployment uses, and whether a shared networked store joins Memory, SQLite and JSONL.
  • Independent reports of Pi Durable resuming long runs after real process kills, outside the launch-day account.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories