Skip to content

Build1 publisher3 min readPublished

Cloudflare moves durable execution under the harness, and the platform starts choosing it

The Agents SDK now carries durable execution, sandboxed code execution and a durable filesystem. Cloudflare's argument is that harnesses cannot own those problems, which makes them a platform decision.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Cloudflare published a post announcing it is bringing more agent harnesses and frameworks to Cloudflare, starting with Flue.
  • Cloudflare states that 2026 is the year agent harnesses go to production, and that harnesses have matured to the point where teams deploy agents as load-bearing infrastructure rather than prototypes.
  • Cloudflare names Codex, Claude Code, OpenCode, Pi and Project Think as harnesses, the software that controls a model's access to the outside world.
  • Cloudflare built Project Think as its first-party agent harness.
  • Cloudflare says that working with customers running agents in production surfaced a common set of distributed systems problems: how an interrupted agent automatically and gracefully resumes without losing context or wasting tokens, how agents run untrusted code securely, and how agents use the tools they were trained for.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Cloudflare has pushed durable execution, dynamic code execution, a durable filesystem and dynamic workflows into the Cloudflare Agents SDK, where any harness built on the SDK can use them [7]. The framing matters more than the feature list: Cloudflare argues these problems cannot be solved inside a harness at all, because they are tied to state, storage and compute, and therefore to the platform the agent runs on [6].

That claim comes out of the company's own harness work. Cloudflare built Project Think as its first-party agent harness [4], and says that running agents in production with customers surfaced the same distributed systems problems every time: how an interrupted agent resumes without losing context or re-spending tokens, how it runs untrusted code securely, and how it gets the tools the model was trained for [5]. The published version of the argument is wrapped in the usual calendar rhetoric, that 2026 is the year agent harnesses go to production [2], which can be ignored without losing the engineering point.

Cloudflare draws the result as three layers: a framework on top providing project structure, conventions, integrations, CLI and developer experience; a harness in the middle running the agentic loop that calls tools, reads results and manages context; and the runtime underneath supplying compute, state and storage [8]. Flue, an open-source framework from the team behind Astro, is the first to build on that bottom layer [9]. It shipped 1.0 Beta this week on the Pi harness, the same harness OpenClaw uses [10].

Flue's approach is declarative: rather than scripting an orchestration loop, you describe the context an agent needs, its model, skills, sandbox and instructions [11]. Cloudflare's worked example is a triage agent that intercepts a bug report, reproduces it in a sandbox and diagnoses it in under 25 lines [12]. Around that sit pre-configured Channels for Slack, GitHub, Linear and Discord that absorb event verification and dispatch boilerplate [13], React hooks in @flue/react that stream agent state, tool execution and live messages into a frontend [14], and a CLI that generates a Markdown blueprint your own coding agent can edit when you run something like `flue add channel slack` [15]. Crash tolerance is handled by Durable Streams, an append-only log of every prompt, tool response and model choice, so a host crash, an LLM provider timeout or a restart does not wipe the turn [16].

The consequence for anyone shipping this is coupling, and it runs in the direction most teams do not plan for. If resumability, sandboxing and durable filesystem semantics live in the runtime, then picking the runtime narrows the set of harnesses that can actually use them. Cloudflare names five harnesses as production-grade in its own post, Codex, Claude Code, OpenCode, Pi and Project Think [3], and only two of the five, Pi and Project Think, appear in the stack it drew [17]. The primitives are advertised as available to any harness on the Agents SDK [7], which is a real offer and also a boundary condition: the harness has to be on the Agents SDK.

Worth pricing before committing: whether a harness you already run can adopt these primitives without a rewrite, whether durable state written through the SDK is portable off Cloudflare, and how the durable log behaves on replay when a tool call has already had an external side effect.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories