Build1 publisher3 min readPublished
Pi 1.0 runs MCP tools inside a sandbox so their output stays out of context
Earendil's Pi 1.0 coding agent ships native MCP support after more than a year of criticising the protocol, running tool calls inside a WASM sandbox. The team credits a more stateless July 2026 spec and a sandbox it already needed for non-LLM models, so MCP support was a small addition.
The Engineer · Build desk

What happened
- For that year, Pi's objection was that MCP loads tool schemas into the context window up front, before the model has called anything.
- Codemode, the 1.0 feature that carries MCP support, also lets Pi plug in non-LLM models such as classifiers and image models.
- Pi Durable, released alongside 1.0, checkpoints every step of an agent so a new process can resume when the old one dies.
- Earendil, which took Pi over in spring 2026, says 'hundreds of thousands' of people use the agent each week.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability Batch jobs against an MCP server, such as scanning an issue tracker, no longer grow context use with every record the server returns.
- cost The model has to write correct sandbox code against tool docs, and records its own filter drops never reach it, so a filtering bug yields a wrong answer with no visible error.
- decision Teams that turned down MCP over context cost can treat that cost as a harness design question, since Pi keeps MCP on the wire and changes only what the model is handed.
- constraint Tool authors on Pi Durable must classify every tool: an unmarked read-only tool loses its automatic crash retry, and a deploy cannot fire twice.
Under Codemode, a Pi model gets a programmable environment where it used to get a list of MCP tool schemas [4]. The harness runs a JavaScript sandbox compiled to WASM [4]. The model finds tools through their docs, writes code that calls them like an SDK, runs the calls in parallel and filters the output [4]. Only the filtered result goes back into context. MCP stays the wire protocol, according to a dev.to walkthrough of the release [4].
Mario Zechner, Pi's original creator, still steers the technology under Earendil [7]. His standard example of the old cost was Perplexity, where three MCP servers reportedly took 143k tokens of a 200k window [6]. It works out to 71.5% of the window [1]. At the commonly cited 550 to 1,400 tokens per tool definition [1], reaching 143k takes between about 102 and 260 definitions [2]. The 71.5% figure applies to a session carrying that many tools at once.
The post gives deferred tool loading one line: tools are not loaded into context up front [5]. Deferred loading addresses the schema cost. Codemode addresses the cost of results.
The announcement's demo combined the Linear MCP server with a sentiment classifier to find frustrated commenters across 167 open issues, and only the final answer reached the model [9]. I treat that as a claim about one workload. It transfers when the output can be reduced in code before the model has to read it, as in a batch scan with a short answer. When the model must read each result to choose its next step, the result enters context anyway and the sandbox saves less.
"Anything that makes the model read tool output it doesn't need is a bug," the post's author wrote [10]. Of the old context complaint, the author added: "Much of it was a harness problem." [10] I agree for read-heavy jobs like the Linear demo.
The stated reasons for the reversal are narrower than the drama. The July 2026 spec revision made MCP more stateless, and the sandbox Pi needed anyway for non-LLM models made MCP support a small addition, according to the post [11]. Both fit the rule in the 1.0 post: "We wait until something has proven itself, and only then do we consider adopting it; weighing its true functionality against its inherent added complexity." [12] The protocol improved and the added complexity shrank. The post does not show that user demand or MCP's spread forced the decision.
Pi Durable, the experimental framework for long-running, crash-resistant agents released the same day [13], has the best engineering in the release. Each tool declares a replay policy. A tool marked `replay: "safe"` re-runs after a crash [15]. An unmarked tool, such as a deploy or a payment, does not; the model is told the call was interrupted [15]. Interrupted model requests are resent, and a `requestId` gives exactly-once submission so retries cannot double-fire work [15].
What to watch
- Whether Earendil documents how deferred tool loading decides when a tool's schema enters context.
- Token measurements from Pi users running Codemode on interactive tasks, where the model reads each result, as opposed to batch scans like the 167-issue demo.
- Whether other minimal harnesses route MCP through a code sandbox after the July 2026 spec revision.