Skip to content

Build1 publisher3 min readPublished

Pi adds MCP through a JavaScript sandbox that skips tool-schema preloading

Pi made MCP a core feature in version 0.99.0 on September 29, 2026, after its creator Mario Zechner spent over a year arguing against the protocol. In Pi, models call MCP tools from JavaScript in a QuickJS sandbox, so tool schemas are not preloaded into context.

The Engineer · Build desk

Illustration accompanying Pi adds MCP through a JavaScript sandbox that skips tool-schema preloading

What happened

  • Pi gives the model four tools, read, write, edit and bash, and had left subagents, plan mode, permission popups and MCP for extensions to add.
  • Perplexity's CTO said in March 2026 that three MCP servers took 143,000 tokens of a 200,000-token window, after which Perplexity dropped MCP internally.
  • The July 2026 MCP spec revision made stateless operation the primary focus and removed the initialize handshake and session layer.
  • Pi's harness already needed a JavaScript sandbox for its Codemode feature, so adding MCP on top took little new work, according to a dev.to account of the release.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Harness builders weighing MCP now have to decide how tool definitions get loaded, because Pi shows protocol support can sit behind code execution with no schema preload per session.
  • cost Harnesses that lack a JavaScript sandbox would have to build and run one before they could copy Pi's approach to MCP.
  • precedent Minimal coding agents that leave MCP out have lost their best-known ally on the token-cost argument, and the remaining objection is narrower: it is about preloading schemas.

Zechner's case against MCP was a token count. He cited a popular Playwright MCP server that put 21 tool definitions and 13,700 tokens into the model's context on every session [3]. That comes to about 652 tokens per tool [1]. The dev.to account puts a single MCP tool definition at 550 to 1,400 tokens, so 652 sits near the low end [8]. Pi's entire system prompt and tool definitions measured under 1,000 tokens at launch [7]. The Playwright server alone loaded more than 13 times that [4]. His alternative was a CLI tool with a good README, which he argued costs nothing until the agent needs it [3].

In the preload pattern, every connected server loads full schemas for all of its tools before the agent does any work [8]. The model picks a tool, emits a JSON call and waits. The result lands in context and stays there for the rest of the session [16]. A chain of ten calls costs ten model turns [14].

Codemode moves the chaining into code. The model writes a script, and the script calls the MCP tools from inside the sandbox [12]. The model finds tools by reading their documentation. Intermediate results stay in the sandbox, and only the script's final value returns to context [12]. This is Zechner's README argument applied to MCP servers. A tool's description costs tokens when the model reads it.

The dev.to author says the 72 percent tax "becomes a rounding error" under Codemode [14]. The post does not include a measured token count for a Codemode session. Perplexity's figure is 71.5 percent before rounding [2], and the account uses it as its example of what preloading costs [9]. For that number to transfer, your tool count has to look like Perplexity's. If its definitions fell in the 550 to 1,400 range, its three servers exposed roughly 100 to 260 tools between them [3]. A harness with one server exposing a dozen tools would preload about 6,600 to 16,800 tokens [5].

Adoption is the other half of the case. MCP went to the Linux Foundation in December 2025 and has more than 17,000 public servers and hundreds of millions of monthly SDK downloads, according to the dev.to account [11]. Its author wrote, "Arguing against a protocol that your potential users already run in production is a losing position regardless of technical merit." [15] The evidence shows one prominent holdout adopting the protocol on its own loading terms.

The team titled its announcement "You Said No MCP!" [5], a fair amount of self-mockery for a release note. I think Codemode is the right design for a harness whose rule is that anything not needed is not built [17]. The model calls MCP tools from code. That lets Pi connect more servers without adding their schemas to every prompt. The savings depend on the model writing scripts that run.

What to watch

  • A measured token count for a Pi Codemode session against the Playwright MCP server Zechner used as his example.
  • Whether harnesses that preload MCP schemas switch to code-execution calling now that the July 2026 stateless spec revision is out.
  • Whether Perplexity, which dropped MCP internally after its 143,000-token measurement, takes the protocol back up.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories