Skip to content

Build1 publisher2 min readPublished

Pi's v0.99.0 adds MCP by running tool calls in a sandbox outside the model's context

Pi shipped MCP in v0.99.0 using a sandbox that keeps tool schemas out of context, where three servers took 143,000 of Perplexity's 200,000 tokens. Other harnesses can copy the design if they are willing to host an interpreter that runs model-written code.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Pi's v0.99.0 adds MCP by running tool calls in a sandbox outside the model's context
Generated illustration

What happened

  • Pi's homepage used to read "Pi does not support MCP", and v0.99.0 arrived on September 29 under the post title "You Said No MCP!".
  • Perplexity's CTO reported the 143,000-token figure at a conference in March 2026, and the company then dropped MCP internally.
  • In Pi's Codemode, the model writes JavaScript that runs in a QuickJS sandbox inside the harness, with MCP tools exposed as functions and only the return value sent back to context.
  • Pi cited the July 2026 MCP revision, which removed the initialize handshake and session layer, as a reason for adding support.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams with many connected servers have to choose between pruning tools and moving tool calls into a harness interpreter, because under default loading every added tool raises the cost of every session.
  • constraint The pattern was cheap for Pi because its sandbox already existed. A harness without one has to build an interpreter and accept model-written code running against live tools before it sees any saving.
  • exposure When one script can reach every connected vendor's tools, a single piece of model-written code can act across all of them, and the sandbox's permissions become the security boundary.

Perplexity's number tells you about Perplexity's tool count. Tool definitions reportedly run 550 to 1,400 tokens each [6]. If schemas alone filled 143,000 tokens across three servers [3], those servers carried roughly 100 to 260 definitions [2]. Pi's own example is smaller. A Playwright MCP server loads 21 definitions and 13,700 tokens per session [7]. That is about 650 tokens a tool [3], or 6.85% of a 200,000-token window [4]. The post's hand-counted GitHub create_issue definition comes to about 280 tokens, under the reported range [5]. The author says these are not lab benchmarks of Pi or Perplexity [8].

The figures share one client default. The full schema for every connected tool loads at session start, and the model pays for it whether or not the tool is ever called [5]. Codemode drops that default. The model sees tools as discoverable documentation [11] and writes a script for the harness to run. In the post's example, a list_labels result goes straight into create_issue inside the sandbox. The model gets back only an issue number and a URL [17].

The best engineering decision here is where the sandbox sits. Earlier code-execution designs shipped one sandbox per MCP server, and those sandboxes could not call into each other [12]. Pi put a single sandbox in the harness, so one script can chain tools from different vendors and return one result [13]. I'd copy that placement before anything else.

Pi's own reasons give the protocol some of the credit. The first, the July revision's removal of the handshake and session layer [14], fixed a deployment cost inside the protocol. The context cost sat in the client [5]. The second reason was that Pi's QuickJS sandbox already existed, so MCP support was cheap to add [15]. The third was stated as philosophy. "We believe the best way to positively influence something is to embrace it," Pi's team wrote [16].

So the context half of the claim holds up, and the reversal does not rest on it alone. I'd expect the saving to be largest where a session touches few of its connected tools, since any documentation the model looks up still has to be read. The post does not include a token count for a Codemode session to set against the Playwright server's 13,700 tokens [7].

What to watch

  • A published token count for a Codemode session against the Playwright server's 13,700-token default load.
  • Whether mainstream MCP clients switch to deferred schema loading by default, cutting the per-session cost without a sandbox.
  • How Pi limits what a single QuickJS script may call when several vendors' servers are connected at once.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories