Skip to content

Build1 publisher3 min readPublished

Your MCP tool definitions cost 64% more than your token counter says

A standing measurement of 14 MCP servers finds Claude's tokenizer counts schema text a median 64.1 percent above tiktoken, the counter every published cost study uses.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Your MCP tool definitions cost 64% more than your token counter says
Generated illustration

What happened

  • The author previously ran 72 trials on what an MCP tool should return, leaving open the question of how much context attached servers consume before an agent does any work.
  • Of the six vendor-published 2026 numbers on MCP context cost that get cited, exactly one is a real measurement study: StackOne's, dated 2026-03-31, which measured GitHub and Atlassian and worked through Cloudflare's code-mode case.
  • Of the other five cited sources, one uses hypothetical unnamed servers, three contain no server-specific token study at all, and one could not be located at any domain.
  • The published state of the art is therefore a single snapshot still showing March's numbers, while MCP servers change every few weeks.
  • loadline is a standing measurement of 14 MCP servers under methodology 0.2.0, with one run dated 2026-08-18, monthly runs planned from there, every row reproducible from the repo with the reader's own credentials, and a stack calculator as its front door.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Someone has finally put a repeatable measurement behind a line item most teams estimate by feel: how much of an agent's context window the attached MCP servers consume before the first user message [5]. The finding that matters is not the size of the bill but the accuracy of the meter, because on one fixed schema string Claude's tokenizer counts a median 64.1 percent higher than tiktoken's o200k_base [8].

Start with the prior art, because it is thinner than the citation count suggests. The author checked the six vendor-published numbers that get cited and reports that exactly one, StackOne's from 2026-03-31, is a real measurement study, covering GitHub and Atlassian and working through Cloudflare's code-mode case [2]. Of the other five, one uses hypothetical unnamed servers, three contain no server-specific token study at all, and one could not be located at any domain [3]. That leaves the published state of the art as a single March snapshot of a surface that changes every few weeks [4].

The replacement, called loadline, covers 14 servers under methodology 0.2.0, with one run dated 2026-08-18, monthly runs from here, and every row reproducible from the repo with your own credentials [5]. All three counters are pointed at the same canonical serialization from methodology section 1.5: same bytes, three counters [6]. Across the 12 rows that produced counts, Claude's premium over o200k runs from 47.0 percent on github to 70.5 percent on kubernetes, with 8 of the 12 between 60 and 66 [7]. Gemini stays within 11 percent of o200k on every row, which rules out a bigger-models-count-more effect and points at Claude's tokenizer on schema text specifically [9]. Every MCP cost study the author has seen counts with tiktoken, the only one of the three counters that runs offline and free [10]. If your loop runs on Claude, those studies are describing a context load roughly 60 percent smaller than the one you pay for [11].

The github row makes the gap concrete: 47 tools at 59,084 tokens under o200k_base and 86,843 under claude-opus-5 [12], a 27,759-token difference [13] that is 43 percent of a 200k window spent before the first user message [14].

Two qualifiers travel with that number. The 47 tools are what the server exposed to the classic personal access token this run used, and the spec lets a surface vary with the authorization presented, so the real GitHub surface is larger; loadline publishes auth scope per row for that reason [15]. And nobody on a tool-search client pays the naive figure: Claude Code has defaulted to progressive disclosure since January 2026, which on this surface models out to roughly 4,300 to 6,800 tokens, a 500-token stub plus 3 to 5 tools at the measured 1,257-token average [16] - nine to fourteen times less than the naive load [17]. The author flags "modeled" as load-bearing there, and it is.

Cheap surfaces exist. Cloudflare's aggregate endpoint answers tools/list with three tools, docs at 362 tokens, search at 572, execute at 660, totalling 1,596 under o200k, 37 times smaller than github's naive load, with hygiene holding at B (85.19) [18]; the qualifier is that Cloudflare also ships 16 product-scoped endpoints this run did not enumerate [19]. Google's Chrome DevTools MCP exposes 52 tools at 7,984 tokens, 153 per tool, against Playwright's 167, filesystem's 192 and linear's 335, with hygiene at B (80.28) [20]. Github's 1,257 tokens per tool is about eight times that [21].

Separately, fetch and postgres both failed to launch from a clean install at 15:38 UTC on 2026-08-18 [22], because both declare unbounded dependencies on the MCP Python SDK and a resolve landing on mcp 2.0.0 breaks them in two different places: a renamed McpError symbol for fetch, a missing mcp.server.fastmcp module for postgres [23].

Watch whether the 64 percent premium survives the next monthly run, and whether anyone measures progressive disclosure instead of modelling it.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories