Build1 distinct publisher3 min readPublished
One measurement puts a 50-server tool catalog at 56 percent of a 128K context window before the first prompt. The per-server arithmetic is the part worth budgeting.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Divide the headline pair and the per-unit figures turn out to be more useful than either number alone. 71,929 tokens across 255 tools works out to roughly 282 tokens of schema per tool [12], against about half a token per tool in the names-only manifest [13]. Averaged over the 50 servers, each one installed carries a standing charge of about 1,439 tokens on every session it loads into [14]. That is the number to budget with, because it is the one that moves when you uninstall something.
The ratios in the writeup do not agree with each other, which matters for anyone about to quote them in a planning doc. The headline band is 4x to 32x [5], and Scalekit is cited as landing on the same 32x ceiling [6]. But the two anchor numbers divide out to about 585x [11], and the Firecrawl pair of ~200 tokens against ~44K divides out to about 220x [15]. Neither sits inside the band. Catalog-load cost and task-completion cost are plainly different measurements, but the piece does not show that arithmetic, so treat the 32x band as spend per task and the 71,929 as context pressure, and keep them apart.
On provenance: the 71,929 figure comes from the author of mcptoon, a CLI whose entire pitch is the gap between those two numbers [10]. Counting tokens in a JSON blob with tiktoken is about the most reproducible measurement available in this area [2], so anyone with 50 servers configured can check it in an afternoon. What the piece does not contain is a third party who did that against the same 255 tools.
The weakest link in the cost argument is the compounding step. The claim is that schemas re-enter every request, converting a one-time load into a per-request fee [9]. That is asserted rather than measured, and the piece does not say how the clients it names, Claude Code and Cursor, treat a prompt prefix that never changes between turns [19]. If that prefix is cached, the money argument shrinks and the context argument survives intact, because 56 percent of a 128K window is gone either way [3].
Which leaves the context number as the durable one. Anthropic's engineering blog is cited as describing a cut from 150K tokens to 2K [7], a 75x reduction [16], which is the company behind the protocol conceding that whole-catalog injection does not scale. Meanwhile 71,929 tokens is 112 percent of a 64K window [17], so the cheaper model tier is unavailable to a 50-server config, and the alternative the piece names is paying for a bigger-context model purely to absorb boilerplate [18]. The schema itself earns its keep twice a session, once when the model picks a tool and once when it fills in arguments [8]. Everything between those two moments is rent on a document nobody is reading.
None of this requires a new protocol to fix. It requires a shorter server list, or discovery by index instead of by encyclopedia.
Ranked by verification strength, evidence, and original report placement.
The author says the counts were measured with tiktoken, OpenAI's tokenizer.
On a 128K context window, 71,929 tokens of tool definitions consume about 56 percent of the window before the first user message is processed.
On a 64K context window, common for cheaper and faster models, the tool definitions do not fit at all.
The tool schema matters twice per session: once when the model picks a tool and once when it fills in arguments.
The writeup addresses users running multiple MCP servers in Claude Code, Cursor, or similar clients.
The writeup names two options when tools do not fit: uninstall servers, or pay for a big-context premium model purely to absorb boilerplate, which it calls the quiet budget killer.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One self-run measurement, unreplicated
The quantitative core is a single tokenizer run on the author's own machine and config, reported with a reproduction command but no published config, raw output or third-party replication. Supporting citations to Firecrawl, Scalekit and Anthropic are paraphrased without links or methodology. The internal arithmetic (per-tool, per-server, window-share) checks out against the stated anchors, which is why this is not lower, but everything upstream of those anchors rests on one interested author.
No adoption data supplied
The cluster contains no installs, downloads, stars, users, deployments or third-party usage of mcptoon, and no measured population data on how many agent users run large MCP configs. The only observations are the author's own benchmark, his release writeup, and a secondhand retelling of another benchmark — none of which measure uptake.
Overstated relative to what is shown
Direction is plausible and the arithmetic is internally consistent, but presentation runs ahead of evidence: a headline 4x-32x band sits beside the author's own numbers implying ~585x and ~220x; 'independent groups keep arriving at the same conclusion' is supported only by unlinked paraphrases; compounding cost is asserted at 'typical frontier pricing' with no model, and the fix is the author's own product. The measurable per-window share (56 percent of 128K, overflow of 64K) is the part that holds up, which keeps this from being higher.
Author sells the remedy
The sole source is written by someone who works on mcptoon, published on a dev.to account named for the problem being described ('mcptokensaver'), with the anchor measurement produced by that project's own tooling and the reproduction path routed through the product's CLI. The two cited corroborators are themselves commercial vendors in adjacent tooling. Disclosure exists but is minimal.
Low-moderate
Confidence is limited by single-source, single-publisher coverage from an interested author, unlinked corroboration, and absent adoption data. It is not lower because the mechanism (schemas injected wholesale per request) is described concretely, the arithmetic is reproducible from the stated anchors, and the directional finding is consistent across the figures cited even where magnitudes conflict.
build
Ten MCP servers, 847 tool schemas, 112K tokens before you type: curation is a cost line1 distinct publisher
build
Every MCP server you add costs about 11,000 tokens before anyone types a word1 distinct publisher
build
Five MCP servers, 22,185 tokens: server count is now a context budget line1 distinct publisher
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 25, 2026