Build1 publisher3 min readPublished
255 tool schemas, 91K tokens: pricing the two MCP costs nobody budgets
An engineering diary for a small MCP client puts numbers on schema bloat in the context window and on the roughly 950 hand-written lines it took to ship with zero dependencies.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A developer publishing as mcptokensaver spent six weeks building mcptoon, a CLI tool that sits between AI agents (Claude Code, Cursor, Codex) and MCP servers, and framed the write-up as an engineering diary rather than a product pitch.
- MCP tool schemas get injected into the context window as JSON; 255 tools equal roughly 91K tokens of JSON braces, brackets, quotes and commas before any actual work happens.
- mcptoon ships with zero dependencies: pyproject.toml declares dependencies = [], described as not minimal or few but zero.
- mcptoon keeps schemas out of context: the agent runs shell commands and only the compact result enters context.
- The trigger for the zero-dependency decision was the uv security incident, in which a transitive dependency in a popular Python tool had a supply chain vulnerability and thousands of projects were affected not through their own fault but because someone upstream did something wrong.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer writing as mcptokensaver spent six weeks building mcptoon, a CLI that sits between coding agents such as Claude Code, Cursor and Codex and the MCP servers they call, then published the engineering diary rather than the pitch [1]. Two of its numbers are the ones teams running MCP rarely put on a spreadsheet: what tool schemas cost before any work happens, and what it costs to refuse other people's dependency trees [2][3].
The first cost is context. MCP tool schemas are injected into the context window as JSON, and by the author's count 255 tools come to roughly 91,000 tokens of braces, brackets, quotes and commas before an agent does anything useful [2]. That is an average of about 357 tokens per tool, paid on every turn that carries the schema block [1]. mcptoon's answer is indirection: keep the schemas out of context, have the agent run shell commands, and let only the compact result enter the window [4].
The second cost is the interesting one, because it is usually invisible. The trigger, according to the author, was the uv security incident, in which a transitive dependency inside a popular Python tool carried a supply chain vulnerability and thousands of projects were affected through no action of their own [5]. He checked his own history: hundreds of packages installed in a year, none of their dependency trees audited [6]. The rule that followed was not "minimal" but zero third-party imports, standard library only, with `dependencies = []` in pyproject [3][7].
Here is what that bought and what it cost. HTTP plumbing, including SSE streaming, error handling, retries and auth, came to about 200 lines against roughly 30 with `requests`; a single POST is eight lines with `urllib` and one with `requests` [8]. The argparse dispatch function for 15 subcommands ran to about 400 lines against roughly 150 with `click` [9]. Hand-rolled response validation added about 300 lines that `pydantic` models would have done themselves [10]. Terminal styling with raw ANSI codes was about 50 lines, and the author says `rich` was simply not needed [11]. Total: roughly 950 lines of substitute code [2], of which about 420 lines are pure surplus over the library versions where he gave both figures [3]. His own verdict is split rather than triumphant: worth it for SSE, because he learned how the protocol works, and not worth it for basic HTTP, which was just plumbing [12].
The payoff is measurable too. The install is 250KB in 0.3 seconds [13], against roughly 5MB for `requests` plus its dependencies and roughly 15MB for `pydantic` plus its dependencies [14] - about 80 times the wheel size from two libraries alone [4]. The 486 tests run in 0.5 seconds because there are no heavy fixtures [15], near 970 tests a second [5]. And zero is a runtime claim, not a toolchain claim: pytest remains a dev dependency, though without pytest-mock, pytest-cov, responses or httpx [16].
Worth watching: every figure here is self-reported by the tool's author [1], the 91K token count scales with how verbose your particular servers are [2], and the published size comparison breaks off mid-list at `click`, so the library-side footprint is understated rather than complete [14]. The open question for anyone copying the approach is whether roughly 950 lines of hand-maintained plumbing [2] costs less over a few years than triaging other people's CVEs [5].