Build1 distinct publisher3 min readPublished
A killed side project produced the useful result: 11 of 31 pilot pairs, $5.60, and zero Java fixes from the cheap model. A savings figure without a pass rate is not a number.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
A token counter cannot tell a compressed context apart from an agent that read almost nothing and stopped early. Both register as savings. According to the dev.to writeup, the 97 percent figure the author was looking at was the second case, and the token-savings numbers were actively misleading [2]. The metric he had built the stack around was scoring the absence of work as efficiency.
That leaves completion as the only usable denominator, and the pilot arithmetic gets awkward there. About $5.60 bought 11 finished task pairs [9], which is roughly 51 cents per pair [2]. If a pair means one run per stack version, that is 22 runs at about 25 cents each, and the full 4,800-run design [8] lands near $1,220 [3]. The author's own estimate was past $1,200 on the cheapest model [11]. Two independently reached numbers agreeing is the one piece of good news in the receipt: an 11-pair sample is small enough to be unrepresentative, and it wasn't.
Then apply the gate. A zero in the fix column makes tokens-per-verified-fix undefined, so the Java portion of that spend bought nothing to compare against [13]. A trustworthy run therefore has to happen on a model that can land a patch, and Sonnet is 2x Haiku on both input and output, $2/$10 per million against $1/$5, which puts the same test past $2,400 [12], or about 50 cents a run [4]. That is the standing price of a defensible savings claim, at roughly 218 times the volume the author actually completed [5], and it excludes EC2 time, Docker builds, and the multi-day work of getting the harness running on both EC2 and an Apple Silicon Mac [10].
The two tools that were cut failed the same way the metric did. Headroom's docs said no behavioral changes were needed, but `headroom init claude` only registers an on-demand MCP tool, and `headroom doctor` showed nothing routed through it without a separate proxy process and an ANTHROPIC_BASE_URL override [4]; its `mcp serve` also crashed against a current MCP SDK unless pinned to mcp<2 [5]. LiteLLM's usage-based routing balances load across provider endpoints rather than sending hard steps to stronger models, and Claude Code fixes one model for an entire session in any case [6]. No amount of token measurement surfaces either problem. Asking whether the component is in the request path at all does, and that is the same question a completion gate asks of the agent.
What survived to the bench was Graphify, Serena, LeanCTX and Caveman [7], and none of them ever got a verdict anyone should trust. That is a more honest place to stop than a percentage, and it sets the reporting rule the rest of this tooling category has been ignoring: tokens per passing task, or nothing.
Ranked by verification strength, evidence, and original report placement.
The project ended a few weeks in, not with a working stack and a savings number, but with proof that the token-savings numbers were actively misleading; the 97 percent savings figure was actually a silent failure.
Haiku was not trustworthy in the pilot: it got zero correct fixes on the Java tasks, and it broke two of the three Python tasks the plain baseline had already solved.
The author built two public projects: token-optimization-stack, a repo with setup docs for tools that reduce token spend, and token-stack-benchmarks, a harness to test whether they worked, after repeatedly hitting token limits using Claude Code for engineering work.
The first version of the stack contained five tools: Graphify (codebase knowledge graph), Serena (symbol-level navigation and editing), Headroom (advertised as transparent context compression), LiteLLM (routing layer), and Caveman (compresses the agent's own output).
Headroom's docs said it needed no behavioral changes once installed, but 'headroom init claude' only adds an on-demand MCP tool the agent can call, and 'headroom doctor' confirmed nothing was routed through it unless a separate proxy process with an ANTHROPIC_BASE_URL override was also run.
Headroom's 'mcp serve' command crashed against a current MCP SDK and needed an old pinned mcp<2 dependency just to start.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed first-hand receipts, single unreplicated source
The author supplies concrete, checkable artifacts: named commands and their observed behaviour, a pinned dependency version, model IDs and per-million token prices, a $5.60 receipt, 11-of-31 completion, and per-task outcomes including a num_turns: 1 stop. That is unusually specific for a practitioner writeup and the arithmetic reproduces the author's own cost estimates. But it is one self-reported pilot, incomplete by design, with an n=3 Python correctness sample the author refuses to rate, no tool versions or commit hashes, no vendor response, and no independent replication.
One practitioner's pilot; no external usage data
Observed adoption is limited to a single engineer installing five tools, dropping two, and running a partially completed pilot benchmark on a cheap model. There is no user count, download figure, organizational deployment, or vendor usage disclosure anywhere in the cluster, so real-world uptake of these token-optimization tools cannot be scored above a single-practitioner trial.
Tooling savings claims overstated; the writeup itself is restrained
The gap sits with the token-optimization tooling and the dashboards, not with this article. Headroom's 'no behavioral changes once installed' did not match observed behaviour, LiteLLM's routing did not do what its framing implied, and the single most impressive savings number in the pilot (97%) came from a run that performed no work. Against that, the author actively deflates his own findings, refuses to state a pass rate from three tasks, and names the cheap-model confound, so the positive gap reflects the ecosystem's claims rather than the narrator's.
No disclosed commercial stake; mild self-publishing visibility incentive
The author publishes a negative result that kills his own two public repos and names no employer, vendor relationship, sponsorship or affiliate arrangement with any of the tools assessed, which cuts against promotional distortion. The residual incentive is the ordinary one for a self-published developer post: an attention-earning contrarian framing ('here's why I killed it', a 97% number that was a failure) on a personal dev.to byline, plus the absence of any vendor right of reply.
Confident on the specific incidents, weak on generalization
Confidence is moderate: the concrete events (install behaviour, dependency pin, routing semantics, $5.60 for 11 of 31 pairs, zero Java fixes, the 97% no-work run) are described precisely enough to be trusted as reported, and the cost arithmetic checks out internally. Confidence drops for the broader conclusion that token metrics are systematically misleading, because it rests on one incomplete pilot, one cheap model, three Python tasks, one author and one publisher, with the stack-versus-model confound unresolved.
science
OX Security says MCP command execution is a design choice, so server owners own the risk1 distinct publisher
build
Ornith-1.0's benchmarks are fine. Ollama can't parse its tool calls.1 distinct publisher
build
Zalando's durable agentic engineering win was a proxy, not a model1 distinct publisher
build
The agent did not fail, the client did: 90 logged MCP trials and a validator that ate the calls1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 25, 2026