Build1 publisher2 min readPublished
One developer's step-level audit finds 97 of 200 Claude Code steps fit a local model
One developer's audit of Claude Code found 97 of 200 agent steps could run on a local model, against 4 of 100 whole requests. That makes the agent step the unit to route on, on evidence from one person's sessions and one RTX 4070.
The Engineer · Build desk

What happened
- The author blames the request-level result on planning: a one-line request leaves the agent to read files, run commands and choose next moves, and that planning needed the frontier model.
- WebFetch was the clearest local case: across 20 real fetches some pages never loaded and some answers were partly wrong, but loaded pages were usually extracted correctly.
- A one-line rule sending WebFetch steps local beat a 9B model's probability judgment, so the router now applies rules first and hands the judge only what is left.
- The example config pins Agent and Task sub-agents and commit messages to the frontier, defaults unmatched steps to the frontier, and ships with the probability judge off.
- Reading the judge's Y or N probability requires Ollama's native /api/chat, because the OpenAI-compatible endpoint drops the probabilities without an error.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Copying the example config offloads WebFetch steps alone; getting closer to the 97-step ceiling means switching on the judge with a threshold measured on the same judge model.
- exposure A logprob judge can misroute steps with no error raised, since an answer token outside the top 20 scores as zero and ' Y' and 'Y' arrive as different tokens.
- constraint Claude Code already hands fetched pages to a small, fast model, so a local WebFetch rule displaces a small hosted model and its saving is priced against that model's tokens.
Four percent against 48.5% is a factor of about twelve [1][2][3]. "So much for cancelling the subscription," the author wrote after the first count [18].
Half the steps is a claim about step counts. For it to become half the spend, the steps that pass locally would need to carry as many tokens as the planning steps that stay on the frontier [4]. The count gives a fetch-and-extract step the same weight as a step where the agent decides what to try next [3][5]. The post does not report per-step token counts, how a step was scored as fine, or which model the 200-step pass ran on. Its example request body names qwen3.5:4b, while the WebFetch rule was measured against a 9B model's judgment [12][9].
The router itself is good engineering. Walk a step through route(). It checks the rules top to bottom and returns on the first match [17][10]. It asks the judge for a probability only when a judge and a local threshold are both set [17]. Anything else gets the default [17]. Every Decision records its reason as "rule", "probability" or "default", so a routing log explains each choice [17]. Commit messages stay on the frontier, per the config comment, "until a probability judge is added" [10]. A fetch that fails also goes to the frontier, through on_fetch_error: frontier [7]. I think rules-first is the right order for a config other people will copy, because a rule match needs no model call and anyone can read it. "Rules decide what rules can decide," the author wrote [19].
The judge asks for a single letter, Y or N, and reads that letter's probability from logprobs on Ollama's native /api/chat [12]. Its request body sets temperature 0, num_predict 1, top_logprobs 20 and think: false [12]. The think and num_predict settings matter for thinking models, whose first token starts their reasoning instead of giving the answer [14]. Ollama 0.12.11 or later is what the repo's scripts expect [13]. On the response side, the code strips whitespace from each returned token, sums the probability that landed on the candidate letters, and declines to trust the judgment when that total is small [16].
What to watch
- Per-step token counts from the same sessions, which would turn the 97-of-200 step share into a share of spend.
- A measured threshold for the probability judge on the same judge model, the condition the config sets before probability leaves null.
- Step-level counts from other developers' Claude Code sessions, especially ones heavy on the Agent and Task sub-agents the config pins to the frontier.