Skip to content

Build1 publisher3 min readPublished

Tool-call parsing splits coding agents into tiers on some models, not on native tool-call models

Polyglot's developer ran six coding agents 30 times on each of seven local models, and three never made a tool call on models that write calls as text. The author's own error bars say 30 runs can sort agents into tiers but cannot rank two agents inside one.

The Engineer · Build desk

Illustration accompanying Tool-call parsing splits coding agents into tiers on some models, not on native tool-call models

What happened

  • The toolshim in goose, a second small model that rewrites text into tool calls, scored 28 on qwen2.5-coder 32B but cut goose from 18 to 5 on gpt-oss.
  • Ollama's default 4,096-token context had silently cut the prompts of goose, Hermes and opencode in the author's earlier runs, so those earlier numbers were wrong.
  • Three pre-release builds of Polyglot 0.13.2 scored 179, 173 and 172 out of 210, landing in the same tiers each time.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Anyone running a model that emits native tool calls, such as Qwen3.8-27B, can pick an agent on prompt size and editing behaviour, since completion rates stop separating the field there.
  • constraint A grid of 30-run cells supports sorting agents into tiers only; a gap of a few runs inside a tier, smaller than Polyglot's own build-to-build spread, cannot be read as a ranking.
  • exposure Agent benchmarks run on Ollama's default context can score agents whose instructions were cut off, with no error raised to show it happened.

The failure on qwen2.5-coder starts in parsing. Served through Ollama, all three sizes of that model write their tool calls as plain text inside the reply [9]. The text is a usable `{"name": "edit_file", ...}` block. An agent that listens only on the native tool channel sees no call. It treats the reply as a final answer, and the task ends with nothing done [9]. According to the post, pi, Hermes and opencode made no tool call in any of the 30 runs on those models [9]. goose made some calls but completed none of the tasks [10]. The exception is goose's toolshim. It hands the model's text to a second, small model whose job is to rewrite it as a proper tool call [11]. Running a second model to read the first one's output is a lot of machinery. On qwen2.5-coder 32B it won anyway, scoring 28 against Polyglot's 24 [11]. Under the pre-set tiers, 28 is reliable and 24 is works sometimes [3]. On gpt-oss, which emits calls in its own trained format, the same shim dropped goose from 18 to 5 [c12, c13]. That moved goose from works sometimes to fails [2]. Polyglot does the job with a parser [14]. A parser handles the formats it was written for. Outside the 7B, Polyglot missed 13 of 150 runs, and the author blames the parser for up to 5 of them: a model stuttering an empty tag before the real one, or running past a broken closing tag into an invented next turn [20]. Those formats, and the others found in reruns, are fixed in 0.13.2, according to the author [20]. The tier cut-offs were set before the first run [3]. With 30 runs, 28 successes is consistent with a true rate anywhere from 79% to 98% [4]. "So a 30 next to a 28 is a tie, not a win," the author wrote [5]. Polyglot's own builds show the noise. Three pre-release builds of 0.13.2 scored 179, 173 and 172 out of 210, all in the same tiers [21]. That is a spread of seven runs, or about 85% against 82% [4]. The author attributes the spread to the models, not the code [21]. Where models emit calls through the native channel, the agents converge. Every agent was reliable on Qwen3.8-27B, four of six on qwen3-coder and three on Devstral [7]. "If you run one of these models, pick your agent on other grounds," the author wrote [8]. Prompt size is one. Measured from Ollama's request log, Polyglot opens a task with about 1.2k tokens of fixed prompt and pi with 1.2-1.6k, against 3.9-4.7k for Hermes and 5.5-7.6k for goose and opencode [16]. Ollama's prompt cache is reused better by pi than by Polyglot, a gap the author called "something for me to fix, not to advertise" [c17, c18]. Prompt size also explains the correction to the earlier post. Ollama defaults to a 4,096-token context and silently cuts anything longer [15]. The opening prompts of goose and opencode overshoot that by roughly 1.4k to 3.5k tokens before any work starts [5]. In the first runs, goose, Hermes and opencode were working with most of their instructions missing [15]. Guarding against this is the best engineering in the post. Every model now runs at a 32k context, and the harness refuses to start below 16k [19]. After each agent's runs it scans Ollama's log and marks the results invalid if it finds a truncated prompt [19]. The author builds Polyglot, the only agent that does not fail on any of the seven models [c1, c6]. Harness, raw results and failed-run transcripts are on GitHub [1]. They replace a first version that used three runs per task and a private harness [22]. No rerun by anyone else is reported in the post.

What to watch

  • An independent rerun of the public harness, particularly the goose, Hermes and opencode columns now that every model runs at a 32k context.
  • Whether pi, Hermes or opencode add parsing for text-form tool calls; their zero cells on qwen2.5-coder are the test.
  • Scores for the released Polyglot 0.13.2 against the 179, 173 and 172 out of 210 of its pre-release builds.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories