Build1 publisher3 min readPublished
Ornith-1.0's benchmarks are fine. Ollama can't parse its tool calls.
A 9B MIT-licensed coding model that reportedly matches a 31B rival on SWE-Bench Verified is still unusable as a Claude Code backend, because the runtime never turns its tool-call XML into a file write.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Ornith-1.0 is a June 2026 open-weight coding model; the author could not get it running as a local Claude Code backend.
- During training Ornith first proposes a plan for how to approach each task, then solves the task using that plan; both the plan and the answer are scored, so the model learns to write better plans as well as better answers.
- On paper the 9B Ornith model matches or beats Gemma4-31B, four times its size, on SWE-Bench Verified and Terminal-Bench 2.1; it is MIT-licensed and runs on a laptop.
- 31B divided by 9B is approximately 3.4, not 4, so the source's 'four times its size' is a rounded-up parameter ratio.
- The setup was: Ollama as runtime with the model built from a raw HF GGUF because the registry pull was blocked by Zscaler on the author's network; LiteLLM as a local proxy with an Anthropic-format /v1/messages front end; Claude Code pointed at the local proxy instead of api.anthropic.com.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer writing on dev.to tried to stand up Ornith-1.0, a June 2026 open-weight coding model, as a local Claude Code backend, and stopped short: the model emitted text shaped exactly like a file-write tool call, Claude Code displayed it, and the file never appeared on disk [1][10]. That failure sits in the runtime, not the weights, and it is the whole story of local agentic coding right now.
The model's pitch is genuinely different. Instead of being trained inside a harness a human already built, Ornith proposes a plan for each task during training, solves the task using that plan, and gets scored on both, so plans and answers improve together [2]. The 9B version is MIT-licensed, runs on a laptop, and is claimed to match or beat Gemma4-31B on SWE-Bench Verified and Terminal-Bench 2.1 [3]. The write-up calls that four times the size; 31 divided by 9 is about 3.4, so treat the multiplier as rounded in the model's favour [4].
None of that survives contact with the plumbing. The stack was Ollama running a model built from a raw Hugging Face GGUF because a registry pull was blocked by Zscaler, a local LiteLLM instance exposing an Anthropic-format /v1/messages front end, and Claude Code pointed at the proxy [5]. Two bugs were afternoon work. Building from a raw GGUF drops the chat-template metadata the registry manifest carries, so Ollama falls back to bare prompt passthrough with no turn boundaries, and the model answers, then keeps answering, on a loop; the fix was a hand-written ChatML template plus repeat_penalty 1.3 [7]. The im_start and im_end pairs are the turn boundaries the raw GGUF was missing, and the tool_call tags in that template are Hermes-style, matching what Ornith actually generates [8]. Then Claude Code failed to recognise the custom model ID and requested extended thinking on every call, which Ollama rejected outright because the Modelfile never declared that capability; MAX_THINKING_TOKENS=0 closed it [9].
The third one is structural. Ornith is Qwen-derived, so it speaks Qwen's Hermes-style XML tool-call format, which vLLM consumes with a single flag, --tool-call-parser qwen3_xml [11]. Ollama can render that template outbound perfectly well and has nothing built in to parse a tool_call block back into an executable call [12]. The schema injection and the parse-back logic have to be hand-written into the Modelfile template [11]. This is not an artefact of the manual GGUF build: the author points to a GitHub issue on ollama/ollama showing the same raw, unparsed XML on the official registry tag, and says other Ornith write-ups list it as the most common complaint about running the model on Ollama [13].
Worth noting what did not go wrong. The author has run Gemma4, Qwen2 and Qwen3 locally before, and says none stuck for reasons of trust rather than inference speed [6]. Ornith adds a second failure mode: not a wrong answer you can catch, but a correct-looking action that never happened.
What to watch: whether Ollama ships a named tool-call parser per family, the way vLLM already does [11]. Until it does, a benchmark score on a 9B model tells you what it could do inside a harness that can execute, and local Ollama is not that harness [12].