Skip to content

Build1 publisher2 min readPublished

Ollama lets thinking models skip the JSON schema by answering without thinking

Ollama since 0.34.4 lets Gemma 4 skip the requested JSON schema, returning bare text with HTTP 200 in 8 of 24 test calls. Until the open fix ships in a release, structured output on local thinking models needs a shape check in the client.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Ollama lets thinking models skip the JSON schema by answering without thinking
Generated illustration

What happened

  • Pull request #18479, merged on 2026-09-22 and first released in 0.34.4, folded the format schema and the thinking block into a single grammar.
  • That grammar treats everything before the thinking-close token as unconstrained thinking, so a model that answers directly is never held to the schema.
  • At temperature 0, gemma4:e2b answered "391" after four tokens with done_reason "stop", a JSON number where the schema required an object with an answer field.
  • Sent with think: false, the same requests came back in the right shape every time on both the chat and generate endpoints.
  • In a replay, the open fix #18783 brought all 16 requests into the schema by getting the model to think first every time.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Smoke tests that try a thinking model on reasoning prompts will pass, while the short, easy prompts typical of structured-output calls are the ones that slip the schema.
  • constraint Dashboards built on HTTP status or done_reason will count these replies as successes, so detection has to happen where the reply is parsed.
  • cost Once the fix ships, short structured calls to a thinking model pay for a thinking pass they used to skip, a latency cost on CPU-only hosts like the test LXC.

The bug is described in its own docstring. In 0.35.1, `llm/llama_server.go` checks whether the model's parser reports a thinking-close string and whether the request carries a format. When both hold, it converts the schema to a grammar and wraps it with `thinkingGrammar` [14]. For Gemma 4 the close string is `<channel|>` [16]. The function's docstring in `llm/gbnf.go` ends by noting that a response may end before any closing appears [15].

According to a dev.to write-up that reproduced the upstream report, #18774, the server log had no line about the format when replies came back as bare text [6][2]. `/api/generate` returned the same results as `/api/chat` [17]. A reply that is one word or one number can look like a parsing problem on the client side [18].

The rate depends on the prompt. Sent eight times each at default temperature with seeds 0 to 7, "What is the capital of France? One word." skipped thinking five times, the 17 * 23 question three times, and "Is 91 prime? Answer yes or no." never [11]. On the two short-answer prompts that is half of all calls [2]. Only Gemma 4 was tested: e2b in the write-up, e4b in the upstream report [7]. The write-up's `check-ollama-format-think.sh` reports whether a model on your own server does the same [6].

The `think: false` result makes `think: true` look unsupported with `format`, but the schema is enforced there too, only after the thinking ends [19]. Shape is all the grammar checks. For the temperature-0 request, the `think: false` reply was `{"answer": ")"}` and the reply that thought first was `{"answer": ">391"}` [10][9]. A client check that confirms the object and its `answer` field passes both.

In the same temperature-0 pair, the reply that thought used 248 eval tokens against 4 for the direct answer, 62 times as many [8][9][3]. The write-up does not report token counts for the patched runs. Its test box was a CPU-only Debian 13 LXC running 0.35.1, the current release [7]. On hardware like that, for a one-field classification, I'd take `think: false` and a validator in the client over a thinking pass on every call.

What to watch

  • Whether #18783 merges as written, getting compliance from the model thinking first, or is reworked to constrain a direct answer, and which release carries it.
  • Results from check-ollama-format-think.sh on thinking models other than Gemma 4, since any model whose parser reports a thinking-close string goes through thinkingGrammar.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories