Build1 publisherNot yet confirmed elsewhere2 min readPublished
Renaming a Gemma 4 model in Ollama 0.35.1 changes which prompt template it gets
Ollama 0.35.1 picks Gemma 4's prompt template from the model's name, so a copied or renamed 12B model loses four prompt tokens when thinking is off. Teams that save the library model under their own name can pin RENDERER gemma4-large in the Modelfile to keep the template that matched Google's.
The Engineer · Build desk

What happened
- When a name carries no size, Ollama compares the parameter count against a 12,000,000,000 threshold, and the library 12B model's config reports 11.9B.
- An upstream bug report compared 50 rendered prompts against Google's chat template for the 12B: the large renderer matched all 50 and the small one missed the empty thought block on all 50.
- Running ollama show on the renamed and original models prints identical output, including RENDERER gemma4, and Ollama does not log the per-request resolution.
- Requests without a think field, including those through /v1/chat/completions, ran with thinking on and got the same prompt from both templates.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure The teams affected are those who customise the library 12B under their own model name and send think: false; users of the plain library tag are not.
- decision Putting the size in every derived name keeps the large template only until someone copies the model to a size-less name, while a RENDERER line in the Modelfile holds through renames.
- constraint The pin is size-specific: e2b and e4b map to gemma4-small, so a shared Modelfile fragment that pins gemma4-large would put the 12B-class template on the small models.
The choice happens in `resolveGemma4Renderer` in `server/renderer_resolution.go`, according to a write-up on dev.to that traced the 0.35.1 source [3]. It checks the short name, then the full name, then the parameter count from the model config, and falls back to the small renderer [3]. The name check is a substring match on e2b, e4b, 12b, 26b or 31b [3]. The name is the one you asked for, tag included, so `assistant:12b` resolves large [5].
A copy keeps the base model's config, including `RENDERER gemma4` and a `model_type` of 11.9B [5]. The large threshold is exactly 12,000,000,000 [4]. By Ollama's own count, the 12B model is 100 million parameters too small for the 12B template [16].
Both renderers are the same code, and gemma4-large adds one flag, `emptyBlockOnNothink: true` [6]. With thinking off, that flag ends the generation prompt with `<|turn>model\n<|channel>thought\n<channel|>`. Without it, the prompt stops at `<|turn>model\n` [6].
The test is cheap, and I'd reuse it. The author sent "Hello" with `think: false` and `num_predict: 1`, then read `prompt_eval_count` [8]. The library tag counted 14 prompt tokens. `ollama cp` to mygemma counted 10, and the same weights saved as assistant-12b counted 14 [9]. A Modelfile adding a system prompt and num_ctx counted 21 as helper and 25 as helper-12b [10]. Both gaps are four tokens, the length of the empty thought block [17]. With `think: true`, the name made no difference to the count [9][10].
The upstream report, ollama/ollama#18824, used a third-party GGUF and said the official library tag is fine because its name contains 12b [15][13]. The write-up agrees for the tag. It points out that adding a system prompt usually means saving a new model under a new name, and a name like assistant has no size in it [13][9].
On outputs, the evidence is one quantization and six prompts. The upstream report did not measure answers. The write-up ran the 12B Q4_K_M at temperature 0 on CPU, small renderer against large [1][8]. The model wrote the missing block itself, spending four generated tokens [1]. For that result to hold on your workload, your prompts would need to resemble those six, and the 26B and 31B would need to behave like the 12B. Only the 12B was tested [8][1].
In my context, derived models with system prompts and thinking switched off, I'd pin it. A Modelfile with `RENDERER gemma4-large` counted 14 tokens, the same as the library tag [14][9]. The write-up also publishes a script, `check-gemma4-renderer.sh`, that lists which local models resolve to which renderer [14].
What to watch
- Whether Ollama's fix for ollama/ollama#18824 changes the 12,000,000,000 threshold or records the resolved renderer when a model is copied or created.
- Whether ollama show or the server log starts reporting the renderer each request resolves to.
- Thinking-off output comparisons on the 26B and 31B, or on other quantizations of the 12B, beyond the six prompts tested so far.