Build1 publisher3 min readPublished
Blender 5.0 rejects three in ten scripts that ten LLMs wrote for it
Ten LLMs' Blender 5.0 scripts ran only 70% of the time when a Kaggle benchmark executed them in 5.0, against 91% for scripts targeting 3.6. Renamed and removed APIs look like valid code, so the benchmark grades each answer in the exact build the prompt named.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Ten models, chosen to span vendors, sizes and open and closed weights, produced 1,190 graded answers through the Kaggle Model Proxy.
- Compositor, boolean solver and sequencer tasks scored 0% on Blender 5.0 across all ten models, and animation tasks reached 10%.
- gpt-5.5 and gemini-3.8-flash led on 5.0 at 87%, yet the author found them as lost on the restructured 5.0 areas as gpt-oss-120b.
- gpt-oss-120b ran 93% of its scripts on 3.6 but 53% on 4.2; the author says it knows Blender 3.x well and stopped there.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Reviewing generated bpy code by eye or scoring it by text similarity cannot flag a removed attribute, so the script has to run in a pinned build before anyone trusts it.
- decision Upgrading to a larger model does not repair breakage from a release newer than its training data, so teams targeting Blender 5.0 need a check that sits outside the model.
- cost The model side of version-pinned grading cost under a cent per answer, so the real expense of adopting it is maintaining the Blender builds and the per-task assert scripts.
The harness is the part of this benchmark I would copy. The Kaggle task ships four Blender Linux builds as a dataset, extracts them at the start of a run, and launches `blender -b --factory-startup --python` on the build each prompt named [12]. An answer passes only if the model's script, followed by a per-task assert script, exits 0 [6]. The grader was tested before the models were. A hand-written reference answer for every task and version passes in all four builds, 119 of 119 [13]. Two control tasks use APIs that did not change, so ordinary correct code is not punished [10].
Blender's bpy API changes on every major release. Enums are renamed, attributes removed, node sockets moved [1]. `mesh.use_auto_smooth = True` is plausible Python, and it raises `AttributeError` in 4.1 and later [3]. The author wrote that "Text similarity cannot see this" and that "The only judge that counts is the Blender version you asked for" [4][5].
Running the code has a blind spot of its own, and the render-engine task shows it. The engine id was `BLENDER_EEVEE` in 3.6, became `BLENDER_EEVEE_NEXT` in 4.2, and went back to `BLENDER_EEVEE` in 5.0 when the legacy engine was deleted [24]. Every model failed this task on some version [24]. According to the author, models that write `BLENDER_EEVEE` for every version pass on 5.0 by accident, while models that learned the 4.2 rename fail 5.0 with full confidence [25]. Blender renamed the engine and then renamed it back, so on this task the models that kept up with 4.2 are the ones penalised [24][25]. gpt-5.5 was one of them. Its answer to the 5.0 prompt flagged the engine id as a change made in 4.2 [26].
The benchmark scores awareness for this reason. Each answer must end with a WATCH OUT block that names every API changed in the target version and its replacement, or says none [7]. Awareness fell from 81% on 3.6 to 48% on 4.2, 40% on 4.5 and 35% on 5.0 [17]. Scripts that ran held at 81% and 83% on 4.2 and 4.5 [16]. On 5.0, answers that ran outnumber answers that named the changes by 35 points [2].
The failures that start at 4.2 are a different kind. Pooled over 4.2, 4.5 and 5.0, five changes from the 4.0-to-4.2 window caused 71 failures: `use_auto_smooth` removed (19), the legacy OBJ exporter removed (15), `Mesh.calc_normals` removed (14), the EEVEE rename (14), and the Principled BSDF `Specular` socket renamed to `Specular IOR Level` (9) [21][3]. Frontier models mostly get these right and open-weight models mostly do not, according to the author [22]. The 5.0 restructures are another matter. The author wrote that "The training data has not caught up with 5.0 and no amount of scale fixes that" [19].
Those figures come from one setup: 30 tasks covering 33 verified API changes, each asked for 3.6, 4.2, 4.5 and 5.0, with the same system prompt, temperature 0, no tools and no retrieval [8][9][11]. They transfer to a pipeline where a model writes bpy from memory against a release newer than most of the Blender code it has seen [2][11]. The author says most of that code was written for 2.8 to 3.x [2]. The benchmark did not test a model with the 5.0 API documentation in its context [11].
What to watch
- A rerun of the same 119 prompts with the Blender 5.0 API docs supplied would test whether context moves the compositor, boolean and sequencer tasks off 0%.
- Models trained after Blender 5.0 shipped: whether their 5.0 run and awareness rates close on the 3.6 figures.