Build1 publisher3 min readPublished
Coding agents that rebuild photos as Blender code judge their own geometry near or below chance
LEGO-Bench, from the University of Maryland and AWS, scores the best coding agent at 53.4 percent on indoor scenes rebuilt from photos in Blender code. The agents cannot tell when an edit helped, so their revise loop needs a measured score to decide which edits to keep.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- LEGO-Bench holds 208 images drawn from 104 indoor and outdoor scenes and uses 443 registered assets.
- All six GPT configurations tested returned a working Blender scene almost every time.
- Raising the reasoning budget lifted GPT-6 Astra's score on an office test subset from 32.3 to 61.8 percent.
- Read straight from the scene programs, object detection reached roughly half the performance of the specialized DINO model, with wider gaps on segmentation and depth.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Accuracy on this task rises with reasoning budget, so the per-scene inference bill becomes part of what a team pays to reach a given score.
- decision Measurement-based review helped weak agents most and the top model little, so a cheaper model with LEGO-Plugin is the setup to price against the top model alone.
- constraint Scene programs cannot yet stand in for dedicated detection, segmentation or depth models, so a pipeline that wants queryable 3D from one photo still has to run those models too.
Each pass through the LEGO-Anything loop has four steps: the agent writes Blender code, runs it, looks at the render, and revises [2]. The output is a program, so objects, geometry, layout and camera position are all explicit and editable [3]. Writing and running are mechanical. Looking and revising depend on a judgement the model makes alone: is the new scene closer to the photo than the old one?
The researchers tested that judgement directly. The models were shown two versions of a scene and asked which better matched the original. Their geometric picks landed near or below chance [12].
A keep-or-revert step that performs at chance keeps bad edits about as often as good ones. Iterating on that basis mostly turns tokens into scene files. The step analysis fits that picture. The most common problems were poor first attempts, revisions that undid earlier progress, and unreliable self-assessment [11].
LEGO-Plugin takes that decision away from the model and needs no extra training [14]. It anchors the starting scene in the reference image, replaces self-judgement with concrete measurements, and shields correct progress from regressive edits [14]. This is good engineering. It gives the step the model fails at to something that can be computed. All six models improved. Weaker agents gained up to 62.7 percent, and GPT-6 Astra, the top model, gained about two percentage points [15]. One figure is a percentage and the other is in points, so they do not compare directly. I think the small gain at the top suggests Astra's remaining errors mostly come from somewhere other than accepting bad edits.
The results do not show that a person must check every scene. The researchers conclude that refinement should rely on concrete measurements instead of the agent's own judgement [13]. Outside the benchmark, the only reference those measurements can use is the single input photo [1].
The benchmark design is careful. Real photos have no precise 3D ground truth, and simple synthetic scenes look unrealistic, the researchers say [5]. So LEGO-Bench renders its inputs from professionally built simulator scenes. It keeps the exact geometry, depth and object assignments hidden as the answer key [5].
That choice also limits how far the scores travel. Astra scored 39.6 percent outdoors [8], 13.8 points below its indoor result [1]. Weaker configurations sat around 15 percent [9], and accuracy dropped as scenes grew more complex [18]. For those numbers to hold on someone else's workload, the photos would need to resemble the simulator scenes. The agent would also need to be one of the GPT configurations tested [7].
The larger reasoning budget bought a 29.5-point gain on the office subset [2]. The write-up does not report what the extra reasoning cost.
What to watch
- LEGO-Bench runs on non-GPT model families, to test whether near-chance self-judgement holds outside the six GPT configurations.
- A published specification of LEGO-Plugin's concrete measurements, and a test of them on real photographs with no simulator answer key.
- Token or cost figures for the larger reasoning budget behind Astra's 61.8 percent office score.