Published Build3 min read
Grok Imagine 2.0 sweeps Ernie 8-0, and the losses are all bookkeeping
Runtimewire's image matchup gave Grok Imagine Image 2.0 every one of eight tasks against Ernie Image Lora Turbo, and the judge repeatedly called Ernie the prettier render while marking it down for a flipped left-right...
Written for builders.See today for builders
What happened
- Grok Imagine Image 2.0 won 8 of 8 tasks against Ernie Image Lora Turbo; Ernie never got on the board.
- Aggregate scores: Ernie Image Lora Turbo 52.9, Grok Imagine Image 2.0 73.3.
- Runtimewire reports 100% confidence in the statistical verdict for the image matchup.
- Method: 8 fresh image tasks generated on the fly so neither model could prepare in advance, scored by gpt-5.4, with every task judged twice, once in each presentation order, and all reported numbers being the average of both passes.
- On the spatial layout task, Grok placed the bed against the left wall, the desk under the back-wall window, the round rug near center and the floor lamp in the front-right corner; Ernie drifted off layout and missed the flat-vector isometric look.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Runtimewire ran eight freshly generated image prompts through Ernie Image Lora Turbo and Grok Imagine Image 2.0, and Grok won all eight, 73.3 to 52.9 in aggregate, with the publisher reporting 100% confidence in the statistical verdict [1][2][3]. That is a 20.4-point spread, or roughly 1.39x [15][16], and the interesting part is that almost none of it was about how the pictures looked.
Read the task-level notes and the failures are countable. In the spatial layout prompt, Grok put the bed against the left wall, the desk under the back-wall window, the rug near center and the floor lamp in the front-right corner; Ernie drifted off the layout and missed the flat-vector isometric requirement [5]. The judge, gpt-5.4, said plainly in both passes that Ernie was the more polished render and docked it anyway for the placement and style misses [6]. In attribute binding, Grok held green cube left of red sphere, blue cylinder behind, yellow duck on top; Ernie inverted the key left-right relationship [7]. In counting, Grok produced exactly nine distinct scarf pins with the requested motifs while Ernie duplicated the lightning bolt and effectively lost the teacup [8]. On the tailor prompt, Ernie put the gloves on the tailor instead of having him measure a glove with both bare hands [12].
Those are not taste disputes. A flipped relation, a duplicated motif, a missing object and a misassigned prop are checkable against the text of the brief. The same pattern extended into prompts that look aesthetic on the surface: Grok held a true close-up focal plane with convincing shallow depth of field on the macro beadwork brief while Ernie delivered a conventional shoe beauty shot [9], respected a negation constraint by producing a reading nook without the forbidden clutter [10], and stayed tighter to a restricted four-color flat-vector palette [11]. Runtimewire's own summary concedes Ernie was sometimes the more conventionally attractive renderer and still calls Grok the model to trust when the prompt matters [13][14].
For contrast, the same desk's video matchup between Kandinsky5 and LTX 2.5 Image to Video Pro landed at 28.6 to 29.8, a 1.2-point gap with only 50% confidence that either model is better, which the publisher called a dead heat [17][20]. That comparison split by domain instead: Kandinsky5 on fluid and particle physics, LTX on crowd choreography and multi-actor staging [19]. Two matchups, one judge, one protocol [4][18]. When the axis is instruction compliance, the separation was total; when it was capability mix, there was no ranking to be had.
Caveats worth holding onto: eight tasks is a small board, the scoring is a single language model rather than a human panel, and the protocol's defense against position bias is judging each task twice in swapped order and averaging [4]. Watch whether the adherence gap survives a larger prompt set and human scoring, whether Ernie's polish edge shows up as a win on loose creative briefs where there is no constraint to violate, and whether either vendor starts publishing per-constraint pass rates rather than aggregate scores. Counting, binding and negation are the parts of this that a buyer can test in an afternoon.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Grok Imagine Image 2.0 won 8 of 8 tasks against Ernie Image Lora Turbo; Ernie never got on the board.
- [2]
Aggregate scores: Ernie Image Lora Turbo 52.9, Grok Imagine Image 2.0 73.3.
- [3]
Runtimewire reports 100% confidence in the statistical verdict for the image matchup.
- [4]
Method: 8 fresh image tasks generated on the fly so neither model could prepare in advance, scored by gpt-5.4, with every task judged twice, once in each presentation order, and all reported numbers being the average of both passes.
- [5]
On the spatial layout task, Grok placed the bed against the left wall, the desk under the back-wall window, the round rug near center and the floor lamp in the front-right corner; Ernie drifted off layout and missed the flat-vector isometric look.
- [6]
In both judge passes of the spatial layout task the judge stated that Ernie (Model A) was more polished visually, but marked it down for missing the flat-vector requirement and for placing the bed away from the left wall.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- runtimewire.comRuntimeWire StaffAug 13Head to head: Ernie Image Lora Turbo vs Grok Imagine Image 2.0
Cited in this coverage: runtimewire.com
Cited in this coverage: runtimewire.com, judge gpt-5.4
- runtimewire.comRuntimeWire StaffAug 13Head to head: Kandinsky5 vs LTX 2.5 Image to Video Pro

