Published Build3 min read
Prompt obedience, not polish: Luma Uni-1 Edit Max 69.7, Ernie Image Lora 56.5
Across eight fresh image tasks judged twice by gpt-5.4, Luma won six and lost none. Ernie Image Lora made attractive pictures that ignored counting, negation and layout instructions.
Written for builders.See today for builders
What happened
- In a head-to-head matchup across eight image tasks, Luma Uni-1 Edit Max scored 69.7 and Ernie Image Lora scored 56.5.
- The task ledger was Luma Uni-1 Edit Max 6 task wins, Ernie Image Lora 0, with 2 ties.
- The aggregate score gap between the two models was 13.2 points.
- The review describes the result as a 99% confidence win for Luma Uni-1 Edit Max.
- Method: 8 fresh image tasks were generated on the fly for the matchup so neither model could prepare in advance; gpt-5.4 scored each one; every task was judged twice, once in each presentation order, to cancel position bias; every reported number, including headline totals, is the average of both passes.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
In a head-to-head run published by runtimewire.com, Luma Uni-1 Edit Max scored 69.7 against Ernie Image Lora's 56.5 across eight image tasks [1], and took the task ledger 6 wins to 0 with 2 ties [2]. The aggregate spread of 13.2 points [3] is the less interesting number; a shutout on individual tasks is harder to explain away as house style.
The reviewer reports the result as a win at 99% confidence [4]. The method: eight tasks generated on the fly for the matchup so neither model could prepare, scored by gpt-5.4, with every task judged twice in swapped presentation order and every reported figure the average of both passes [5].
The named wins are about following instructions, not looking good. According to the review, Luma took perspective and scale with a clean one-point-perspective shot down an empty library aisle, negation by actually excluding the forbidden objects, exact counting with a true overhead seven-cup flat lay, and the Marshgate shelf tag by delivering the requested Swiss-style placard and line breaks rather than improvising [6]. It also handled the rainy citrus scooter ad better, landing closer to the specified compact electric scooter, branding and flat-vector treatment [7].
Ernie's problem, in the reviewer's words, is that it often looked good while being wrong [8]. Its library frame was attractive but put a large foreground book and table in the aisle, which breaks an empty-aisle brief [9]. The cozy reading nook ignored negation constraints and left visible wall art and shelving in frame [10]. The shelf tag drifted into a supermarket scene and collapsed the typography [11]. The recurring shape of the near-misses was the same: right count but wrong camera angle, good mood but weaker context, polished composition missing the specified object form [12].
The two ties read as genuine draws rather than narrow losses. On hands and anatomy the judges split on which image better conveyed the bracelet-tying action, which the reviewer takes as neither model fully locking the assignment [13]. On missed-train commuter, Ernie was stronger on narrative props and context while Luma was stronger on subtle facial emotion and cinematic portraiture [14]. The kettle task also split across passes, one preferring Ernie's simpler stovetop product read and the other Luma's cleaner palette and studio restraint [15]. Since five Luma wins and both ties are named individually out of eight tasks, the kettle is the sixth win on the two-pass average, and it is the one the reviewer flags as not effortless [16]. The final call is that Luma is the model to trust when the prompt actually matters [17].
What to watch: eight tasks and one judge model is a thin base, so the question is whether Ernie's counting, negation and typography failures reproduce on a larger set or were sampling luck. The failure modes here are the cheap ones to verify in-house, since object counts, exclusion lists and text blocks are checkable without a judge model at all. Also worth watching is whether a single order swap is enough to neutralise position bias when two of eight tasks flipped between passes [15][13].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
In a head-to-head matchup across eight image tasks, Luma Uni-1 Edit Max scored 69.7 and Ernie Image Lora scored 56.5.
- [2]
The task ledger was Luma Uni-1 Edit Max 6 task wins, Ernie Image Lora 0, with 2 ties.
- [4]
The review describes the result as a 99% confidence win for Luma Uni-1 Edit Max.
- [5]
Method: 8 fresh image tasks were generated on the fly for the matchup so neither model could prepare in advance; gpt-5.4 scored each one; every task was judged twice, once in each presentation order, to cancel position bias; every reported number, including headline totals, is the average of both passes.
- [6]
Luma won on perspective and scale with the empty one-point-perspective library aisle, on negation by actually excluding forbidden objects, on exact counting with a true overhead seven-cup flat lay, and on the Marshgate shelf tag by delivering the requested Swiss-style placard and line breaks instead of improvising.
- [7]
Luma also handled the rainy citrus scooter ad better, getting closer to the requested compact electric scooter, branding, and flat-vector treatment.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- runtimewire.comRuntimeWire StaffAug 12Head to head: Ernie Image Lora vs Luma Uni-1 Edit Max
Cited in this coverage: runtimewire.com head-to-head review
Cited in this coverage: runtimewire.com
Cited in this coverage: runtimewire.com methodology note
Cited in this coverage: runtimewire.com, judge rationale for task 1

