Published Build3 min read
Prompt fidelity beats prettiness: MiniMax H3 sweeps Kandinsky5 four tasks to zero
A four-task head-to-head scored MiniMax H3 Reference-to-Video 34.6 against Kandinsky5's 22.4. The losing model sometimes looked better, which is the useful part of the result.
Written for builders.See today for builders
What happened
- Runtimewire published a head-to-head comparison of Kandinsky5 and MiniMax H3 Reference to Video, judged on whether the model actually delivers the shot the prompt asked for; the publisher's final call was MiniMax H3 Reference to Video as clear winner for faithful scene construction, specific camera behavior, and keeping multiple moving elements coherent.
- The test used 4 fresh video tasks, generated on the fly for the matchup so neither model could prepare in advance, with gpt-5.4 scoring each one.
- To cancel position bias, every task was judged twice, once in each presentation order, and every number reported including the headline totals is the average of both passes.
- Overall scores were 34.6 for MiniMax H3 Reference to Video and 22.4 for Kandinsky5.
- MiniMax H3 Reference to Video took four task wins to zero and was reported as ahead at 97% confidence.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Runtimewire scored MiniMax H3 Reference to Video at 34.6 against Kandinsky5's 22.4 across four video tasks, with four task wins to zero and, by the publisher's account, MiniMax ahead at 97% confidence [4][5]. The reason this matters more than another leaderboard row is the shape of the loss: Kandinsky5 was not beaten on looks, and in several clips it was reasonably attractive and sometimes a touch steadier [6].
The judging criterion was whether the model actually delivered the shot that was asked for [1]. On the crowd task, the prompt specified a busy Tokyo scramble crossing seen from above, dozens of pedestrians crossing in different directions, each moving independently without merging or warping into one another, overcast daylight, 16:9 [7]. MiniMax produced a dense, overhead, legible crowd flow; Kandinsky5 produced a sparser, more frontal street crossing that misses the defining geometry and energy of a scramble [8]. The judge noted that Kandinsky5 was the temporally steadier of the two on that clip and still adhered less well to the requested scene [9]. Stability on the wrong composition is not partial credit.
The same substitution pattern shows up across the other three. On a lighting-transition prompt, MiniMax delivered the requested locked-off living room, a believable dusk shift, and a lamp turning on near the end, while Kandinsky5 drifted into something closer to an outdoor patio and never executed the key lighting beat [10]. The Moon Jelly Orbit prompt is the most demanding of the set: one continuous 16:9 shot, a translucent moon jellyfish pulsing upward beside a rusted research buoy wrapped in fine green algae, spiralling bubbles, a lone pipefish in the background, and a slow exact camera orbit holding a steady radius with no drift into a dolly or pan, no cuts [11]. MiniMax got closer on the subject-centric orbit, the jelly detail, the caustic aquarium mood, and more of the specified staging; Kandinsky5's underwater imagery was decent but read as a looser pan through a related scene [12]. On Reedbed Egret Tangle the failure was countable rather than atmospheric: MiniMax held the requested cluster of birds and dragonflies with coherent independent motion, while Kandinsky5 underdelivered on animal count and showed shakier subject continuity [13].
That is the operational point. A model that returns an adjacent idea cannot be corrected by prompt tuning, because the prompt was already correct; a model that returns the right blocking with imperfect texture can be re-rolled or upscaled. The margin here is 12.2 points, roughly 1.54 times Kandinsky5's total [15][16].
The caveats are real and the publisher states them. Four tasks were generated fresh so neither model could prepare, and gpt-5.4 did the scoring [2]. Each task was judged twice with presentation order swapped, and every reported number is the average of both passes [3]. That controls position bias; it does not turn four samples into a benchmark, and it makes one model's taste the arbiter of another's. The available writeup also reproduces full prompt text for two of the four tasks and describes the rest in prose [17].
Worth watching: whether the fidelity gap survives a larger and more varied task suite, whether Kandinsky5's steadiness advantage becomes worth something in reference-image workflows where composition is fixed outside the prompt, and whether these margins move when the judging model is next updated.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Runtimewire published a head-to-head comparison of Kandinsky5 and MiniMax H3 Reference to Video, judged on whether the model actually delivers the shot the prompt asked for; the publisher's final call was MiniMax H3 Reference to Video as clear winner for faithful scene construction, specific camera behavior, and keeping multiple moving elements coherent.
- [2]
The test used 4 fresh video tasks, generated on the fly for the matchup so neither model could prepare in advance, with gpt-5.4 scoring each one.
- [3]
To cancel position bias, every task was judged twice, once in each presentation order, and every number reported including the headline totals is the average of both passes.
- [4]
Overall scores were 34.6 for MiniMax H3 Reference to Video and 22.4 for Kandinsky5.
- [5]
MiniMax H3 Reference to Video took four task wins to zero and was reported as ahead at 97% confidence.
- [6]
According to the review, Kandinsky5 did not lose because it looks bad: in several clips it was reasonably attractive and sometimes even a touch steadier.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- runtimewire.comRuntimeWire StaffAug 12Head to head: Kandinsky5 vs MiniMax H3 Reference to Video
Cited in this coverage: runtimewire.com
Cited in this coverage: gpt-5.4 judge notes as published by runtimewire.com

