Skip to content

Build1 publisher3 min readPublished

H3's reference path is a different checkpoint, capped at 12 files, and stops at 768p

MiniMax's open-weight H3 base generates at 768p, and the 2K stage is not in the release. Pipelines validated on text-to-video demos will meet ceilings the demos never showed.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A dev.to guide covers MiniMax H3 beyond basic generation, arguing that most H3 tutorials stop after text-to-video or image-to-video works once, which proves the model runs but misses reference-to-video (R2V).
  • R2V uses a different checkpoint from T2V and I2V; in ComfyUI, use the ref2va weights for reference-driven generation, not the fl2va weights.
  • H3-Base generates at 768p.
  • The official ComfyUI integration requires version 0.30.0 or later and provides T2V, I2V, and R2V templates under Template Library -> Video.
  • MiniMax's separate H3-Regenerate-2K stage produces 2K output, but that module is not currently part of the open-weight release.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A ComfyUI walkthrough published on dev.to lays out the operating constraints on MiniMax H3's reference-to-video path, and they are not the constraints a text-to-video demo exercises [1]. Two of them decide whether a pipeline survives a production brief: reference-driven generation runs on a different checkpoint than T2V and I2V, and the open-weight base tops out at 768p [2][3].

Start with the weights. According to the guide, reference work uses the `ref2va` checkpoint, not the `fl2va` weights that back text-to-video and image-to-video [2]. A graph that proved out T2V is therefore not a graph you extend into R2V by adding inputs; you are loading a second model. The official ComfyUI integration requires version 0.30.0 or later and ships T2V, I2V, and R2V templates under Template Library -> Video [4].

Resolution is the harder ceiling. H3-Base generates at 768p [3]. MiniMax's H3-Regenerate-2K stage produces 2K output, but the guide states that module is not currently part of the open-weight release [5]. If your deliverable is 2K, the open weights get you a 768p master and nothing else; the finishing step is not yours to run.

The input ceilings look generous until you multiply them. H3-Base-Ref2VA accepts up to 9 images, up to 3 video clips, up to 3 audio clips, and no more than 12 files in total [6]. The per-type maxima sum to 15, which is three above the total cap, so no configuration hits all three limits at once [7]. Use all 3 video and all 3 audio slots and you have 6 image slots left, not 9 [8]. Duration is capped twice: each video or audio clip must run between 2 and 15 seconds, and the total for each media type cannot exceed 15 seconds [9]. Three video references therefore average 5 seconds each at best [10]. The target clip itself is a 4 to 15 second shot [11].

Then there is the failure mode nobody catches in review. ComfyUI identifies references by the order in which they are connected, so `<Picture 1>` in the prompt must be the first connected image [12]. Mislabel that and the prompt still runs; it just applies the wardrobe reference to the face. Any batch harness that reorders inputs is silently reassigning roles.

The editing mode is where the guide is most useful, because it treats preservation as something you specify rather than hope for. R2V will take a source video and apply a local change: replacing a character or object, swapping a background, relighting, style transfer, a localized effect, or keeping versus replacing the original audio [13]. The recommended prompt carries a retention contract, marking composition, motion, timing, or audio as `fully_preserved` and writing `N/A` where a section does not apply, rather than leaving intent ambiguous [14]. The guide is explicit about why: without that structure the model redesigns the whole shot when only one attribute should change [15]. With an audio reference, H3 generates new speech that follows a speaker's vocal characteristics while producing new audio for the target scene [16].

What to watch: whether MiniMax releases H3-Regenerate-2K weights, since that single decision determines whether an open-weight pipeline can deliver above 768p without an upscaler you source yourself [5][3]. Also worth pinning is the ComfyUI version, given the 0.30.0 floor for the official integration [4]. Anyone scoping R2V work should cost the 12-file, 15-second-per-type budget before promising a shot list, because filling every slot does not reliably improve the result [17].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories