MiniMax H3 Hits ComfyUI on Day 0: Open Weights, Native Audio, 2K Video

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

MiniMax H3 Hits ComfyUI on Day 0: Open Weights, Native Audio, 2K Video

MiniMax's third-generation video model, H3, is now natively supported in ComfyUI from day one. The open-weights model generates video with real stereo sound, up to 2K resolution and 15 seconds per clip, from text, images, video, or audio. It excels at multimodal context understanding, motion transfer, and editing. ComfyUI's optimizations, including pruning modulation weights and int8 quantization, cut memory usage by 66%, enabling local inference on a GPU like the RTX 3060.

This is a next-generation open-weights video model. Feed it text, images, video, or audio and it generates video with real stereo sound, up to 2K, up to 15 seconds a clip.
  1. embedding-shape

    > We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.

    Is this a common approach to reducing weights with "no loss in output quality", assuming this is true? Seems almost too simple to work. If this is doable, would this be applicable to LLMs as well?

    Neat with native frame-to-frame generation, but wonder how easy it is to "link" together clips at the intersection, typically the models kind of lose the "momentum" across these stiches, being able to merge things with frame-to-frame between clips might help with this it feels like.

  2. vblanco

    Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.

  3. sheesdev

    The mouse render is surprisingly good. Several of those clips stood out to be a pretty big leap in terms of current SOTA models.

    The only one that looks "off" is the beverage ad video during the can opening clip, it still has that "AI smoothening" effect. Good thing this can be done pretty well using traditional rendering.

    I feel like for a good while now we'll transition into a process that uses traditional "close-up" rendering/shots + AI generated wide-shots or quick cuts.

    Exciting, but also troubling. This being open-weights is a massive win for the community though.

  4. Mashimo

    > The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.

    Pretty cool.

    But assuming you have a 16GB 3060, how long would it take to generate a 15 second clip?

  5. fodkodrasz

    On one hand: impressive.

    On the other hands aesthetically it all looks painfully bland and generic.

More from this day

2026-08-03