Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Generate up to 2K clips with synced stereo audio straight from a MiniMax H3 node graph in ComfyUI — open weights, no watermark, no API.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Gemini Omni
Gemini Omni Video Generator

Nano Banana2
Best Image Generator
Why Creators Choose comfyui minimax h3 in ComfyUI
MiniMax H3 is a general-purpose, omni-modal model that ships as open weights and plugs into ComfyUI through a ready-made comfyui minimax h3 graph. Because it reads text, images, video, and audio inside one shared context, dialogue, effects, and music are modeled alongside the picture instead of being bolted on afterwards. Clips run to roughly 15 seconds at up to 2K and 24fps, and every node setting stays in your hands.
- Sound Built Into Every FrameVoice, effects, and music are generated in the same pass as the picture, so the comfyui minimax h3 output arrives as one MP4 that already lines up.
- Runs Entirely on Your MachineBecause the comfyui minimax h3 weights are public, you can adjust resolution, clip length, and sampling parameters locally without ever hitting an API ceiling.
- Mix Text, Stills, Clips & VoiceFeed several reference types into a single run and pin down a face, a look, a movement, a camera path, or a voice using the comfyui minimax h3 nodes.
Running comfyui minimax h3: A Three-Step Walkthrough
Follow these three steps to produce open-weight video with built-in audio using the comfyui minimax h3 workflow.
Core Capabilities of the comfyui minimax h3 Workflow
Everything needed for local video work lives in this comfyui minimax h3 package: three ready templates, omni-modal inputs, built-in stereo sound, reference locking, and an optional attention patch for faster runs.
Three Ready-Made Templates
Text-to-video, image-to-video, and reference-to-video samples ship with the comfyui minimax h3 package, so each generation mode works the moment you open it.
One Shared Multimodal Context
Rather than handling each modality separately, the comfyui minimax h3 model reads text, pictures, footage, and sound together and blends them into a single result.
Lock Down Characters, Styles, and Motion
Point the comfyui minimax h3 R2V node at up to 9 images, 3 videos, and 3 audio clips to hold a face, a look, a movement, a camera angle, or a voice steady.
Clean On-Screen Text and Logos
Signage, captions, and brand marks come out legible with the comfyui minimax h3 model, which also follows natural-language instructions about how references relate.
Optional Sage Attention Boost
Slot a Patch Sage Attention KJ node into the comfyui minimax h3 graph and generation time drops by roughly half, with only a slight quality trade-off.
Automatic Resolution and Duration Math
The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapped to the 32-pixel grid and 17-frame blocks at 24fps.
Questions About the comfyui minimax h3 Workflow
Answers to the questions people ask most about the comfyui minimax h3 workflow and about running MiniMax H3 inside ComfyUI.
What exactly is the comfyui minimax h3 workflow?
It is ComfyUI's built-in integration for MiniMax H3, an open-weight omni-modal generation model from MiniMax. With it, the comfyui minimax h3 graph turns text, pictures, footage, and audio references into video that carries its own stereo soundtrack, all in one forward pass.
How high can the output resolution go?
The comfyui minimax h3 workflow reaches 2K at 24fps for clips of about 15 seconds. Its native canvas uses a 768-pixel short edge, tops out at 768x1344, and rounds dimensions to multiples of 32.
Which generation modes ship with it?
Three examples come in the box: text-to-video, image-to-video with optional first- and last-frame control, and reference-to-video for locking a character, style, motion, camera path, or voice.
Does the workflow produce audio as well?
It does. The comfyui minimax h3 model builds stereo voice, effects, and music in the same pass as the visuals, and everything lands synced inside a single MP4.
What do I need to get started?
Update ComfyUI to version 0.30.0 or later, open Template Library > Video, select a comfyui minimax h3 template, and follow the pop-up to pull the weights from the Comfy-Org/MiniMax-H3 repository on Hugging Face.
Is there a way to make it run faster?
Yes. Install SageAttention together with the KJNodes custom nodes, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 graph to roughly halve render time.
Create Your First Clip with comfyui minimax h3
Open weights, synced stereo sound, and every knob exposed — MiniMax H3 runs locally in ComfyUI, and the text, image, and reference video graphs are waiting for your first prompt.
