Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Produce open-weight clips in ComfyUI via the comfyui minimax h3 workflow, using text, image, or reference prompts with synchronized stereo audio at 2K/24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit
Lyrics to Song
Turn your lyrics into professional songs with AI

Suno 5.5
Create Professional Music with AI

AI Music Generator
Create music from videos with AI

AI Song Generator
Create songs from videos with AI

Suno AI Music Generator
Create Professional Music with AI
Seedance2.0
The Future of AI Video Is Here.
Happy Horse 1.0

Veo3.1
Create Stunning Videos with Veo3.1
What You Get from the comfyui minimax h3 Workflow
The comfyui minimax h3 workflow packages MiniMax's open-weight, omni-modal model for ComfyUI. It processes text, imagery, motion, and sound within one unified context, so the generated clip arrives with stereo audio already in sync — speech, effects, and score come from the same run. Output can reach roughly 15 seconds of 2K video at 24fps, and every node remains fully editable.
- Synced Stereo Sound IncludedSpeech, effects, and music are created alongside the visuals and encoded into the same MP4, so everything lines up without extra work inside the comfyui minimax h3 workflow.
- Local Execution with Open WeightsWith open weights you can run the comfyui minimax h3 model on your own machine and tune resolution, length, or any diffusion parameter—no API quotas or remote service limits.
- Multiple Reference Types in One PassPass a blend of text, pictures, clips, and audio into the same generation; the comfyui minimax h3 nodes can hold a character, style, motion, camera angle, or voice consistent throughout.
Three Steps to Create with the comfyui minimax h3 Workflow
Follow this quick guide to produce open-weight clips with synchronized sound through the comfyui minimax h3 workflow.
Everything Built Into the comfyui minimax h3 Workflow
The comfyui minimax h3 workflow brings together ready-made ComfyUI templates, open-weight omni-modal generation, synchronized stereo audio, reference-based control, and optional Sage Attention speedups—a full local video toolchain.
Three Templates, Ready to Use
ComfyUI's template library includes a comfyui minimax h3 example for text-to-video, one for image-to-video, and one for reference-to-video, so each mode has a working starting point.
All Modalities in One Context
The comfyui minimax h3 model evaluates text, photos, video, and audio as a single combined context, letting you mix any of them inside the same generation.
Control via References
Use up to 9 images, 3 videos, and 3 audio files through the comfyui minimax h3 R2V node to fix a character's identity, visual style, movement, camera motion, or vocal tone.
Sharp Text and Brand Rendering
The comfyui minimax h3 model handles spelled-out text and logos crisply, and it follows natural-language instructions that describe how your reference materials relate to each other.
Faster Output with Sage Attention
Insert the Patch Sage Attention KJ node into the comfyui minimax h3 workflow and you can roughly double rendering speed while keeping quality virtually intact.
Precise Resolution and Duration Grid
The comfyui minimax h3 Resolution Selector works from aspect ratio and megapixels to calculate width and height, snapping to the 32-pixel grid and the 17-frame blocks at 24fps for duration.
Quick Answers on the comfyui minimax h3 Workflow
Straightforward responses to common issues around using the MiniMax H3 open-weight model within ComfyUI.
How does the comfyui minimax h3 workflow work?
It's ComfyUI's native integration for MiniMax H3, an omni-modal model released as open weights. Text, image, video, and audio references can all be combined in one run, and the comfyui minimax h3 workflow outputs a synchronized clip with stereo sound.
What resolution and frame rate can I expect?
The comfyui minimax h3 workflow can deliver up to 2K resolution at 24fps for about 15 seconds. Internally it works from a 768px short edge, caps the canvas at 768x1344, and rounds dimensions to multiples of 32.
What types of generation modes come with the templates?
The comfyui minimax h3 template collection includes text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) for locking character, style, motion, camera, or voice.
Does the comfyui minimax h3 workflow create audio?
Yes. Voice, sound effects, and music are synthesized together with the imagery through the comfyui minimax h3 model and stored in one MP4, so the soundtrack stays in sync.
How do I start using it?
Update ComfyUI to 0.30.0 or newer, open the Template Library, choose a comfyui minimax h3 workflow under Video, and download the models from Hugging Face's Comfy-Org/MiniMax-H3 repo when prompted.
Is there a way to make generation faster?
Yes. Install SageAttention and KJNodes, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider inside the comfyui minimax h3 workflow—this can roughly double the rendering speed.
Turn Your Ideas into Video with the comfyui minimax h3 Workflow
Start the comfyui minimax h3 workflow to create open-weight clips with synchronized stereo audio in ComfyUI—T2V, I2V, and R2V templates let you control every parameter and generate locally.
