Omni-Reference Video Generation with Native Stereo Audio — 5–15 s @ 24 fps
MiniMax H3 is an omni-modal video generation model from MiniMaxAI: text, image, video, or audio in → video with synchronized 32 kHz stereo audio out. CNAPS Studio runs the H3-Base-Ref2VA checkpoint (omni-reference mode): a text prompt plus reference media — including a reference clip's own soundtrack — generates a new clip whose picture and audio come out of a single denoising loop, so lip movement and sound land in sync. Clips run 5 – 15 seconds at 24 fps across six trained canvas sizes. 33B-parameter dense transformer, ~144 GB of bf16 weights, offloaded component-by-component; expect minutes-scale generation per clip.
Available in all regions through CNAPS Studio. MiniMax H3 is released under the MiniMax H3 Community License Agreement — a community license, not an OSI-approved open-source license. Usage terms apply to what you generate; see License Note below.
📹 See it in action
What it does
MiniMax H3 turns a text prompt plus reference media into a short video clip with built-in synchronized stereo audio — no separate text-to-video and text-to-speech step, everything in one model. That's the headline capability.
CNAPS Studio deploys the omni-reference (Ref2VA) checkpoint. You write a prompt and attach at least one reference — an image, a video clip, or both — and H3 uses them as guidance for subject, style, and, because the reference clip's soundtrack is read as well, voice timbre and sonic character. References are what this checkpoint is built around: a prompt on its own is not enough to drive it. Output runs 5 – 15 seconds at 24 fps in one of six trained canvas sizes, from 21:9 widescreen to 9:16 vertical.
Note on the other checkpoint. MiniMax also publishes H3-Base-FL2VA (first-frame / last-frame / in-between animation). That checkpoint is not deployed in CNAPS Studio — the Video Generation node is Ref2VA only. If you need frame-to-frame animation, see the other models in the Video Generation lineup.
Problem it solves
- Video + audio in one shot — Most open-source video generators produce silent video. H3 generates the video and matching stereo audio (dialogue, sound effects, ambience, music) in the same forward pass — synchronized, no separate audio pipeline needed
- Reference-driven creative direction — Steer generation with mixed references: text describes the shot, an image anchors the subject or style, a video clip carries both motion/style reference and its soundtrack, which can act as a voice timbre reference so a generated character speaks in that voice
- Lip-sync without a second pass — Because picture and sound share one denoising loop, dialogue lands on the mouth movement rather than being aligned afterward
- Bilingual & multilingual dialogue — 11 languages officially supported (English, Chinese, Japanese, Korean, French, German, Italian, Spanish, Portuguese, Russian, Arabic) with additional partial support
- Up to ~14.4-second clips with sound — long enough for a narrative shot with dialogue, not just a 4–5 s micro-clip
Input/Output
Input
The node has three input ports. In CNAPS Studio a filled connection point means required and a hollow one means optional.
# | Port | Marked as | Actually |
1 | Text prompt — describes the shot | Required (filled) | Required |
2 | Reference image — subject / style / object anchor | Optional (hollow) | See below |
3 | Reference video clip — motion and style reference; its own soundtrack is also read as an audio reference | Optional (hollow) | See below |
⚠️ Connect at least one of the two reference ports. The node marks both the image and the video input as optional, but that reflects each port on its own — neither one is individually mandatory. In practice the model will not run correctly on a text prompt alone: you must supply a reference image, a reference video clip, or both.
A text-only run is not a supported path on this node. If you want text-to-video with no references, use one of the other models in the Video Generation lineup.
<sample inputs>
One image, one clip. Upstream, the Ref2VA checkpoint accepts up to 9 reference images, 3 video clips and 3 audio clips (each 2 – 15 s, ≤ 15 s total per modality). The CNAPS Studio node exposes one image port and one video port, and no separate audio port — audio conditioning comes from the reference clip's own soundtrack.
Options
The node exposes four settings — three on the node itself, one under Advanced.
1. Clip length (5 – 15 s at 24 fps) — set in frames:
Frames | Duration |
124 (default) | ~5.2 s |
158 | ~6.6 s |
175 | ~7.3 s |
192 | ~8.0 s |
226 | ~9.4 s |
243 | ~10.1 s |
277 | ~11.5 s |
294 | ~12.3 s |
328 | ~13.7 s |
345 | ~14.4 s |
2. Canvas (W×H) — the model's six trained aspect ratios:
Canvas | Aspect ratio |
1344 × 768 (default) | 16:9 |
768 × 1344 | 9:16 |
1024 × 768 | 4:3 |
768 × 1024 | 3:4 |
768 × 768 | 1:1 |
1536 × 672 | 21:9 |
3. Diffusion steps (20 – 75) — default 30. Changing the step count changes the result itself, not just its level of detail; higher values also take longer. Start with the default.
4. Random seed (Advanced Options) — -1 = random. Set a fixed value to reproduce a result.
Output
An MP4 video with a synchronized 32 kHz stereo audio track — matching the requested frame count and canvas size.
Performance & Deployment
Parameters | 33B (dense transformer — H3-Omni-Transformer); ~13B of those are in AdaLN branches that can be precomputed and cached |
Precision / weight size | BF16, ~144 GB of weights |
Execution in CNAPS Studio | Weights are offloaded component-by-component; expect minutes-scale generation per clip |
Deployed checkpoint | H3-Base-Ref2VA (omni-reference) only |
Required inputs | Text prompt, plus at least one of: reference image, reference video clip |
Output resolutions | Six fixed canvases, 768 × 768 up to 1536 × 672 (see Options) |
Clip duration | 124 – 345 frames (~5.2 – 14.4 s) |
Frame rate | 24 fps |
Audio output | 32 kHz stereo |
Supported dialogue languages | 11 stable (English, Chinese, Japanese, Korean, French, German, Italian, Spanish, Portuguese, Russian, Arabic) + partial support for others |
Upstream serving frameworks | SGLang, vLLM, Hugging Face Diffusers ( ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-H3")), ComfyUI (native templates: T2V, R2V) |
2K path. MiniMax's H3-Regenerate-2K (native 2K regeneration) and H3-Context-IR (advanced prompt understanding) are hosted MiniMax API services, not part of the open-source release. They are not part of this node. Maximum output here is the canvas you select.
How MiniMax H3 compares
MiniMax H3 sits alongside the other video generators in CNAPS Studio's Video Generation lineup. Its distinguishing capability is video + synchronized native audio in one pass — most alternatives produce silent video and require a separate TTS or Foley step. Trade-offs: the community license carries commercial-use restrictions outside CNAPS Studio (see below), and the model is heavy enough that generation runs in minutes rather than seconds.
Cosmos3-Nano | Wan2.2 TI2V 5B | Helios-Base | MiniMax H3 (Ref2VA) | |
Video + native audio | ✅ Yes (optional) | ❌ No | ❌ No | ✅ Yes (native stereo) |
Reference-guided (image + clip) | Partial | Partial | ❌ | ✅ Yes (image + clip, clip audio included) |
First & last frame mode | ❌ | ❌ | ❌ | ❌ (FL2VA checkpoint not deployed) |
Max duration | ~7.9 s | ~5 s | ~45 s | ~14.4 s |
License | OpenMDW 1.1 | Apache 2.0 | Apache 2.0 | Community (restricted) |
All models run inside CNAPS Studio — compare them side-by-side with the Video Compare tool.
License Note (Important)
- License: MiniMax H3 Community License Agreement — this is not an OSI-approved open license
- Regional availability: MiniMax H3 is available to CNAPS Studio users in all regions. If you instead download the open weights (e.g., from Hugging Face) and run them independently, the community license restricts commercial use in the USA / EU / UK / South Korea, and you would need your own approved license from MiniMax via the application form.
- Encoder license: The H3 text encoder reuses Qwen3-VL-32B, which is licensed under Apache 2.0 — that attribution applies to the encoder component
- Different from the rest of the lineup: These terms differ significantly from the permissive licenses on Wan2.2 and Helios, and from the OpenMDW 1.1 license on Cosmos3-Nano; review the full agreement before relying on the model outside CNAPS Studio
Model Source
- Hugging Face
- HF Collection: MiniMax-H3 Collection — 46 items, includes LoRAs, quantizations, and finetunes
- GitHub: https://github.com/MiniMax-AI/MiniMax-H3
- API docs: MiniMax Platform
- Consumer app: Hailuo AI
- License: MiniMax H3 Community License Agreement
- License application form (USA/EU/UK/South Korea): https://platform.minimax.io/h3-license
Compliance & Provenance
Provider | MiniMaxAI (commercial provider with open weights under community license) |
Provider type | Specialized |
License | MiniMax H3 Community License Agreement — not OSI-approved. Encoder (Qwen3-VL-32B) under Apache 2.0 |
Commercial use | Permitted; commercial products/services exceeding USD 20M annual revenue require separate written authorization from MiniMax (§IV.1) |
EU AI Act risk class | Limited Risk |
Art. 50 transparency | Required — outputs are marked with a visible AI-generation label and machine-readable provenance. See AI Policy §2. |
Region availability | Through CNAPS Studio: all regions. Independent open-weights commercial use in USA / EU / UK / South Korea requires the user's own approved MiniMax application. |
Training data summary | Pending — provider has not yet published per Art. 53(d) |
For more on how we classify models and mark outputs, see our AI Policy.