MiniMax H3: the omni model you can actually use today
Updated July 31, 2026
What actually shipped
H3 is MiniMax's third generation of the Hailuo line (Hailuo 01 → Hailuo 02 → H3) and the first sold as a general-purpose omni model rather than a video generator. One network takes a mixed context — text, images, video, audio — and returns video with natively generated stereo audio, up to 15 seconds at 2K.
| MiniMax H3 | |
|---|---|
| Released | July 31, 2026 — live on launch day |
| Resolution | Native 2K (2560×1440) at 24fps; 768p tier |
| Duration | 5–15 seconds |
| Audio | Native stereo, generated with the picture (dialogue, effects, music) |
| Omni-reference | Text + up to 9 images, 3 video segments, 3 audio clips in one request |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Modes | Text-to-video, image-to-video, reference-to-video, V2V motion transfer, instruction editing |
| Open weights | Promised “within the next few days” |
| Parameters | Not disclosed |
Where you can get it — checked July 31, 2026
- Hailuo app — live. The site carries an “H3 is live” banner and exposes an Omni Reference slot in the create flow.
- MiniMax's own platform — the model page offers API access and “Try in MiniMax Hub”.
- fal.ai — three endpoints, published the same day, filed under the name Hailuo 03:
minimax/hailuo-03/text-to-video,/image-to-videoand/reference-to-video. Day-0 availability on the aggregator most Western teams build on is the part that separates this launch from every other one this summer. - Atlas Cloud (reseller) — the same three endpoints, and the only place we found a published 768p rate.
- kie.ai — not yet. Its API rejects every MiniMax model id we tried.
- Hugging Face — no H3 repository on the MiniMaxAI organisation yet, so the open-weights promise is still a promise.
What it costs
Both providers publish a rate, and they agree within a cent:
| Price per second | 15-second clip | |
|---|---|---|
| H3 at 2K (fal.ai) | $0.13 | ≈ $1.95 |
| H3 at 2K (Atlas Cloud) | $0.14 | ≈ $2.10 |
| H3 at 768p (Atlas Cloud) | $0.10 | ≈ $1.50 |
| Seedance 2.0 at 720p (fal) | $0.303 | ≈ $4.55 |
| Veo 3.1 standard | $0.40 | ≈ $6.00 |
MiniMax's own framing is that H3 runs under a third of mainstream per-second pricing at 2K and under half at 768p — and at $0.13 against Veo 3.1's $0.40, that holds. It is the cheapest 2K row in our cost calculator, and it is cheaper at 2K than Seedance 2.0 is at 720p.
The interesting idea: one context, not one task
Every other flagship splits the work into endpoints — text-to-video, image-to-video, style reference, motion reference, video editing. H3 collapses them. You hand it a mixed bag of media and describe the relationship between the piecesin ordinary language. MiniMax's own demo prompt is the clearest illustration:
Reference the Hitchcock camera movement from video 1, make the person in image 2 sing, with the singing voice from audio 3.
No mode switch, no separate motion-transfer model. MiniMax argues that the isolation of tasks and modalities is what limits generalisation, and that language is the bridge that unifies them — so the pre-training set deliberately mixes text-to-image, text-to-video with joint stereo audio and native multi-shot, text-to-audio that does not separate speech, effects and music, and reference/editing pairs across every combination of media.
How they got 2K cheaply
- H3-VAE — a rebuilt tokenizer whose higher compression buys a 4× sequence-length saving. That is what makes native 2K affordable rather than a premium tier.
- In-context Regeneration — instead of bolting on an upscaler, H3 regenerates its own low-resolution result with the original multimodal context still attached. An upscaler guesses at small text; regeneration with context can actually reconstruct it.
- H3-Omni Transformer — MiniMax threw away the Hailuo-02 architecture that had been its advantage, on the argument that architecture tricks should give way to task generalisation. Heterogeneous understanding/generation training lifted throughput by about 30%.
H3 versus FLUX 3
These two landed eight days apart with nearly the same pitch: one model for several modalities, native audio, open weights to follow. The difference is what you can do about it.
| MiniMax H3 | FLUX 3 | |
|---|---|---|
| Availability | Live on launch day, buyable now | Gated early access, apply and wait |
| Pricing | $0.13/s at 2K on fal, day one | None published |
| Max clip | 15s at 2K | 20s at 720p |
| Audio | Native stereo | Native, multilingual dialogue |
| Also generates | Images, audio | Images, audio, robot action predictions |
| Open weights | Promised within days | “Later in 2026”, undated |
Read the other side on our FLUX 3 tracker. If you want the omni-model idea in production this week, H3 is the only one of the two you can pay for.
What to watch next
The weights. MiniMax said days, not months, and it has an incentive to move: an openly licensed 2K video model with native audio would be the first of its kind, and would land before FLUX 3 Dev. We check the MiniMaxAI Hugging Face organisation and the aggregator catalogs every day, and this page changes the same day anything appears.
Frequently asked questions
- Can I use MiniMax H3 right now?
- Yes — and that makes it unusual this summer. We checked on July 31, 2026, the day it launched: it is live in the Hailuo app, MiniMax lists API and MiniMax Hub access on its own model page, and the reseller Atlas Cloud already serves three endpoints (text-to-video, image-to-video, reference-to-video). fal.ai listed all three endpoints the same day under the name “Hailuo 03”; kie.ai has not picked it up.
- How much does MiniMax H3 cost?
- $0.13 per second at 2K on fal.ai, or $0.14 at 2K and $0.10 at 768p on Atlas Cloud — so a 15-second 2K clip is roughly $1.95. References are billed on top at fal: audio is free, the first five images are free, extra images are $0.04 each and a reference video costs another $0.13/s. That makes H3 cheaper at 2K than Seedance 2.0 is at 720p. MiniMax has published no English rate card of its own.
- Is MiniMax H3 open source?
- Not yet. MiniMax says it will release the weights 'within the next few days, subject to the relevant laws and regulations', and says it designed H3 for compatibility with several domestic Chinese accelerators from the start. As of July 31 no H3 repository exists on the MiniMaxAI Hugging Face organisation. If it lands, it would be the strongest openly licensed video model available.
- What is omni-reference?
- Instead of separate modes for image-to-video, style reference and motion reference, H3 takes text, images, video and audio as one context and lets you describe the relationship in plain language. MiniMax's own example prompt: 'reference the Hitchcock camera movement from video 1, make the person in image 2 sing, with the singing voice from audio 3.' Reported limits are nine images, three video segments and three audio clips per request.
- How does H3 compare to Seedance and Veo?
- On price it undercuts both: $0.13/s at 2K on fal against roughly $0.30/s for Seedance 2.0 at 720p on fal and $0.40/s for Veo 3.1 standard. On quality nothing is settled — H3 has not appeared on the Artificial Analysis boards we track, so the ranking claims circulating on launch day are not independently verified. What is genuinely different is the input side: one context of text, image, video and audio instead of one task per endpoint.
- What can't H3 do?
- MiniMax lists the limits itself: multimodal context understanding still has 'huge room to improve' and will be fused with its M-series language models in a later version; model scale is limited and scaling is the stated direction; and picture fidelity in some scenes needs work. Clips also top out at 15 seconds, shorter than Seedance 2.5's promised 30.