VidModelHub

MiniMax H3: the omni model you can actually use today

Updated July 31, 2026

ReleasedAnnounced and shipped on July 31, 2026. 2K video with native stereo sound, text, image, video and audio read as one context — and open weights promised within days. After a summer of models nobody outside a partner list could touch, this one went live the same morning it was announced.

What actually shipped

H3 is MiniMax's third generation of the Hailuo line (Hailuo 01 → Hailuo 02 → H3) and the first sold as a general-purpose omni model rather than a video generator. One network takes a mixed context — text, images, video, audio — and returns video with natively generated stereo audio, up to 15 seconds at 2K.

MiniMax H3
ReleasedJuly 31, 2026 — live on launch day
ResolutionNative 2K (2560×1440) at 24fps; 768p tier
Duration5–15 seconds
AudioNative stereo, generated with the picture (dialogue, effects, music)
Omni-referenceText + up to 9 images, 3 video segments, 3 audio clips in one request
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16
ModesText-to-video, image-to-video, reference-to-video, V2V motion transfer, instruction editing
Open weightsPromised “within the next few days”
ParametersNot disclosed

Where you can get it — checked July 31, 2026

  • Hailuo app — live. The site carries an “H3 is live” banner and exposes an Omni Reference slot in the create flow.
  • MiniMax's own platform — the model page offers API access and “Try in MiniMax Hub”.
  • fal.ai — three endpoints, published the same day, filed under the name Hailuo 03: minimax/hailuo-03/text-to-video, /image-to-video and /reference-to-video. Day-0 availability on the aggregator most Western teams build on is the part that separates this launch from every other one this summer.
  • Atlas Cloud (reseller) — the same three endpoints, and the only place we found a published 768p rate.
  • kie.ai — not yet. Its API rejects every MiniMax model id we tried.
  • Hugging Face — no H3 repository on the MiniMaxAI organisation yet, so the open-weights promise is still a promise.

What it costs

Both providers publish a rate, and they agree within a cent:

Price per second15-second clip
H3 at 2K (fal.ai)$0.13≈ $1.95
H3 at 2K (Atlas Cloud)$0.14≈ $2.10
H3 at 768p (Atlas Cloud)$0.10≈ $1.50
Seedance 2.0 at 720p (fal)$0.303≈ $4.55
Veo 3.1 standard$0.40≈ $6.00

MiniMax's own framing is that H3 runs under a third of mainstream per-second pricing at 2K and under half at 768p — and at $0.13 against Veo 3.1's $0.40, that holds. It is the cheapest 2K row in our cost calculator, and it is cheaper at 2K than Seedance 2.0 is at 720p.

Read the reference pricing too
fal bills references separately: audio references are free, the first five reference images are free, each additional image is $0.04, and a reference video costs another $0.13 per second at 2K. A 10-second omni-reference shot with a 5-second reference clip is therefore about $1.30 + $0.65, not $1.30. MiniMax has published no English rate card of its own, so these are the prices you can actually buy at today.

The interesting idea: one context, not one task

Every other flagship splits the work into endpoints — text-to-video, image-to-video, style reference, motion reference, video editing. H3 collapses them. You hand it a mixed bag of media and describe the relationship between the piecesin ordinary language. MiniMax's own demo prompt is the clearest illustration:

Reference the Hitchcock camera movement from video 1, make the person in image 2 sing, with the singing voice from audio 3.

No mode switch, no separate motion-transfer model. MiniMax argues that the isolation of tasks and modalities is what limits generalisation, and that language is the bridge that unifies them — so the pre-training set deliberately mixes text-to-image, text-to-video with joint stereo audio and native multi-shot, text-to-audio that does not separate speech, effects and music, and reference/editing pairs across every combination of media.

How they got 2K cheaply

  • H3-VAE — a rebuilt tokenizer whose higher compression buys a 4× sequence-length saving. That is what makes native 2K affordable rather than a premium tier.
  • In-context Regeneration — instead of bolting on an upscaler, H3 regenerates its own low-resolution result with the original multimodal context still attached. An upscaler guesses at small text; regeneration with context can actually reconstruct it.
  • H3-Omni Transformer — MiniMax threw away the Hailuo-02 architecture that had been its advantage, on the argument that architecture tricks should give way to task generalisation. Heterogeneous understanding/generation training lifted throughput by about 30%.
Unverified on launch day
Ranking claims are circulating — top of the video-editing board, second in text-to-video with audio. H3 does not appear on the Artificial Analysis boards we track as of July 31, so we cannot confirm any of it. MiniMax also reports its own limits: context understanding with “huge room to improve”, limited model scale, and picture fidelity that still needs work in some scenes. A technical report is promised.

H3 versus FLUX 3

These two landed eight days apart with nearly the same pitch: one model for several modalities, native audio, open weights to follow. The difference is what you can do about it.

MiniMax H3FLUX 3
AvailabilityLive on launch day, buyable nowGated early access, apply and wait
Pricing$0.13/s at 2K on fal, day oneNone published
Max clip15s at 2K20s at 720p
AudioNative stereoNative, multilingual dialogue
Also generatesImages, audioImages, audio, robot action predictions
Open weightsPromised within days“Later in 2026”, undated

Read the other side on our FLUX 3 tracker. If you want the omni-model idea in production this week, H3 is the only one of the two you can pay for.

What to watch next

The weights. MiniMax said days, not months, and it has an incentive to move: an openly licensed 2K video model with native audio would be the first of its kind, and would land before FLUX 3 Dev. We check the MiniMaxAI Hugging Face organisation and the aggregator catalogs every day, and this page changes the same day anything appears.

Frequently asked questions

Can I use MiniMax H3 right now?
Yes — and that makes it unusual this summer. We checked on July 31, 2026, the day it launched: it is live in the Hailuo app, MiniMax lists API and MiniMax Hub access on its own model page, and the reseller Atlas Cloud already serves three endpoints (text-to-video, image-to-video, reference-to-video). fal.ai listed all three endpoints the same day under the name “Hailuo 03”; kie.ai has not picked it up.
How much does MiniMax H3 cost?
$0.13 per second at 2K on fal.ai, or $0.14 at 2K and $0.10 at 768p on Atlas Cloud — so a 15-second 2K clip is roughly $1.95. References are billed on top at fal: audio is free, the first five images are free, extra images are $0.04 each and a reference video costs another $0.13/s. That makes H3 cheaper at 2K than Seedance 2.0 is at 720p. MiniMax has published no English rate card of its own.
Is MiniMax H3 open source?
Not yet. MiniMax says it will release the weights 'within the next few days, subject to the relevant laws and regulations', and says it designed H3 for compatibility with several domestic Chinese accelerators from the start. As of July 31 no H3 repository exists on the MiniMaxAI Hugging Face organisation. If it lands, it would be the strongest openly licensed video model available.
What is omni-reference?
Instead of separate modes for image-to-video, style reference and motion reference, H3 takes text, images, video and audio as one context and lets you describe the relationship in plain language. MiniMax's own example prompt: 'reference the Hitchcock camera movement from video 1, make the person in image 2 sing, with the singing voice from audio 3.' Reported limits are nine images, three video segments and three audio clips per request.
How does H3 compare to Seedance and Veo?
On price it undercuts both: $0.13/s at 2K on fal against roughly $0.30/s for Seedance 2.0 at 720p on fal and $0.40/s for Veo 3.1 standard. On quality nothing is settled — H3 has not appeared on the Artificial Analysis boards we track, so the ranking claims circulating on launch day are not independently verified. What is genuinely different is the input side: one context of text, image, video and audio instead of one task per endpoint.
What can't H3 do?
MiniMax lists the limits itself: multimodal context understanding still has 'huge room to improve' and will be fused with its M-series language models in a later version; model scale is limited and scaling is the stated direction; and picture fidelity in some scenes needs work. Clips also top out at 15 seconds, shorter than Seedance 2.5's promised 30.