VidModelHub

FLUX 3: one model for image, video, audio and action

Updated July 30, 2026

AnnouncedLaunched July 23, 2026 — and almost nobody can use it yet. Black Forest Labs' first video model arrives inside a single multimodal network, in gated early access, with no public API and no published price. Here's what is real, what is claimed, and what to watch next.

What FLUX 3 actually is

FLUX 3 is Black Forest Labs' first model that leaves the image lane. One network generates images, video with natively generated audio, and action predictions for robots — jointly trained rather than assembled from separate specialists behind a shared API. BFL calls the alignment approach Self-Flow.

The research argument is worth taking seriously even if you never touch the robotics side: images capture spatial structure at one instant, video adds time and physical dynamics, audio exposes the causal link between a mechanical event and the sound it makes. Train on all three at once and each modality constrains the others — which is why the hammer blow in a FLUX 3 clip rings on the exact frame the metal is struck, instead of on a soundtrack pasted over the top.

For this site, FLUX 3 is also the first model that belongs on both sides of the house — it appears in our video model tracker and our image model tracker at the same time.

What you can and can't get today

Status (Jul 30, 2026)How you get it
FLUX 3 VideoOpen — gated early accessApply at bfl.ai/models/flux-3; approved users work through Discord
FLUX 3 ImageNot open — “in the following weeks”Separate early-access phase, no date given
FLUX 3 ActionPartners onlyFLUX-mimic, built with mimic robotics; in testing at Audi
FLUX 3 Dev (open weights)Confirmed, undated — “later in 2026”Hugging Face, when it lands
Public API & pricingNot publishedBFL says “in the following weeks”
No pricing exists
There is no per-second rate, per-clip price or subscription tier for any FLUX 3 tier. Pages quoting FLUX 3 costs today are extrapolating from FLUX.2 or inventing numbers. The moment real pricing appears on BFL, fal or kie, it goes into our cost calculator.

Video specs, as announced

FLUX 3 Video
Max length20 seconds in a single generation
AudioNative — dialogue (multilingual), SFX, music, generated with the picture
Resolution720p during early access; 1080p expected shortly after
Input modesText, up to 10 image references, start frames, keyframes, reference video, audio
Generation modesText-to-video, image-to-video, video-to-video, keyframe transitions, video/audio continuation
Longer piecesAgentic chaining of clips into multi-shot sequences
Aspect ratios9:16 through 21:9
ParametersNot disclosed

What it looks like

These are Black Forest Labs' own published samples, mirrored to our CDN and re-encoded — audio intact, so turn the sound on. BFL did not publish the prompts behind them, so treat them as a capability reel, not prompt/output pairs.

Stop-motion bakery4s

Clay-animation look with the frame-rate stutter and thumbprint texture of the real technique — not a smooth 3D render wearing a clay shader.

Alpine descent15s

The longest sample BFL published (15s): a continuous aerial-to-tracking descent that has to keep terrain, snow spray and scale coherent throughout.

Ballroom, chandelier light4s

Period drama in low key: dozens of specular chandelier reflections on a polished floor while couples move through them.

Galloping horse5s

Animal gait is the classic tell — leg order, hoof contact and mane inertia at speed, in a lateral tracking shot.

Ocean swell at the cliffs5s

Fluid simulation without a simulator: breaking wave, spray, foam dispersal and the sound of the impact generated together.

Motorcycle POV, wet road4s

Handlebar POV with rain beading on the screen, road spray and a lead rider holding a consistent line — the reference-frame stability case.

Official FLUX 3 samples · Black Forest Labs · source

The benchmark numbers, and how to read them

BFL published a preference evaluation — human raters picking between FLUX 3 and a competitor on the same prompt, run on 10-second 720p clips. FLUX 3's win rate:

FLUX 3 win rate
vs Luma Ray 3.293%
vs Runway Gen-4.577%
vs Grok Imagine Video69%
vs Kling v3 Pro60%
vs HappyHorse 1.0 / 1.159% / 57%
vs Seedance 2.052% — a coin flip
vs Gemini Omni Flash52% — a coin flip
Read the fine print
These are BFL's own numbers, and BFL itself labels the chart a preliminary evaluation of an early FLUX 3 candidate — a pre-release checkpoint, not the model in early access. The honest headline is the bottom two rows: against the models actually leading the field, this is a tie, not a leapfrog. No independent benchmark exists yet and FLUX 3 has not entered the public arenas.

What early testers report

  • Dialogue and drama hold up. Character scenes with spoken lines are where FLUX 3 impresses most — lip timing and the audio landing with the picture.
  • Fast action is not its edge. Rapid camera movement and kinetic sequences still look better out of Seedance.
  • Image-to-video references are inconsistent. Supplied reference images do not reliably stick — an early-access rough edge BFL will need to fix before this is a production tool.
  • Reasoning-ish prompt handling. Testers report it works out setting, period and event order from very sparse prompts — one widely shared clip came from nothing more than "Visualization of ‘Annabel Lee’ by Edgar Allan Poe."

What changes for FLUX.2 users

Nothing, yet. FLUX.2remains BFL's shipping image family and the open-weight photorealism option — FLUX 3 Image is not open, and FLUX 3 Dev has no date. When the image tier does open, the promised gains are complex-prompt handling, multilingual text rendering and a wider style range, all from a model that also understands motion and sound.

FLUX.2 (today)FLUX 3 (when it opens)
ModalitiesImage only (T2I, I2I)Image, video, audio, action
AvailabilityPublic API + open weightsGated early access, video first
Open weightsDev (32B) and Klein (Apache-2.0)Dev planned, undated
PricingPublishedNone yet

Where it will show up first

Both major aggregators have already staged the page. fal lists FLUX 3 text-to-video and image-to-video as coming soon; kie has an upcoming placeholder that explicitly says model IDs, pricing and limits are not announced. That is the same pattern we watched before the Seedance releases — the tell is the moment a real per-second price appears on either page.

Our monitor checks BFL, fal and kie every day and this page is updated the same day anything moves.

The three things still missing

  1. Public API and pricing — decides whether FLUX 3 is a real competitor for production work or a demo everyone admires.
  2. FLUX 3 Image early access — the tier most people actually want, promised within weeks.
  3. FLUX 3 Dev weights and license — whether a full multimodal backbone can realistically be run locally, and on what terms.

Get the three announcements, nothing else

One email per milestone: image early access, public API with real prices, and open weights. We test each on day one and send what we measured.

Related

  • FLUX 3 vs Seedance 2.5 — the two headline video releases of this summer, compared on what is actually known.
  • FLUX 3 prompts — originals written for native audio and 20-second multi-shot chaining.
  • Release calendar — every tracked model and every dated milestone, video and image.

Frequently asked questions

Is FLUX 3 released?
It was officially launched on July 23, 2026, but 'launched' does not mean available. Only FLUX 3 Video is open, through a gated early-access program you have to apply for at bfl.ai/models/flux-3. Image early access is promised 'in the following weeks', action prediction ships through partners only, and the open-weight FLUX 3 Dev release is planned for later in 2026.
How much does FLUX 3 cost?
Nothing has been published — no per-second rate, no per-image price, no subscription tier. Black Forest Labs has said the public API is coming in the following weeks. Anyone quoting FLUX 3 pricing today is guessing or making it up.
How long can FLUX 3 videos be?
Up to 20 seconds in a single generation, with audio generated in the same pass. Longer pieces are built by chaining clips into multi-shot sequences. Resolution is capped at 720p during early access, with 1080p expected shortly after.
Does FLUX 3 generate sound?
Yes — native audio is generated jointly with the video, not dubbed on afterwards: dialogue (including non-English), sound effects tied to on-screen events, and music. That joint training is the whole architectural argument behind the model.
Is FLUX 3 better than Seedance or Veo?
Unproven. In Black Forest Labs' own preference test, FLUX 3 won 93% against Luma Ray 3.2 and 77% against Runway Gen-4.5 — but only 52% against Seedance 2.0 and Gemini Omni Flash, which is a coin flip. BFL labels those numbers a preliminary evaluation of an early, pre-release checkpoint. No independent benchmark exists yet, and FLUX 3 has not entered the public arenas.
Will FLUX 3 have open weights?
BFL has confirmed an open-weight FLUX 3 Dev release of the full multimodal backbone, planned for later in 2026, with no date, parameter count or license terms announced. Judging by FLUX.1 and FLUX.2, expect a non-commercial license rather than a permissive one. Until then, FLUX.2 remains the open-weight option.
How do I get FLUX 3 access?
Apply through the official model page at bfl.ai/models/flux-3; approved applicants are currently working through Discord. Beware lookalike domains — BFL has stated that consumer sites branded 'flux-3' are not affiliated with the company.