FLUX 3: one model for image, video, audio and action
Updated July 30, 2026
What FLUX 3 actually is
FLUX 3 is Black Forest Labs' first model that leaves the image lane. One network generates images, video with natively generated audio, and action predictions for robots — jointly trained rather than assembled from separate specialists behind a shared API. BFL calls the alignment approach Self-Flow.
The research argument is worth taking seriously even if you never touch the robotics side: images capture spatial structure at one instant, video adds time and physical dynamics, audio exposes the causal link between a mechanical event and the sound it makes. Train on all three at once and each modality constrains the others — which is why the hammer blow in a FLUX 3 clip rings on the exact frame the metal is struck, instead of on a soundtrack pasted over the top.
For this site, FLUX 3 is also the first model that belongs on both sides of the house — it appears in our video model tracker and our image model tracker at the same time.
What you can and can't get today
| Status (Jul 30, 2026) | How you get it | |
|---|---|---|
| FLUX 3 Video | Open — gated early access | Apply at bfl.ai/models/flux-3; approved users work through Discord |
| FLUX 3 Image | Not open — “in the following weeks” | Separate early-access phase, no date given |
| FLUX 3 Action | Partners only | FLUX-mimic, built with mimic robotics; in testing at Audi |
| FLUX 3 Dev (open weights) | Confirmed, undated — “later in 2026” | Hugging Face, when it lands |
| Public API & pricing | Not published | BFL says “in the following weeks” |
Video specs, as announced
| FLUX 3 Video | |
|---|---|
| Max length | 20 seconds in a single generation |
| Audio | Native — dialogue (multilingual), SFX, music, generated with the picture |
| Resolution | 720p during early access; 1080p expected shortly after |
| Input modes | Text, up to 10 image references, start frames, keyframes, reference video, audio |
| Generation modes | Text-to-video, image-to-video, video-to-video, keyframe transitions, video/audio continuation |
| Longer pieces | Agentic chaining of clips into multi-shot sequences |
| Aspect ratios | 9:16 through 21:9 |
| Parameters | Not disclosed |
What it looks like
These are Black Forest Labs' own published samples, mirrored to our CDN and re-encoded — audio intact, so turn the sound on. BFL did not publish the prompts behind them, so treat them as a capability reel, not prompt/output pairs.
Clay-animation look with the frame-rate stutter and thumbprint texture of the real technique — not a smooth 3D render wearing a clay shader.
The longest sample BFL published (15s): a continuous aerial-to-tracking descent that has to keep terrain, snow spray and scale coherent throughout.
Period drama in low key: dozens of specular chandelier reflections on a polished floor while couples move through them.
Animal gait is the classic tell — leg order, hoof contact and mane inertia at speed, in a lateral tracking shot.
Fluid simulation without a simulator: breaking wave, spray, foam dispersal and the sound of the impact generated together.
Handlebar POV with rain beading on the screen, road spray and a lead rider holding a consistent line — the reference-frame stability case.
Official FLUX 3 samples · Black Forest Labs · source
The benchmark numbers, and how to read them
BFL published a preference evaluation — human raters picking between FLUX 3 and a competitor on the same prompt, run on 10-second 720p clips. FLUX 3's win rate:
| FLUX 3 win rate | |
|---|---|
| vs Luma Ray 3.2 | 93% |
| vs Runway Gen-4.5 | 77% |
| vs Grok Imagine Video | 69% |
| vs Kling v3 Pro | 60% |
| vs HappyHorse 1.0 / 1.1 | 59% / 57% |
| vs Seedance 2.0 | 52% — a coin flip |
| vs Gemini Omni Flash | 52% — a coin flip |
What early testers report
- Dialogue and drama hold up. Character scenes with spoken lines are where FLUX 3 impresses most — lip timing and the audio landing with the picture.
- Fast action is not its edge. Rapid camera movement and kinetic sequences still look better out of Seedance.
- Image-to-video references are inconsistent. Supplied reference images do not reliably stick — an early-access rough edge BFL will need to fix before this is a production tool.
- Reasoning-ish prompt handling. Testers report it works out setting, period and event order from very sparse prompts — one widely shared clip came from nothing more than "Visualization of ‘Annabel Lee’ by Edgar Allan Poe."
What changes for FLUX.2 users
Nothing, yet. FLUX.2remains BFL's shipping image family and the open-weight photorealism option — FLUX 3 Image is not open, and FLUX 3 Dev has no date. When the image tier does open, the promised gains are complex-prompt handling, multilingual text rendering and a wider style range, all from a model that also understands motion and sound.
| FLUX.2 (today) | FLUX 3 (when it opens) | |
|---|---|---|
| Modalities | Image only (T2I, I2I) | Image, video, audio, action |
| Availability | Public API + open weights | Gated early access, video first |
| Open weights | Dev (32B) and Klein (Apache-2.0) | Dev planned, undated |
| Pricing | Published | None yet |
Where it will show up first
Both major aggregators have already staged the page. fal lists FLUX 3 text-to-video and image-to-video as coming soon; kie has an upcoming placeholder that explicitly says model IDs, pricing and limits are not announced. That is the same pattern we watched before the Seedance releases — the tell is the moment a real per-second price appears on either page.
Our monitor checks BFL, fal and kie every day and this page is updated the same day anything moves.
The three things still missing
- Public API and pricing — decides whether FLUX 3 is a real competitor for production work or a demo everyone admires.
- FLUX 3 Image early access — the tier most people actually want, promised within weeks.
- FLUX 3 Dev weights and license — whether a full multimodal backbone can realistically be run locally, and on what terms.
Get the three announcements, nothing else
One email per milestone: image early access, public API with real prices, and open weights. We test each on day one and send what we measured.
Related
- FLUX 3 vs Seedance 2.5 — the two headline video releases of this summer, compared on what is actually known.
- FLUX 3 prompts — originals written for native audio and 20-second multi-shot chaining.
- Release calendar — every tracked model and every dated milestone, video and image.
Frequently asked questions
- Is FLUX 3 released?
- It was officially launched on July 23, 2026, but 'launched' does not mean available. Only FLUX 3 Video is open, through a gated early-access program you have to apply for at bfl.ai/models/flux-3. Image early access is promised 'in the following weeks', action prediction ships through partners only, and the open-weight FLUX 3 Dev release is planned for later in 2026.
- How much does FLUX 3 cost?
- Nothing has been published — no per-second rate, no per-image price, no subscription tier. Black Forest Labs has said the public API is coming in the following weeks. Anyone quoting FLUX 3 pricing today is guessing or making it up.
- How long can FLUX 3 videos be?
- Up to 20 seconds in a single generation, with audio generated in the same pass. Longer pieces are built by chaining clips into multi-shot sequences. Resolution is capped at 720p during early access, with 1080p expected shortly after.
- Does FLUX 3 generate sound?
- Yes — native audio is generated jointly with the video, not dubbed on afterwards: dialogue (including non-English), sound effects tied to on-screen events, and music. That joint training is the whole architectural argument behind the model.
- Is FLUX 3 better than Seedance or Veo?
- Unproven. In Black Forest Labs' own preference test, FLUX 3 won 93% against Luma Ray 3.2 and 77% against Runway Gen-4.5 — but only 52% against Seedance 2.0 and Gemini Omni Flash, which is a coin flip. BFL labels those numbers a preliminary evaluation of an early, pre-release checkpoint. No independent benchmark exists yet, and FLUX 3 has not entered the public arenas.
- Will FLUX 3 have open weights?
- BFL has confirmed an open-weight FLUX 3 Dev release of the full multimodal backbone, planned for later in 2026, with no date, parameter count or license terms announced. Judging by FLUX.1 and FLUX.2, expect a non-commercial license rather than a permissive one. Until then, FLUX.2 remains the open-weight option.
- How do I get FLUX 3 access?
- Apply through the official model page at bfl.ai/models/flux-3; approved applicants are currently working through Discord. Beware lookalike domains — BFL has stated that consumer sites branded 'flux-3' are not affiliated with the company.