Written for early access · rendered the day we get in
FLUX 3 prompts — eight originals for a model that hears
FLUX 3 generates sound in the same pass as the picture, holds 20 seconds in one shot, and chains clips into sequences. These eight are written to push exactly those things — hammer strikes that must ring on the right frame, a two-hander with two distinct voices, Japanese dialogue, a soundscape that pans as the camera walks.
2. Environment + light — where, and what the light is doing.
3. An explicit Audio: line — name every sound you want, in the order it happens. Say "no music" when you mean it, or you will get a score.
4. A short style line — look, stock, lens.
BFL has published no official prompt guide for FLUX 3. This is the convention that works across audio-native video models today; we update it the moment there is an official one.
Blacksmith, sound on the frameAudio causality
Medium close-up on a blacksmith's anvil in a dark forge, orange light from the coals raking across the smith's forearms. He strikes the glowing bar four times with an even rhythm, turning it a quarter after the second strike; sparks scatter on each impact and die on the dirt floor. Handheld, slight breathing movement. Audio: four hammer strikes ringing off the anvil with decaying metallic tail, spark hiss, low bellows roar, no music. Style: naturalistic documentary, shallow depth of field.
What it tests: The single hardest thing to fake by dubbing: each ring must land on the exact frame the hammer contacts, and the fourth strike must sound different because the bar has cooled.
Lighthouse two-handerDialogue · 20s
Twenty seconds, one shot. Inside a lighthouse lamp room during a storm, two keepers face each other across the rotating lamp housing. The older one says, "You logged it as a fishing boat." The younger one, not looking up: "That's what the book says." The older keeper turns to the window; the beam sweeps across both faces once every four seconds, alternating who is lit. Rain hammers the glass. Audio: two distinct voices with the acoustics of a small glass room, storm outside, the mechanical grind of the rotating lamp, no music. Style: 1970s film stock, cool green cast.
What it tests: Dialogue is FLUX 3's reported strength — this tests it under load: two voices with different timbre, lip timing across a 20-second take, and a light cycle that must stay on its four-second beat.
Tokyo counter, Japanese dialogueMultilingual
A nine-seat ramen counter at midnight, seen from the customer's side. The chef sets a bowl down and says in Japanese, "お待たせしました" — the subtitle is not shown. The customer nods, breaks the chopsticks apart and lifts noodles through the steam. Warm tungsten light, condensation on the window behind. Audio: spoken Japanese with natural room reverb, broth ladle, chopsticks separating, the extractor fan hum, faint street traffic through glass, no music. Style: quiet observational realism, 35mm.
What it tests: Multilingual dialogue is an announced capability and rarely tested honestly — non-English speech has to be intelligible, correctly accented, and lip-matched, with no English bleeding into the mouth shapes.
ASMR: knife through candied fruitAudio-first
Extreme close-up, top-down. A very sharp knife lowers through a candied crystal-glazed persimmon on a black stone slab; the sugar shell fractures first, then the blade passes through the soft interior and taps the stone. Two halves fall apart slowly. Soft directional light, no other objects in frame. Audio: the glass-like crack of the sugar shell, the dull wet cut, the knife tip clicking on stone, faint room tone, absolutely no music. Style: macro food photography, high contrast.
What it tests: An audio-first prompt where the sound carries the whole clip: three distinct materials in under five seconds, each with its own acoustic signature, in the order the blade meets them.
Twenty-second market walkContinuous audio bed
One unbroken 20-second Steadicam shot following a woman with a canvas bag through a covered market. She passes a fishmonger hosing ice, a caged radio playing, a man calling out prices, and steps out through the far arch into open street noise and daylight. The light changes from cold fluorescent under the roof to warm afternoon sun as she exits. Audio: the soundscape shifts with her position — water on ice fading behind, the radio passing left to right, the vendor's call peaking as she passes, then the acoustic opening up to traffic as she leaves the building. Style: naturalistic, handheld energy.
What it tests: Spatial audio over time: sources must pan and fall away with the camera's position across 20 seconds. Get this wrong and it sounds like one static mix under a moving picture.
Three-shot rooftop chaseMulti-shot chain
Three shots, cut on action, same character throughout: a courier in a yellow jacket. Shot one, wide: she vaults a rooftop air-conditioning unit, pigeons scattering. Shot two, low angle close: her boots land on gravel, camera whip-pans right to follow. Shot three, drone pull-back: she runs to the roof edge and stops, city haze behind her. Evening light, long shadows. Audio: continuous across all three cuts — footfalls on gravel, breathing, wing beats, distant traffic, a single low sustained tone that carries under the cuts. Style: grounded action, no slow motion.
What it tests: Agentic chaining is the announced route to longer pieces. The test is whether the jacket, face and light stay identical across cuts while the audio runs continuously rather than restarting per clip.
Vertical street answer9:16
Vertical 9:16 framing. A man in his sixties stands on a windy pier, held in a medium shot from chest height, a microphone just in frame at the bottom edge. He thinks, then says, "Forty-one years. Every morning except the day my daughter was born." He looks off camera and half-smiles. Grey sea and gulls behind him. Audio: his voice slightly wind-buffeted, the microphone catching gusts, gulls, waves against pilings, no music. Style: documentary interview, natural light.
What it tests: Vertical is where this content actually gets watched, and wind-affected speech is a specific audio-realism test — a clean studio voice over a windy pier is the giveaway.
Keyframe: closed to bloomedKeyframe control
Start frame: a tight bud on a peony stem against a dark background, one water droplet on the sepal. End frame: the same peony fully open, the droplet gone, a single petal fallen onto the surface below. Between them, generate a continuous time-lapse bloom with the camera drifting almost imperceptibly closer. Light warms slightly as the flower opens. Audio: no music — only a low room tone, the faint crackle of unfurling petals, and the single soft tap of the falling petal at the end. Style: botanical macro, black backdrop.
What it tests: Keyframe-to-video is an announced control mode. Two fixed endpoints leave nowhere to hide: the interpolation has to be botanically plausible, and the one scripted sound must land on the frame the petal touches down.
Get the rendered results
One email when we get FLUX 3 access, with all eight prompts rendered — sound on, unedited, what worked and what didn't. No other emails.
Until then: models you can prompt today
The audio line in these prompts works on any model that generates native sound. Our Veo 3 prompt library has 44 originals in JSON and prose form, the Seedance library pairs official prompts with their outputs, and the prompt generator builds either format for you. For where FLUX 3 stands against the model everyone is comparing it to, see FLUX 3 vs Seedance 2.5.
Frequently asked questions
How do you write a FLUX 3 prompt?
Black Forest Labs has not published a prompt guide. The structure that works across audio-native video models — and the one these prompts use — is four parts: subject and action first, then environment and light, then an explicit 'Audio:' line naming every sound you want (and saying 'no music' if you mean it), then a short style line. Write the audio line as deliberately as the picture: it is generated in the same pass, not dubbed on.
Have these prompts been tested on FLUX 3?
No, and we won't pretend otherwise — FLUX 3 Video is in gated early access and there is no public API. Each prompt is written to expose a specific announced capability: sound landing on the right frame, 20-second continuity, multilingual dialogue, keyframe control, multi-shot chaining. The day we get access we render all eight and publish the outputs here, unedited, the failures included.
Can I use these prompts on other models today?
Mostly yes. Veo 3.1 and Seedance also generate native audio, so the audio line carries over; drop the 20-second framing where the model can't hold it and split the multi-shot prompt into separate shots. Our Veo 3 and Seedance prompt libraries have model-specific versions you can copy right now.
How long can a FLUX 3 clip be?
Up to 20 seconds in one generation with audio, at 720p during early access. Longer sequences are built by chaining clips — which is why one of these prompts is written as a three-shot chain rather than a single take.