FLUX 3 Video Prompt Generator
Write prompts for FLUX 3 Video, the Black Forest Labs model that generates picture and synchronized sound in one pass. This generator follows the structure in BFL's own prompting guides and uses their published camera and lighting vocabulary, so the terms in your prompt are ones the model was documented to understand. Covers native audio direction, multi-shot hard cuts, keyframe briefs, and the 5 to 20 second timing window.
Describe the main visual scene — subjects, environment, mood, and key visual details
Quick Palettes
Generated Prompt
Fill in the form and click "Generate" to create an optimized FLUX 3 Video video prompt.
Tip: Describe the motion and temporal progression of your scene. Think in terms of "what happens over time" rather than a static description.
FLUX 3 Video Tips
- • Name the format, not just the subject. "A 1987 local news report about teenagers at the mall" carries its own camera work, lighting and edit rhythm
- • One camera idea per shot. Stacking moves is the fastest way to get mushy motion
- • Audio is on by default, so write one sound line or the soundtrack is left to chance
- • Put spoken words in quotes and name a visible speaker, otherwise the line can render as on-screen text
- • There is no negative prompt in FLUX 3. Describe what you want, and keep negation for text, sound and delivery only
- • Repeat a character description word for word across shots. Paraphrasing is what breaks identity
- • Short prompts often beat long ones. Draft at low cost, then extend only the part that failed
How to Use the Prompt Generator
Set the mode and the shape
Choose text-to-video, a keyframe mode, or video continuation, then pick the prompt format you want back. If your idea has a recognizable media language, put it in the Era and Medium field, since that one line does more work than any other input on the page.
Direct one shot
Describe the subject and what changes, then pick a single camera move, a lighting term and a motion quality. Add one audio line naming sounds you can see in frame. Resist stacking camera moves, because that is the fastest route to mushy motion.
Draft, then fix one thing
Copy the prompt into the BFL playground, the flux-3-video API, fal.ai, Replicate, OpenRouter or ComfyUI Partner Nodes. Render a draft first, since it costs a fraction of a full render. When something is off, change the one element that failed rather than rewriting the whole prompt.
Four prompt formats FLUX 3 accepts
Black Forest Labs documents more than one valid way to write a prompt. Each has a job, and picking the wrong one is a common reason results drift.
Natural one-liner
Your default. Best for fast exploration and single subjects.
A low tracking shot of a fox sprinting through wet pine undergrowth at dawn. Mist drifts between the trees as the camera keeps pace beside it. Cool blue morning light, fast but controlled motion, cinematic naturalism.
Labeled fields
Best when iterating, because you can change one line and re-roll without disturbing the rest.
Camera shot: wide shot, low angle Subject + action: a lone rider crosses a shallow desert river Depth of field: shallow (sharp on subject, blurred background) Lighting + palette: warm backlight with soft rim, amber, cream, walnut Motion: water splashes around the horse's legs, orange dust hangs in the light Style: epic western realism
Timecoded beats
Paces moments inside one continuous take. Does not create cuts.
0.0-1.5s: locked wide of a still harbor at dawn, boats motionless on glassy water 1.5-3.0s: a slow push-in begins as gulls lift off the water 3.0-5.0s: the sun breaks the horizon, warm light spreads and the camera settles
Multi-shot hard cuts
Real angle changes inside one generation. Save it for 15 to 20 second clips.
SHOT ONE: wide aerial of a desert highway at dawn, a single red car speeding through. HARD CUT. SHOT TWO: interior close-up, the driver's hands drumming the wheel. HARD CUT. SHOT THREE: from the roadside, the car shrinks into the heat haze. One music bed across all three shots.
FLUX 3 Video specs and configs
What the model actually exposes, taken from the Black Forest Labs API reference. Several of these differ from what image model habits would lead you to expect.
| Clip length | 5 to 20 whole seconds, or auto. Video continuation caps at 15s |
|---|---|
| Frame rate | 24fps, fixed. No fps parameter |
| Resolution | hd 720p (default) or fhd 1080p via upsampler. 1920 x 1088 at 16:9 |
| Aspect ratios | auto, 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Modes | Text-to-video, image-to-video (keyframes), video continuation, draft enhance |
| Keyframes | 1 image sets the opening frame, 2 interpolate start to end, up to 10 act as timed waypoints |
| Audio | Native and synchronized, on by default. Speech, ambience, effects and music. 13+ lip-synced languages |
| Multi-shot | SHOT ONE and HARD CUT tokens produce angle changes inside a single generation |
| Negative prompt | Not supported. Trailing clauses work only for on-screen text, audio and delivery |
| Not available | No seed, no guidance scale, no step count, no motion strength, no camera parameters |
Camera terms that are confirmed to work
Black Forest Labs publishes a controlled camera vocabulary, which is unusual and worth using. Pair a term with a concrete subject and one motion detail, the way their own examples do, rather than dropping the term on its own.
Movement
pan, tilt, dolly in, dolly out, tracking shot, orbit, arc shot, crane, boom, handheld, whip pan, dolly zoom, steadicam follow, push through, Snorricam, camera roll, pedestal, trucking, locked-on, Lazy Susan
Shot size and angle
macro, extreme close-up, close-up, medium shot, cowboy shot, full shot, two shot, wide shot, establishing shot, aerial, low angle, high angle, POV, over-the-shoulder, Dutch angle, worm's eye, bird's eye, ground level, profile, tableau, fourth wall
Focus and optics
shallow depth of field, deep focus, rack focus, split diopter, tilt shift, wide angle 24mm, telephoto compression, fisheye, anamorphic flares, macro lens, probe lens, halation, parallax, vignette
Lighting
rim light, chiaroscuro, golden hour, neon practicals, volumetric light, hard light, haze, spotlight, light flash, projections, underwater light. Describe light by quality, direction and emotional tone rather than by fixture name
Common mistakes, and what to do instead
| Mistake | Fix |
|---|---|
| Writing a photograph in words, with nothing changing | Describe change over time. Give the subject a verb and the environment its own motion |
| Stacking three camera moves onto one shot | One camera idea. Short compounds like "low tracking shot" are fine |
| Adding "no warping, no extra limbs, no bad hands" | Simplify the motion and let the action play over a longer duration |
| Leaving audio out of the prompt | Audio is on by default, so name one or two layers or the mix is guesswork |
| A quoted line with no speaker on camera | Name the visible speaker, or write voiceover, or the line may render as on-screen text |
| Rewording a character between shots | Repeat the description word for word. Paraphrasing breaks identity |
| Ending with "Render: 1080p, 16:9, 20 seconds" | Those are API parameters, not prompt text. In the prompt they can leak into the picture |
| Comma separated tag soup | Write natural sentences. The text encoder is a language model, not a tag bag |
Frequently Asked Questions
What is FLUX 3 Video?
FLUX 3 Video is the video mode of FLUX 3, the multimodal model Black Forest Labs announced on July 23, 2026, with video reaching general availability in early August 2026. It is not a separate product from the FLUX image line. One set of weights, built on a flow matching approach the team calls Self-Flow, generates image, video and natively synchronized audio together. Video clips run 5 to 20 seconds at a fixed 24 frames per second, at 720p or 1080p, through the flux-3-video API endpoint.
How does this FLUX 3 video prompt generator work?
You fill in the parts of a shot that matter: scene, action, camera move, lighting, motion quality, audio and continuity anchors. The generator then writes them into the structure Black Forest Labs documents in its own prompting guides, using the camera and lighting terms from its published vocabulary. You pick the output shape too, so you can get a natural one-liner, a labeled field block, timecoded beats, or a multi-shot script with hard cuts.
What is the best prompt structure for FLUX 3 Video?
Black Forest Labs documents five elements: subject and action, camera direction, scene and atmosphere, motion qualities, and continuity constraints. The fourth one catches people out. FLUX 3 wants an explicit quality of movement (slow, abrupt, weightless, chaotic, precise, cinematic, documentary) as its own slot, separate from the action itself. The fifth matters most for edits and extensions, where you state what has to stay stable across the clip.
Why does naming a year and a format work so well?
Because a media format is a compressed instruction. Testing published by fal.ai found that "a 1969 documentary about Woodstock" beats an elaborate shot list, since the format carries its camera work, lighting, edit rhythm and period audio mix all at once. The model reproduces the editing grammar of the era rather than applying a filter over modern footage. If your idea has a recognizable media language, lead with era plus medium and keep the rest short.
Does FLUX 3 support negative prompts?
No. There is no negative_prompt parameter in the API, and Black Forest Labs states plainly that FLUX does not support them, so you should describe what you do want. There is a real nuance, though: short trailing negation clauses do work for three specific things. "No on-screen text, no subtitles" stops a quoted line rendering as a graphic, "No sound" suppresses the audio track, and "no announcer delivery" steers voice performance. What does not work is Stable Diffusion habit, so skip lists like "no warping, no extra limbs". Fix those by simplifying the motion instead.
How long can a FLUX 3 video be?
Between 5 and 20 whole seconds per generation, or auto, which fits the length to your content. Video continuation is the exception and caps at 15 seconds. For anything longer than 20 seconds you chain several calls, which is orchestration rather than one long render. Fit the writing to the length: for 5 to 8 seconds stick to one subject, one camera idea and one sound idea, and save narrative arcs and hard cuts for 15 to 20 seconds.
Can FLUX 3 cut between shots in a single generation?
Yes, and this is one of its more distinctive abilities. SHOT ONE and HARD CUT are literal control tokens. Writing "SHOT ONE: wide aerial of a desert highway at dawn. HARD CUT. SHOT TWO: interior close-up of the driver's hands on the wheel." produces real angle changes inside one clip. Add a line like "one music bed across all three shots" so the score does not restart at every cut. Timecodes such as 0.0-1.5s work differently: they pace beats within a single continuous take rather than creating cuts.
How do I keep a character consistent across shots?
Repeat the subject description in identical wording in every shot block. If shot one says "the detective in a rumpled tan coat", shot two says exactly that again. Rewording it for variety is the single most common cause of identity drift, because the model has nothing to anchor to. The same applies to products and locations.
Does FLUX 3 generate audio, and how do I prompt for it?
Audio generation is on by default, so every prompt is an audio prompt whether you wrote one or not. Leave it out and the soundtrack is guesswork. There are four layers: speech, ambience, effects and music, and you should use only the one or two that matter, because a crowded mix lowers quality. Put exact spoken words in quotation marks and name the visible speaker, which is what triggers lip-sync. A quoted line with no speaker on camera can be treated as text and painted onto the screen. Name sound sources you can see in frame, and never ask for silence, since naming the quiet sound (room tone, wind) works better.
Which languages does FLUX 3 lip-sync?
More than 13, including English with several dialects, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi and Punjabi. Label a line with its language when it is not English. Early hands-on reports suggest pronunciation and emotional tone are worth checking, so generate a couple of takes when the exact delivery matters.
What is the difference between the generation modes?
Text-to-video writes the whole scene from the prompt. Image-to-video is not a separate endpoint, it is a keyframes array, and the number of images changes what happens: one image is the exact opening frame, two images become a start and end pair that the model fills in between, and three to ten act as ordered waypoints along a timeline that you can pin to exact seconds. Video continuation takes an existing clip and carries it forward, so you describe only the next beat. Draft mode is a cheap fast preview you can re-render at full quality once you like it.
How should image-to-video prompts differ from text-to-video?
Do not describe what the picture already shows. The image supplies subject, composition, lighting and style, so spending words on those wastes your control and can fight the image. Describe change instead of state. The official examples read as transformations, using phrasing like "transitions from bright midday to glittering night" followed by the mechanism of the change. For start and end pairs, keep the two frames related, with the same subject, scene or camera setup, so there is a plausible path between them.
How long should a FLUX 3 prompt be?
Shorter than most people expect. There is no documented character limit, but Black Forest Labs warns that over-stuffing reduces coherence, and fal.ai reported that almost every prompt in its example set was a single sentence with no shot list and no lighting direction. Start with a one-liner, then extend only the part that failed rather than rewriting from scratch. Long prompts earn their length for multi-shot and timecoded work, where the structure is doing real work. Write in natural sentences, not comma separated tags, since the text encoder is a language model rather than a tag bag.
What resolutions and aspect ratios does FLUX 3 Video support?
Two resolution tiers: hd at 720p, which is the default, and fhd at 1080p, produced by an upsampler rather than native full resolution sampling. For 16:9 the documented full resolution output is 1920 by 1088. Aspect ratios are auto, 21:9, 2:1, 16:9, 4:3, 1:1, 3:4 and 9:16. Frame rate is fixed at 24 and there is no fps parameter, so leave it out of your prompt.
Is there a seed or a guidance scale in FLUX 3 Video?
The video endpoint exposes far fewer knobs than image models trained people to expect. There is no guidance scale, no step count, no CFG and no motion strength slider, and camera control is prompt language rather than a parameter. Reproducibility is limited, so treat re-rolls as part of the workflow and use draft mode to keep that cheap. What you do get is mode, prompt, keyframes, duration, resolution, aspect ratio, an audio toggle and a safety tolerance setting.
How much does FLUX 3 Video cost to run?
Black Forest Labs prices per second of output. At the time of writing that is about $0.17 per second at 720p and $0.29 per second at 1080p for text-to-video and image-to-video, with draft mode near $0.06 per second, and video continuation priced higher. Prices change, so check the official pricing page before you budget. This prompt generator is free either way, since it writes the prompt rather than rendering the video.
Are the model weights available to download?
Not for FLUX 3 Video. It is API only, and the FLUX API terms do not grant a right to download or self host it. A FLUX 3 Dev open weight version has been announced but not released, with no published date, license or parameter count, so treat any specific claim about it as speculation. Note also that the API terms let Black Forest Labs use your inputs and outputs, including for training, so read them before putting sensitive material through it.
How does FLUX 3 Video compare to Veo, Sora and Kling?
Carefully is the honest answer. The published win rates come from Black Forest Labs' own human preference evaluations on 10 second 720p clips from a pre-release checkpoint, and they show FLUX 3 ahead of Runway Gen-4.5 and Kling v3 Pro, and at roughly statistical parity with Seedance 2.0 and Gemini Omni Flash. Black Forest Labs published no comparison against Veo, Sora or Wan, so claims in either direction there are unverified, and the model does not yet appear on independent leaderboards. Where it clearly differs on specifications is the 20 second single generation, native synchronized audio, and multi-shot cuts inside one clip.
Will these prompts work with other AI video tools?
Mostly, with edits. The shot structure, camera vocabulary and lighting language carry over well to Veo, Sora, Kling, Runway, Seedance and Wan. The parts that are specific to FLUX 3 are the SHOT ONE and HARD CUT tokens, the four layer audio block, and the absence of a negative prompt. If you are moving a prompt to a model that has a negative prompt field, pull your constraints out of the trailing clause and put them there instead.
Is this tool free to use?
Yes. You get 3 free prompt generations per day, and no signup is required. For unlimited access, a Promptslove membership covers all the AI tools plus 20,000+ premium prompts.
More video prompt generators
Want Unlimited AI Prompt Generation?
Get unlimited access to all AI tools, 20,000+ premium prompts, courses, and resources designed to maximize your creative output.
FLUX and Black Forest Labs are trademarks of their respective owners. Promptslove is not affiliated with or endorsed by Black Forest Labs. Model specifications and pricing described here reflect published documentation at the time of writing and may change.
