Vidu Video Prompt Generator
Write prompts for Vidu Q4 Preview and Vidu Q3, ShengShu's video models with native sound and 16-second clips. This generator follows Vidu's own prompt guides: subject, action, camera, scene and style in that order, one camera move per clip, @1 tags for reference images, quoted dialogue for lip sync, and exclusions rewritten as what should be on screen.
Describe the main visual scene — subjects, environment, mood, and key visual details
Quick Palettes
Generated Prompt
Fill in the form and click "Generate" to create an optimized Vidu video prompt.
Tip: Describe the motion and temporal progression of your scene. Think in terms of "what happens over time" rather than a static description.
Vidu Tips
- • Follow Vidu's order: subject, action, camera, scene, style. Put the shot call first when framing matters most
- • One location, one subject, one action, and one camera move per clip. Stacked moves make the camera overcorrect
- • Use a single speed word for motion: subtle, gentle, slow or steady keep it smooth
- • Tag uploads as @1, @2, @3 in upload order and give each one a job
- • Audio is on by default, so write it. Quote dialogue, name the speaker, and write "no music" if you want none
- • Text-to-video is a Q3 job. Q4 Preview takes images and references and adds 2K and 4K
- • There is no negative prompt. Describe what should be there instead
Vidu Prompt Templates
Copy-ready shot briefs built around what Vidu rewards. Swap the [BRACKETED] parts for your own scene, then paste straight into the tool.
Single Cinematic Shot
CinematicVidu's order: shot call first, then subject, action, scene, style, with one camera move at the end.
[ANGLE, e.g. Low-angle shot]: a [SUBJECT WITH CLOTHING AND ONE DETAIL] [ONE ACTION WITH ONE SPEED WORD, e.g. walks steadily toward camera] in [ONE LOCATION]. [SPECIFIC LIGHT, e.g. side key light, cool shadows]. [STYLE, e.g. cinematic realism, teal and amber grade]. [SOUND TIED TO THE ACTION]. No background music. [ONE CAMERA MOVE, e.g. Slow push-in.] [DURATION] seconds, [RATIO], [RESOLUTION].
Dialogue with Lip Sync
AudioNamed speaker, delivery note and a quoted line sized at about 2.5 words per second.
Medium close-up: [CHARACTER] sits at [PLACE], [LIGHT]. [He/She] looks up and speaks. [CHARACTER] ([DELIVERY, e.g. warm, unhurried]): "[LINE OF ABOUT 2.5 WORDS PER SECOND]" [AMBIENT SOUND]. [MUSIC LINE or "No background music."] Static locked-off shot. [DURATION] seconds, [RATIO], [RESOLUTION].
Reference Characters
ReferencesReference-to-video with @ tags in upload order, each reference given one job.
@1 [CHARACTER ROLE] and @2 [CHARACTER ROLE] [SHARED ACTION] in @3 [LOCATION]. @1 [SMALL ACTION]; @2 [REACTION]. [LIGHT AND STYLE]. [SOUND OR DIALOGUE]. [ONE CAMERA MOVE]. [DURATION] seconds, [RATIO], [RESOLUTION].
Photo Come to Life
Image-to-VideoImage-to-video: describe only what changes, not what the image already shows.
[THE ONE THING THAT MOVES, e.g. Her hair lifts gently in the wind] as [SECONDARY MOTION, e.g. petals drift slowly past]. [SOUND, e.g. soft wind and distant birdsong]. No background music. [ONE CAMERA MOVE WITH SPEED WORD, e.g. Gentle push-in.] [DURATION] seconds, [RESOLUTION].
Anime Moment
AnimeVidu has no style switch on Q3 and Q4, so anime is named plainly in the style slot.
[ANGLE]: [CHARACTER WITH HAIR, OUTFIT AND COLOUR] [ONE ACTION] on [LOCATION]. Japanese anime style, clean cel shading, expressive eyes, painted background, [TIME OF DAY] light. [SOUND]. [MUSIC LINE]. [ONE CAMERA MOVE]. [DURATION] seconds, [RATIO], [RESOLUTION].
Start-to-End Transition
KeyframesStart and end frames constrain the camera better than words; describe the one continuous change between them.
The scene moves in one continuous action from the first frame to the last: [WHAT CHANGES, e.g. the empty café fills with morning light as the barista flips the sign to open]. [ONE SPEED WORD] motion, [CAMERA: static or one move]. [SOUND]. No background music. [DURATION] seconds, [RESOLUTION].
The Vidu prompt pattern
Vidu says the order of the parts matters more than the length of the prompt. Keep to one subject, one action and one camera move per clip.
| Subject | Who or what, with clothing, colour and one distinguishing detail |
|---|---|
| Action | One clear action with one speed word: "walks steadily toward camera" |
| Camera | One angle and one move: "low-angle shot", ending with "Slow push-in." |
| Scene | One location with specific light: "wet alley, side key light, cool shadows" |
| Style | Cinematic, realistic or anime, plus the grade and mood |
| Sound | Quoted dialogue with the speaker, sound effects tied to actions, and music or "no music" |
Drifts and overcorrects
A moody cinematic scene of a woman in a city with dynamic movement, slow dolly forward with a slight pan left and a zoom, beautiful lighting, no people.
Vidu style
Low-angle shot: a courier in a yellow rain jacket walks steadily toward camera down an empty neon alley at night. Wet cobblestones, side key light, cool shadows. Cinematic realism. Rain taps the awning. No background music. Slow push-in.
Vidu specs and configs
| Latest model | Vidu Q4 Preview, viduq4-preview, launched October 7, 2026. Image-to-video and reference-to-video only |
|---|---|
| Text-to-video model | Vidu Q3 (January 30, 2026): viduq3-pro and viduq3-turbo, also used for start and end frame |
| Duration | 1 to 16 seconds on Q3 and Q4. Some reference models run 3 to 15 seconds |
| Resolution | Q3: 540p, 720p, 1080p. Q4 Preview: 540p, 720p, 1080p, 2K, 4K |
| Frame rate | 24fps |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4. Image-to-video follows the input image |
| References | Q4: up to 15 images and 3 audio clips. Q3: up to 7 images. Tagged @1, @2 or @Name |
| Audio | Native dialogue, voiceover, sound effects and music, on by default, with lip sync |
| Not available | No negative prompt, no style switch, no movement_amplitude, no bgm field on Q3 or Q4 |
| Prompt limit | 5,000 to 20,000 characters on the Vidu API, 2,000 on fal.ai |
How to Use the Prompt Generator
Pick the model and mode
Use Vidu Q3 Pro for text-to-video and start-plus-end frames. Use Vidu Q4 Preview for image-to-video and reference-to-video, or when you need 2K or 4K. For references, list each upload as @1, @2 and so on with its job.
Direct one shot and its sound
Describe one subject doing one thing in one place, pick a single camera move and a speed word, then write the dialogue, sound effects and whether you want music.
Generate, then draft at 720p
Paste the prompt into vidu.com, the Vidu API or fal.ai with the same duration and aspect ratio. Draft at 720p, or use off-peak mode on Q3 to halve the cost, then render your pick at full resolution.
Frequently Asked Questions
What is Vidu?
Vidu is the AI video model family from ShengShu Technology, developed with Tsinghua University. It runs on vidu.com, the Vidu API at platform.vidu.com, and hosts such as fal.ai. The current models are Vidu Q4 Preview, launched on October 7, 2026, and Vidu Q3, launched on January 30, 2026. Both generate video with native sound in one pass at 24fps and run up to 16 seconds per clip.
What is new in Vidu Q4 Preview?
Vidu Q4 Preview (API ID viduq4-preview) is ShengShu's new flagship. It accepts up to 15 reference images plus up to 3 voice or audio references, renders from 540p up to 2K and 4K, and runs 1 to 16 seconds. For now it only does image-to-video and reference-to-video. On the Artificial Analysis image-to-video arena it sits in a statistical tie for second place, alongside Seedance 2.0 and Gemini Omni Flash, behind MiniMax H3.
Which Vidu model should I pick?
For text-to-video and start-plus-end-frame clips, use Vidu Q3 Pro, or Q3 Turbo when speed matters more. For image-to-video, reference-to-video, large reference sets or anything above 1080p, use Vidu Q4 Preview. Q3 also has specialised reference models: viduq3-mix, viduq3-ad for adverts and viduq3-drama for character-driven short drama with named subjects and voices.
How should a Vidu prompt be structured?
Vidu's own formula is subject, then action, camera, scene and style, and Vidu says that order matters more than length. When framing is the point of the shot, put the shot call first, as in "Low-angle shot: a courier walks toward camera down a narrow alley at night." Keep each clip to one location, one subject and one action, and describe specific light ("side key light, cool shadows") instead of mood words.
How do I write camera movement for Vidu?
One camera action per generation, written as a modifier at the end of the shot, like "Slow push-in." Vidu's guides warn that stacking moves such as "slow dolly forward with slight pan left" makes the camera overcorrect. Pair the move with a single speed word, since subtle, gentle, slow and steady reduce variance. Zoom in is a movement, not an angle, so state the angle as well. For tight control, start and end frames constrain the camera better than words.
How do reference images work?
In reference-to-video, refer to uploads with the @ syntax in upload order: @1, @2, @3, or [@1] in brackets. Vidu's own example reads "@1 and @2 are eating hot pot together." Named subjects work too, as in "@Cowboy sits at the bar in @Bar." Give every reference a clear job, such as the character, the outfit, the product or the location. Q4 Preview takes up to 15 images but does not yet support named subjects, so numbered tags are safest there. Q3 takes up to 7.
Does Vidu generate sound and dialogue?
Yes. On Q3 and Q4 audio is on by default and covers dialogue, voiceover, sound effects and music, with lip sync. Write the sound explicitly, because a missing cue does not mean silence. Name the speaker, describe the delivery and quote the line. Keep lines to roughly two and a half words per second so they fit the clip, and use a close or medium shot for clean lip sync. If you want no music, say so in the prompt. Q3 lists English, Chinese and Japanese voices.
Is there a negative prompt, style or motion-strength setting?
Not on Q3 or Q4. The current Vidu API has no negative prompt field, no general or anime style switch, no movement_amplitude and no bgm parameter. Those appear only on older Q1 and Q2 endpoints on fal.ai. On the current models, style, motion strength, music and exclusions all go in the prompt text, which is exactly what this generator writes.
What durations, resolutions and aspect ratios does Vidu support?
Q3 and Q4 Preview both run 1 to 16 seconds at 24fps, with some reference models limited to 3 to 15 seconds. Q3 renders 540p, 720p or 1080p. Q4 Preview adds 2K and 4K. Aspect ratios are 16:9, 9:16, 1:1, 4:3 and 3:4, and image-to-video follows your input image.
How long can a Vidu prompt be?
It depends on where you run it. The Vidu API accepts 5,000 characters on most endpoints and 20,000 on Q4 and Q3 Pro, while fal.ai caps Q3 at 2,000 characters. Length is not the goal though. This generator keeps prompts tight, about 60 to 150 words for a single shot, so they paste anywhere.
What does Vidu cost?
On the Vidu API, pricing is in credits per second and scales with resolution. Q4 Preview runs from 9 credits per second at 540p to 78 at 4K, and Q3 Pro from 9 at 540p to 24 at 1080p. Off-peak mode roughly halves the price on Q3 in exchange for delivery within 48 hours.
Is this tool free to use?
Yes. You get 1 free prompt generation per day with no signup required. For unlimited access, sign up for a Promptslove membership which includes all AI tools and 29,000+ premium prompts.
More video prompt generators
Want Unlimited AI Prompt Generation?
Get unlimited access to all AI tools, 29,000+ premium prompts, courses, and resources designed to maximize your creative output.
Vidu is a trademark of ShengShu Technology. Promptslove is not affiliated with or endorsed by ShengShu. Model specifications and pricing described here reflect published documentation at the time of writing and may change.
