Vidu Video Prompt Generator

Write prompts for Vidu Q4 Preview and Vidu Q3, ShengShu's video models with native sound and 16-second clips. This generator follows Vidu's own prompt guides: subject, action, camera, scene and style in that order, one camera move per clip, @1 tags for reference images, quoted dialogue for lip sync, and exclusions rewritten as what should be on screen.

Describe the main visual scene — subjects, environment, mood, and key visual details

0/4

Quick Palettes

Generated Prompt

Fill in the form and click "Generate" to create an optimized Vidu video prompt.

Tip: Describe the motion and temporal progression of your scene. Think in terms of "what happens over time" rather than a static description.

Vidu Tips

  • • Follow Vidu's order: subject, action, camera, scene, style. Put the shot call first when framing matters most
  • • One location, one subject, one action, and one camera move per clip. Stacked moves make the camera overcorrect
  • • Use a single speed word for motion: subtle, gentle, slow or steady keep it smooth
  • • Tag uploads as @1, @2, @3 in upload order and give each one a job
  • • Audio is on by default, so write it. Quote dialogue, name the speaker, and write "no music" if you want none
  • • Text-to-video is a Q3 job. Q4 Preview takes images and references and adds 2K and 4K
  • • There is no negative prompt. Describe what should be there instead

Vidu Prompt Templates

Copy-ready shot briefs built around what Vidu rewards. Swap the [BRACKETED] parts for your own scene, then paste straight into the tool.

Single Cinematic Shot

Cinematic

Vidu's order: shot call first, then subject, action, scene, style, with one camera move at the end.

[ANGLE, e.g. Low-angle shot]: a [SUBJECT WITH CLOTHING AND ONE DETAIL] [ONE ACTION WITH ONE SPEED WORD, e.g. walks steadily toward camera] in [ONE LOCATION]. [SPECIFIC LIGHT, e.g. side key light, cool shadows]. [STYLE, e.g. cinematic realism, teal and amber grade]. [SOUND TIED TO THE ACTION]. No background music. [ONE CAMERA MOVE, e.g. Slow push-in.]

[DURATION] seconds, [RATIO], [RESOLUTION].

Dialogue with Lip Sync

Audio

Named speaker, delivery note and a quoted line sized at about 2.5 words per second.

Medium close-up: [CHARACTER] sits at [PLACE], [LIGHT]. [He/She] looks up and speaks. [CHARACTER] ([DELIVERY, e.g. warm, unhurried]): "[LINE OF ABOUT 2.5 WORDS PER SECOND]" [AMBIENT SOUND]. [MUSIC LINE or "No background music."] Static locked-off shot.

[DURATION] seconds, [RATIO], [RESOLUTION].

Reference Characters

References

Reference-to-video with @ tags in upload order, each reference given one job.

@1 [CHARACTER ROLE] and @2 [CHARACTER ROLE] [SHARED ACTION] in @3 [LOCATION]. @1 [SMALL ACTION]; @2 [REACTION]. [LIGHT AND STYLE]. [SOUND OR DIALOGUE]. [ONE CAMERA MOVE].

[DURATION] seconds, [RATIO], [RESOLUTION].

Photo Come to Life

Image-to-Video

Image-to-video: describe only what changes, not what the image already shows.

[THE ONE THING THAT MOVES, e.g. Her hair lifts gently in the wind] as [SECONDARY MOTION, e.g. petals drift slowly past]. [SOUND, e.g. soft wind and distant birdsong]. No background music. [ONE CAMERA MOVE WITH SPEED WORD, e.g. Gentle push-in.]

[DURATION] seconds, [RESOLUTION].

Anime Moment

Anime

Vidu has no style switch on Q3 and Q4, so anime is named plainly in the style slot.

[ANGLE]: [CHARACTER WITH HAIR, OUTFIT AND COLOUR] [ONE ACTION] on [LOCATION]. Japanese anime style, clean cel shading, expressive eyes, painted background, [TIME OF DAY] light. [SOUND]. [MUSIC LINE]. [ONE CAMERA MOVE].

[DURATION] seconds, [RATIO], [RESOLUTION].

Start-to-End Transition

Keyframes

Start and end frames constrain the camera better than words; describe the one continuous change between them.

The scene moves in one continuous action from the first frame to the last: [WHAT CHANGES, e.g. the empty café fills with morning light as the barista flips the sign to open]. [ONE SPEED WORD] motion, [CAMERA: static or one move]. [SOUND]. No background music.

[DURATION] seconds, [RESOLUTION].

The Vidu prompt pattern

Vidu says the order of the parts matters more than the length of the prompt. Keep to one subject, one action and one camera move per clip.

SubjectWho or what, with clothing, colour and one distinguishing detail
ActionOne clear action with one speed word: "walks steadily toward camera"
CameraOne angle and one move: "low-angle shot", ending with "Slow push-in."
SceneOne location with specific light: "wet alley, side key light, cool shadows"
StyleCinematic, realistic or anime, plus the grade and mood
SoundQuoted dialogue with the speaker, sound effects tied to actions, and music or "no music"

Drifts and overcorrects

A moody cinematic scene of a woman in a city with dynamic movement, slow dolly forward with a slight pan left and a zoom, beautiful lighting, no people.

Vidu style

Low-angle shot: a courier in a yellow rain jacket walks steadily toward camera down an empty neon alley at night. Wet cobblestones, side key light, cool shadows. Cinematic realism. Rain taps the awning. No background music. Slow push-in.

Vidu specs and configs

Latest modelVidu Q4 Preview, viduq4-preview, launched October 7, 2026. Image-to-video and reference-to-video only
Text-to-video modelVidu Q3 (January 30, 2026): viduq3-pro and viduq3-turbo, also used for start and end frame
Duration1 to 16 seconds on Q3 and Q4. Some reference models run 3 to 15 seconds
ResolutionQ3: 540p, 720p, 1080p. Q4 Preview: 540p, 720p, 1080p, 2K, 4K
Frame rate24fps
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4. Image-to-video follows the input image
ReferencesQ4: up to 15 images and 3 audio clips. Q3: up to 7 images. Tagged @1, @2 or @Name
AudioNative dialogue, voiceover, sound effects and music, on by default, with lip sync
Not availableNo negative prompt, no style switch, no movement_amplitude, no bgm field on Q3 or Q4
Prompt limit5,000 to 20,000 characters on the Vidu API, 2,000 on fal.ai

How to Use the Prompt Generator

1

Pick the model and mode

Use Vidu Q3 Pro for text-to-video and start-plus-end frames. Use Vidu Q4 Preview for image-to-video and reference-to-video, or when you need 2K or 4K. For references, list each upload as @1, @2 and so on with its job.

2

Direct one shot and its sound

Describe one subject doing one thing in one place, pick a single camera move and a speed word, then write the dialogue, sound effects and whether you want music.

3

Generate, then draft at 720p

Paste the prompt into vidu.com, the Vidu API or fal.ai with the same duration and aspect ratio. Draft at 720p, or use off-peak mode on Q3 to halve the cost, then render your pick at full resolution.

Frequently Asked Questions

What is Vidu?

Vidu is the AI video model family from ShengShu Technology, developed with Tsinghua University. It runs on vidu.com, the Vidu API at platform.vidu.com, and hosts such as fal.ai. The current models are Vidu Q4 Preview, launched on October 7, 2026, and Vidu Q3, launched on January 30, 2026. Both generate video with native sound in one pass at 24fps and run up to 16 seconds per clip.

What is new in Vidu Q4 Preview?

Vidu Q4 Preview (API ID viduq4-preview) is ShengShu's new flagship. It accepts up to 15 reference images plus up to 3 voice or audio references, renders from 540p up to 2K and 4K, and runs 1 to 16 seconds. For now it only does image-to-video and reference-to-video. On the Artificial Analysis image-to-video arena it sits in a statistical tie for second place, alongside Seedance 2.0 and Gemini Omni Flash, behind MiniMax H3.

Which Vidu model should I pick?

For text-to-video and start-plus-end-frame clips, use Vidu Q3 Pro, or Q3 Turbo when speed matters more. For image-to-video, reference-to-video, large reference sets or anything above 1080p, use Vidu Q4 Preview. Q3 also has specialised reference models: viduq3-mix, viduq3-ad for adverts and viduq3-drama for character-driven short drama with named subjects and voices.

How should a Vidu prompt be structured?

Vidu's own formula is subject, then action, camera, scene and style, and Vidu says that order matters more than length. When framing is the point of the shot, put the shot call first, as in "Low-angle shot: a courier walks toward camera down a narrow alley at night." Keep each clip to one location, one subject and one action, and describe specific light ("side key light, cool shadows") instead of mood words.

How do I write camera movement for Vidu?

One camera action per generation, written as a modifier at the end of the shot, like "Slow push-in." Vidu's guides warn that stacking moves such as "slow dolly forward with slight pan left" makes the camera overcorrect. Pair the move with a single speed word, since subtle, gentle, slow and steady reduce variance. Zoom in is a movement, not an angle, so state the angle as well. For tight control, start and end frames constrain the camera better than words.

How do reference images work?

In reference-to-video, refer to uploads with the @ syntax in upload order: @1, @2, @3, or [@1] in brackets. Vidu's own example reads "@1 and @2 are eating hot pot together." Named subjects work too, as in "@Cowboy sits at the bar in @Bar." Give every reference a clear job, such as the character, the outfit, the product or the location. Q4 Preview takes up to 15 images but does not yet support named subjects, so numbered tags are safest there. Q3 takes up to 7.

Does Vidu generate sound and dialogue?

Yes. On Q3 and Q4 audio is on by default and covers dialogue, voiceover, sound effects and music, with lip sync. Write the sound explicitly, because a missing cue does not mean silence. Name the speaker, describe the delivery and quote the line. Keep lines to roughly two and a half words per second so they fit the clip, and use a close or medium shot for clean lip sync. If you want no music, say so in the prompt. Q3 lists English, Chinese and Japanese voices.

Is there a negative prompt, style or motion-strength setting?

Not on Q3 or Q4. The current Vidu API has no negative prompt field, no general or anime style switch, no movement_amplitude and no bgm parameter. Those appear only on older Q1 and Q2 endpoints on fal.ai. On the current models, style, motion strength, music and exclusions all go in the prompt text, which is exactly what this generator writes.

What durations, resolutions and aspect ratios does Vidu support?

Q3 and Q4 Preview both run 1 to 16 seconds at 24fps, with some reference models limited to 3 to 15 seconds. Q3 renders 540p, 720p or 1080p. Q4 Preview adds 2K and 4K. Aspect ratios are 16:9, 9:16, 1:1, 4:3 and 3:4, and image-to-video follows your input image.

How long can a Vidu prompt be?

It depends on where you run it. The Vidu API accepts 5,000 characters on most endpoints and 20,000 on Q4 and Q3 Pro, while fal.ai caps Q3 at 2,000 characters. Length is not the goal though. This generator keeps prompts tight, about 60 to 150 words for a single shot, so they paste anywhere.

What does Vidu cost?

On the Vidu API, pricing is in credits per second and scales with resolution. Q4 Preview runs from 9 credits per second at 540p to 78 at 4K, and Q3 Pro from 9 at 540p to 24 at 1080p. Off-peak mode roughly halves the price on Q3 in exchange for delivery within 48 hours.

Is this tool free to use?

Yes. You get 1 free prompt generation per day with no signup required. For unlimited access, sign up for a Promptslove membership which includes all AI tools and 29,000+ premium prompts.

More video prompt generators

Want Unlimited AI Prompt Generation?

Get unlimited access to all AI tools, 29,000+ premium prompts, courses, and resources designed to maximize your creative output.

Vidu is a trademark of ShengShu Technology. Promptslove is not affiliated with or endorsed by ShengShu. Model specifications and pricing described here reflect published documentation at the time of writing and may change.