Google Omni 1.1 Prompt Generator
Generate optimized prompts for Gemini Omni 1.1 Flash — Google's any-to-any video model that turns text, images and video into high-resolution clips with native synchronized audio. Built on Google's official prompting documentation, with 1.1's new scene extension, first/last-frame control, video references and draft mode baked in.
Describe the main visual scene — subjects, environment, mood, and key visual details
Generated Prompt
Fill in the form and click "Generate" to create an optimized Gemini Omni 1.1 Flash video prompt.
Tip: Describe the motion and temporal progression of your scene. Think in terms of "what happens over time" rather than a static description.
Gemini Omni 1.1 Flash Tips
- • Quotation marks are for on-screen text, not dialogue. Write speech unquoted after a colon — "A woman says: My name is Clara." Quoting it burns the line into the frame as text.
- • Omni multi-shots by default. If you want one unbroken take you have to say so: "In a single continuous shot" or "No scene cuts."
- • Less is more — 30 to 80 words is the sweet spot. Google says Omni needs less prescription than Veo; past ~100 words you start fighting the model.
- • Skip pixel-level camera specs. ISO, f-stop and mm focal lengths actively degrade coherence. "Shallow depth of field" works; "85mm at f/1.4" does not.
- • Always state audio intent — even if it is "natural ambient sound only, no music, no dialogue." Explicit negation stopped unwanted stock music in about half of documented tests.
- • For edits, one instruction per turn, then the exact phrase "Keep everything else the same." Drift starts around turn 5.
- • 16:9 and 9:16 only — there is no 1:1 on Omni. Clips run 3-10s each, extendable to 40s total.
- • Draft at 360p, finish at 4K. Same prompt, one tenth the cost while you explore variations.
Gemini Omni 1.1 Prompt Templates
Copy-ready shot briefs built around what this model actually rewards — including Google's own published prompts. Swap the [BRACKETED] parts for your own scene, then paste straight into the tool.
Whip-Pan Reveal (Google 1.1 demo)
Google OfficialGoogle's own 1.1 launch prompt — a designed camera move with two subjects.
A close-up low-angle shot of [SUBJECT A, e.g. a stylish drummer in a beige suit playing a red drum kit] in [LOCATION, e.g. a grand hall] transitions as the camera whip-pans to the side, revealing [SUBJECT B, e.g. an older saxophonist] alongside [SUBJECT C, e.g. a ballet dancer spinning in a white outfit] under [LIGHTING, e.g. soft purple stage lights]. 16:9, 10 seconds, natural performance audio, no dialogue.
Claymation Knowledge Explainer
Google OfficialGoogle's flagship world-knowledge prompt. The word "accurate" does the real work.
claymation explainer of [SCIENTIFIC PROCESS, e.g. protein folding], everything is made out of clay, no hands, stop motion, accurate Sound design: soft tabletop foley only. No music, no dialogue. In a single continuous shot.
Kinetic Typography Hook
Text on ScreenQuoted glyphs render on screen. Note there is no dialogue here at all.
One word on the screen at a time: "[WORD 1], [WORD 2], [WORD 3], [WORD 4], [WORD 5]" Each word appears for 1s with a different animated style, perfect pacing to a rhythm. Background: [CLEAN BACKDROP, e.g. a matte black studio]. Audio: soft electronic clicks timed to each word. No dialogue. No music. 9:16 vertical, 8 seconds.
Diegetic In-World Text
Text on ScreenText that lives inside the scene — signs, storefronts, plates. Quote the exact glyphs.
[SCENE DESCRIPTION, e.g. A slow push down a rain-slicked street at dusk]. There is a street sign that says: "[SIGN TEXT]", there is a storefront that says: "[STOREFRONT TEXT]", there's a [OBJECT] with the [LABEL TYPE]: "[LABEL TEXT]". Keep all on-screen text correct and readable. Sound design: [AMBIENT SOUND]. No dialogue.
Dialogue — Unquoted, After a Colon
AudioThe rule most guides get wrong. Quotes render text; a colon speaks the line.
[SPEAKER DESCRIPTION, e.g. A woman in her thirties wearing a grey overcoat] in [SETTING], [FRAMING, e.g. medium shot, eye level]. In a voice that is [VOICE QUALITIES, e.g. crisp and clear, with a thoughtful, analytical tone and a standard American accent], [SPEAKER NAME] says: [LINE OF DIALOGUE, UNQUOTED] Sound design: [AMBIENT BED] kept low under the voice. No music. In a single continuous shot.
Surgical Conversational Edit
EditingGoogle's documented pattern — short instruction plus one exact preserve phrase.
[ONE CHANGE, e.g. Make the phone invisible. / Change the car color to metallic blue. / Put a fashionable hat on this person.] Keep everything else the same.
Event-Conditional Edit (two turns)
EditingTurn 1 establishes timed beats, turn 2 edits relative to them. A documented power move.
Turn 1: [SCENE, e.g. A person stands in front of a mirror]. At 0:05, [TRIGGER ACTION, e.g. she touches the mirror] then she does it again at 0:07, 0:08 and 0:09. Turn 2: [WHAT CHANGES, e.g. The style of the video changes] every time [TRIGGER ACTION, e.g. the person touches the mirror].
Single-Variable A/B Chain
EditingOne variable per turn gives you a clean matched set instead of unrelated renders.
Turn 1: [BASE SCENE, e.g. A person opening a sleek black product box on a wooden desk in a minimalist apartment]. [LIGHTING, e.g. Morning light]. [CAMERA, e.g. Slow zoom in]. In a single continuous shot. Turn 2: Make it [VARIABLE 1 CHANGED, e.g. sunset light instead of morning]. Keep everything else the same. Turn 3: Now make [VARIABLE 2 CHANGED, e.g. the box white instead of black]. Keep everything else the same.
Scene Extension (new in 1.1)
ExtensionExtends past 10s toward the 40s cap. Timecodes reset to the new portion.
Continue the video. [WHAT HAPPENS NEXT, e.g. The camera slightly pans and we now see she is talking to a man with curly hair, we see the man's back]. [SPEAKER] says: [LINE, UNQUOTED, IF ANY] Audio: [HOW THE SOUND CHANGES, e.g. the music continues into the chorus]. Keep the character, lighting and location consistent with the previous shot.
Extension — Cinematic Dolly Zoom
ExtensionGoogle's 1.1 demo. The move that used to break at the extension seam.
Continue the video. Execute a cinematic optical dolly-zoom shot. The camera dollies forward while simultaneously zooming out, keeping [SUBJECT]'s [EXPRESSION/POSE, e.g. frozen shocked face] locked at the exact same size. Audio: [SCORE OR AMBIENCE, e.g. a rising dramatic score under room tone]. No dialogue.
First + Last Frame Interpolation
ReferencesYou own both ends; Omni invents only the middle. Same image twice = seamless loop.
[# Sources <FIRST_FRAME>@[YOUR START IMAGE] <LAST_FRAME>@[YOUR END IMAGE]] A smooth cinematic transition from [WHAT THE FIRST FRAME SHOWS] to [WHAT THE LAST FRAME SHOWS]. In between, the camera [EXACT TRANSITION PATH, e.g. pushes forward and arcs gently right]. No jump cuts. Use this image as the starting frame. [DURATION], [ASPECT RATIO], [AUDIO INTENT].
Multi-Reference Sequence with Timecodes
ReferencesGoogle's official 6-reference example. Tags are 0-indexed.
[0-3s] [SCENE TYPE, e.g. A studio fashion sequence]. Starting with [SUBJECT] <IMAGE_REF_0>, [ACTION] <IMAGE_REF_1> [3-6s] Then we see [SUBJECT] <IMAGE_REF_2> [ACTION] <IMAGE_REF_3> [6-10s] And finally [SUBJECT] <IMAGE_REF_4> [ACTION] <IMAGE_REF_5> Use the given images as references for video generation. The images should not be used as literal initial frames.
Role-Assigned Multi-Modal Mix
ReferencesLabel what every reference is FOR — unlabelled references get blended together.
[# References <IMAGE_REF_0>@[IDENTITY] <IMAGE_REF_1>@[PRODUCT] <VIDEO_REF_0>@[CAMERA MOVE]] [SUBJECT] <IMAGE_REF_0> [ACTION] holding [PRODUCT] <IMAGE_REF_1>, with the camera language of <VIDEO_REF_0>. Use the given image(s) as references for video generation. The images should not be used as literal initial frames. Use the given video(s) as references. Do not use them as a source for video editing.
Motion Transfer — Identity from Image
ReferencesVideo reference carries motion, image reference carries identity. Separates performance from casting.
Apply the pose and motion from <VIDEO_REF_0> to the character in <IMAGE_REF_0>. Place them in [ENVIRONMENT]. Keep the character's face, hair and wardrobe exactly as the reference image establishes. [CAMERA]. Audio: [SOUND]. Use the given video(s) as references. Do not use them as a source for video editing.
Sketch-to-Video
Google OfficialThe final clause is what stops your drawing appearing in the output.
turn this into [REALISM LEVEL, e.g. realistic footage], using the drawing only as a guide for movement, do not show the drawing in the final video [SUBJECT AND SETTING DETAIL]. [LIGHTING]. Sound design: [AUDIO]. No dialogue.
Single Continuous Shot (oner)
CinematicOmni multi-shots by default — continuity is opt-in and needs a literal phrase.
[SCENE AND SUBJECT, e.g. A young product designer sits at a small desk beside a rainy window and opens a sketchbook]. [ONE OR TWO ACTIONS MAXIMUM]. The camera [CAMERA MOVE, e.g. starts close on the pencil tip, pulls back to a medium shot, then orbits gently left]. [LIGHTING, e.g. Warm desk lamp light, cool blue rain outside, shallow depth of field]. In a single continuous shot. No scene cuts. Sound design: [AMBIENCE]. No music, no dialogue. 16:9, 10 seconds.
Multi-Shot Sequence with Timecodes
CinematicWhen you actually want cuts, author the beats explicitly rather than leaving it to chance.
[0-3s] [BEAT ONE] [3-6s] [BEAT TWO] [6-10s] [BEAT THREE] Consistent throughout: [WHAT MUST NOT CHANGE, e.g. the same character, the same warm tungsten lighting, the same location]. [STYLE]. Sound design: [AUDIO ACROSS ALL BEATS]. 16:9, 10 seconds.
Physics-Grounded Demonstration
ExplainerOmni simulates momentum, deformation and friction rather than animating an approximation.
[OBJECT] [PHYSICAL ACTION, e.g. bouncing on a hardwood court]. Each [EVENT] has realistic physics: [SPECIFY — e.g. ball deformation on impact, friction slowing the bounces naturally]. [SPEED, e.g. real time, not slow motion]. Camera: [ANGLE AND MOVE]. Lighting: [LIGHTING]. Sound design: realistic impact audio synchronised to the moment of contact. No music. In a single continuous shot.
Explainer with Native Narration
ExplainerNarration is generated in the same pass and timed to the animation — no separate VO step.
A [STYLE, e.g. skeuomorphism stop motion] explainer about [TOPIC, e.g. how the brain hippocampus works] with a compelling voiceover. Accurately visualize [WHAT MUST BE CORRECT]. Camera: [FRAMING]. Lighting: [LIGHTING]. A narrator says: [SHORT SCRIPT, UNQUOTED, SIZED TO THE CLIP LENGTH] Quiet ambient pad underneath. No sound effects. 16:9, 10 seconds.
Product Hero from One Photo
Product & AdsImage-to-video keeps the real SKU. Prompt motion only — do not re-describe the photo.
[# Sources <FIRST_FRAME>@[YOUR PRODUCT PHOTO]] The [PRODUCT] rotates slowly on [SURFACE, e.g. a matte studio pedestal], [LIGHTING, e.g. soft key light from the left], shallow depth of field. Keep the logo sharp and readable throughout the rotation. In a single continuous shot. Sound design: [AUDIO, e.g. calm ambient bed]. No dialogue. [ASPECT RATIO], [DURATION].
Audience Variant from One Master
Product & AdsConversational state holds the product constant while pace and palette are re-targeted.
Turn 1: Generate a [DURATION] product ad for [PRODUCT]. [TREATMENT, e.g. Cinematic close-ups, premium feel, deep blacks]. [AUDIO MOOD]. End on the product alone, centered. Turn 2: Now generate a variant of the same ad targeted at [AUDIENCE]. [NEW TREATMENT, e.g. Faster cuts, brighter colors, social-feed energy]. Keep the product exactly the same.
UGC-Style Creator Ad
Product & AdsHandheld realism plus lip-synced delivery — the format that reads as genuine.
[DURATION] clip in casual handheld [DEVICE] vlog style. [PERSON DESCRIPTION] walks down [LOCATION] holding the camera in selfie mode, talking energetically about [TOPIC]. Natural skin texture, no filter, slight motion blur. Background: [BACKGROUND]. She says: [SHORT SCRIPT, UNQUOTED — keep it to what fits the clip length] Sound design: [STREET AMBIENCE] under the voice. No music. 9:16 vertical.
Restyle While Preserving Motion
EditingStyle transfer framed as a verb on the existing scene. Use one style reference only.
Take the attached clip and restyle it as [TARGET STYLE, e.g. a hand-painted animated scene]. Preserve the camera motion and timing exactly. Replace [WHAT CHANGES, e.g. the modern setting with a 1980s countryside]. Keep the subject's posture and expression coherent across frames. Minimal version: Remake this video in [STYLE] aesthetic. Keep everything else the same.
Vertical Reel from Reference Stills
SocialMulti-image reference preserves the real room, signage and products.
[# References <IMAGE_REF_0>@[WIDE SHOT] <IMAGE_REF_1>@[DETAIL] <IMAGE_REF_2>@[HERO ITEM]] A [DURATION] vertical 9:16 clip set in [PLACE] from the reference images. Handheld, [LIGHTING, e.g. natural morning light]. The camera drifts from [START POINT], past [MIDPOINT], and settles on [END POINT]. Keep the interior, signage and products exactly as in the references. In a single continuous shot. Sound design: [SPECIFIC AMBIENT SOUNDS]. No music, no dialogue.
Trigger-Anchored ASMR Sound Design
AudioAttaching a sound to a specific visual event beats vague audio direction measurably.
[SCENE, e.g. A hand moves slowly through a wall of ferns in a dim greenhouse]. Add [SOUND, e.g. soft harp notes] synchronized to exactly when [TRIGGER EVENT, e.g. the hand touches each fern leaf]. Underneath: [AMBIENT BED, e.g. faint dripping water and distant birdsong]. No music, no dialogue, no narration. Camera: [FRAMING]. In a single continuous shot. [ASPECT RATIO], [DURATION].
Interior Walkthrough from a Photo
CinematicReal-estate and architecture — the camera arc Google's own Home Tour demo showcased.
[# Sources <FIRST_FRAME>@[YOUR ROOM PHOTO]] A single continuous [DURATION] interior walkthrough starting from this photograph. The camera pushes slowly forward through [ROOM], arcs [DIRECTION] past [FEATURE], and pulls back to reveal [FINAL REVEAL]. [LIGHTING, e.g. Late-afternoon golden-hour light]. No people. Keep the furniture, flooring and fixtures exactly as photographed. Use this image as the starting frame. Sound design: quiet room tone only. No music, no dialogue. No scene cuts.
What's new in Omni 1.1
Released August 27, 2026. Google frames 1.1 as a suite of creative controls rather than a new base model — but the controls change how you write prompts.
Scene extension to 40 seconds
The headline change. Clips extend past the old 10-second ceiling to a 40-second cumulative maximum, and 1.1 reads up to 10 seconds of prior context when extending — versus roughly the final second in earlier models. Append-only, and the timecode origin resets in each extension.
First-frame / last-frame control
Specify the opening and closing frames of a shot and Omni interpolates between them. This is also how camera control is exposed — orbits, dollies and transitions — rather than through a dedicated parameter. Use the same image in both slots for a seamless loop.
360p draft mode
Up to 60% faster at about a third the cost of 720p. Google's only official speed claim for 1.1, and the basis of the draft-then-finish workflow their own Draft Room demo is built around.
1080p and 4K output — upscaled
New high-resolution tiers, but Google's API docs label both "upscaled" rather than natively generated. Worth knowing before you pay 1.5x or 3x the per-second rate.
Working video references
Up to 3 video clips of 3 seconds each as reference input. At the June launch these were accepted but not correctly processed; 1.1 makes reference-to-video with clips a real path. Audio inside a video reference is ignored.
What did NOT change
Google claims no improvement to character consistency or text rendering. The model card updated for 1.1 still lists both as open limitations — along with complex motion. Third-party claims of consistency gains in 1.1 have no source.
What is Gemini Omni?
Gemini Omni is Google DeepMind's any-to-any generative video family, announced at Google I/O on May 19, 2026 under the banner "create anything from anything." It collapses a previously fragmented multimodal stack — text-to-video, image-to-video, video editing, audio generation — into a single model with a single editing surface, combining an intuitive grasp of physics (gravity, kinetic energy, fluid dynamics) with Gemini's knowledge of history, science and cultural context.
Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) shipped on August 27, 2026. It runs in the Gemini app, Google Flow, YouTube Shorts, the YouTube Create app, Google Vids, Google AI Studio, the Gemini API and Gemini Enterprise Agent Platform. On the Artificial Analysis video arena, Omni Flash currently sits second on text-to-video and fourth on image-to-video — competitive rather than dominant, and comfortably ahead of Veo 3.1 on both boards.
A note on accuracy: much of the third-party writing about this model repeats claims Google has never made — a "Gemini Omni Pro" variant, consistency improvements in 1.1, specific audio codecs, per-generation credit costs. None of those have a primary source. Everything on this page is drawn from Google's own documentation, model card and pricing pages.
Core Capabilities
Any-to-Any Input
Mix text, images and video references in one prompt — Omni reasons across all of them. Audio reference upload is not yet supported in the API.
Conversational Editing
Stateful multi-turn edits. The model remembers the video and applies changes while preserving what you did not mention. Reliable to about 4 turns.
World-Knowledge Grounding
Accurate explainers for technical topics — physics, biology, history, chemistry, architecture — from a named metaphor rather than exhaustive description.
Physics Reasoning
Gravity, kinetic energy, fluid dynamics and cloth simulation are modelled, not approximated. Internal reasoning plans motion before frames are synthesized.
Native Synchronized Audio
Ambient sound, SFX, dialogue and music generated in the same forward pass as the video — not dubbed on afterwards.
Style Transfer
Apply claymation, anime, watercolour, risograph or chalk-on-blackboard while preserving the original motion. One style reference at a time.
SynthID + C2PA
Every output carries an imperceptible pixel and audio watermark plus C2PA Content Credentials, verifiable in the Gemini app, Chrome and Google Search.
Avatar Mode
Create videos featuring a digital version of yourself with your own voice. Enrolment requires explicit consent and is 18+; a single selfie is not enough.
Omni 1.1 Flash Specs
Straight from Google's API documentation and pricing pages. These are the constraints the generator writes against.
| Model ID | gemini-omni-1.1-flash |
| Released | August 27, 2026 |
| Duration | 3–10s per generation, extendable to 40s cumulative |
| Frame rate | 24 FPS |
| Aspect ratios | 16:9 (default) and 9:16 only — no 1:1 |
| Resolutions | 360p · 720p (default) · 1080p (upscaled) · 4K (upscaled) |
| Task modes | text_to_video · image_to_video · reference_to_video · edit · extend |
| Reference tags | <FIRST_FRAME> · <LAST_FRAME> · <IMAGE_REF_N> · <VIDEO_REF_N> |
| Video references | Max 3 clips × 3 seconds each; reference audio ignored |
| Audio reference upload | Not supported in the current API version |
| Unsupported params | No negative prompt, no seed, no temperature, no system instructions |
| API pricing | ~$0.03 / $0.10 / $0.15 / $0.30 per second (360p / 720p / 1080p / 4K) |
| Free access | Google Flow free tier (50 daily credits); YouTube Shorts & Create at no cost |
| Provenance | SynthID watermark + C2PA Content Credentials on every output |
The Official Omni Prompting Framework
Drawn from Google DeepMind's Omni prompt guide, the Gemini API documentation and Google's video best-practices pages — plus the measured findings from published hands-on testing.
Quotes render text; colons speak
The highest-impact rule on Omni, and the one nearly every guide gets wrong. Google: "To prevent the model from rendering text in the video, use a colon (:) after the speaker's action to denote speech and avoid using quotation marks." Write "A woman says: My name is Clara." for audio. Save double quotes for glyphs you want in frame.
Less prescriptive than Veo, on purpose
DeepMind states the divergence outright: "With Veo, you need to share precise instructions to get the best results. But with Gemini Omni, you don't have to be as prescriptive." Cover the five elements — framing and motion, style, lighting, location, action — in 30 to 80 words and let world knowledge fill the rest. Ported Veo prompts generally need trimming by 40 to 60%.
Single shot is opt-in
Omni builds a short multi-shot narrative by default. If you want one unbroken take you must ask: "In a single unbroken scene," "In a single continuous shot," or "No scene cuts." For deliberate sequences, author beats with timecodes — "[0-3s] ... [3-6s] ..." — but never chain three distinct events in one short clip; Google documents that as producing muddled output.
Reference tags assign roles
Official 0-indexed syntax: <FIRST_FRAME>, <LAST_FRAME> (only valid alongside FIRST_FRAME), <IMAGE_REF_N>, <VIDEO_REF_N>. For complex mixes, declare them up front: "[# References <IMAGE_REF_0>@character <VIDEO_REF_0>@camera_move]". Label what each reference is for — unlabelled references get blended. In image-to-video, prompt motion only; re-describing the image degrades results.
Edits are surgical, and there is one preserve phrase
Google: "Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes." The documented idiom is exactly "Keep everything else the same." One instruction per turn. Their own example: not a 40-word removal request, just "Make the phone invisible. Keep everything else the same." Reliable ceiling is around 4 turns before drift compounds.
World knowledge is the moat
For explainers, name the concrete visual metaphor and let Gemini supply accurate detail. Google's flagship prompt is six words of style plus one crucial word: "claymation explainer of protein folding, everything is made out of clay, no hands, stop motion, accurate." Use cultural touchstones, historical eras and scientific terms directly instead of granular description.
Audio is generated, not dubbed
Always state audio intent — even when the answer is "natural ambient sound only, no music, no dialogue." Google's own examples use a "Sound design:" prefix. Anchor sound to visual events where you can ("harp notes synchronized to each time she touches a leaf"); trigger-anchored audio measurably beats vague direction, and explicit negation stopped unwanted stock music in roughly half of documented tests.
Draft at 360p, finish at 4K
The 1.1 draft tier is 60% faster at about a third the cost. Run 3 or 4 variations changing one variable each, compare, then render the winner at high resolution — $0.03/sec versus $0.30/sec is a 10x saving on the exploration phase. Google's own "Draft Room" demo app is built around this exact loop.
Prompts Published by Google
Verbatim prompts from Google's own launch posts, DeepMind pages and API documentation — the clearest signal of what this model expects.
Eight Documented Omni Mistakes
Every one of these is named in Google's own documentation or reproduced across published hands-on testing.
A woman says: "My name is Clara."
A woman says: My name is Clara.
Quotation marks tell Omni to render glyphs on screen. Quoting dialogue is why lines end up burned into the frame instead of spoken.
Please remove the cell phone that the person is holding in their hand and fill in the background so it looks like they are just holding their hand empty.
Make the phone invisible. Keep everything else the same.
Google’s own before/after. Overly descriptive edit prompts confuse the diffing engine and cause unintended changes.
Shot on 85mm at f/1.4, ISO 400, anamorphic.
Shallow depth of field, warm practical lighting.
Pixel-level camera specs actively degrade coherence on Omni. Cinematographic facts steer it; technical metadata does not.
Assuming you get one unbroken take.
In a single continuous shot. No scene cuts.
Omni builds a short multi-shot narrative by default. Continuity is opt-in, and there are three literal phrasings that trigger it.
A negative prompt field with "wall, frame, blurry".
No text overlay on screen. No extra sound effects.
Omni has no negative-prompt parameter. Negatives go inside the prompt as instructions — the exact opposite of Veo’s convention.
She walks in, then sits down, then the lights change, then she stands.
She walks in and sits down as the lights dim.
Chaining three or more distinct events in one short clip is documented to produce muddled or incomplete video. One or two actions per 10 seconds.
Re-describing everything already visible in your reference image.
She turns toward the window and smiles.
In image-to-video, redundant description confuses the model. Prompt motion only and refer to the subject generically.
A 1:1 square clip for Instagram.
9:16 vertical, or 16:9 landscape.
Omni supports 16:9 and 9:16 only. There is no square output — crop in post if you need one.
Frequently Asked Questions
What is Gemini Omni 1.1 Flash?
Gemini Omni 1.1 Flash (model ID gemini-omni-1.1-flash) is Google's any-to-any generative video model, released on August 27, 2026. It accepts text, images and video freely mixed in a single prompt and outputs video with native synchronized audio generated in the same pass. The original Gemini Omni was announced at Google I/O on May 19, 2026 and reached the API on June 30. Version 1.1 is framed by Google as "a new suite of creative controls and generative video capabilities" rather than a new base model.
What is new in Omni 1.1 versus the original Omni Flash?
Five things. (1) Scene extension — clips now extend to a 40-second cumulative maximum, and 1.1 reads up to 10 seconds of prior context when extending, versus roughly the final second in earlier models. (2) First-frame / last-frame control for keyframe interpolation and seamless loops. (3) A 360p draft mode that is up to 60% faster at about a third the cost of 720p. (4) 1080p and 4K output — though both are upscaled, not natively generated. (5) Working video references, up to 3 clips of 3 seconds each. Notably, Google claims no improvement to character consistency or text rendering; the model card updated for 1.1 still lists both as open limitations.
How long can Gemini Omni 1.1 videos be?
Each generation produces 3 to 10 seconds at 24 FPS. From there you extend in further passes up to a 40-second cumulative maximum. Extension is append-only — you cannot prepend or insert mid-clip — and the timecode origin resets, so "after 2s" in an extension prompt means 2 seconds into the new portion, not the original clip.
Should dialogue go in quotation marks?
No — and this is the rule most prompt guides get backwards. Google's best-practices documentation says to use a colon after the speaker and avoid quotation marks, because quotes cause the model to render the line as on-screen text instead of speaking it. Write "A woman says: My name is Clara." unquoted for audio. Reserve double quotation marks for text you actually want rendered in frame, like a street sign or a kinetic-typography sequence.
What aspect ratios and resolutions does Omni 1.1 support?
16:9 (default) and 9:16 only — there is no 1:1, 4:3 or 21:9 option. Resolutions are 360p, 720p (default), 1080p and 4K, with the caveat that Google's own API docs label 1080p and 4K as "upscaled" rather than natively generated.
Does Gemini Omni generate audio, and can I upload an audio reference?
Omni generates synchronized native audio — ambient sound, SFX, dialogue and music — in the same forward pass as the video. But uploading an audio reference is not supported in the current API version, so describe the audio you want in words instead. Audio input does work on the consumer surfaces (Google Flow, the Gemini app). Voice editing is not supported anywhere, and you cannot add dialogue when extending an uploaded video of someone talking.
How do negative prompts work on Omni?
There is no negative-prompt parameter at all — nor seed, temperature, or system instructions. Google's guidance is to put negatives directly in the prompt as instructions: "No dialogue," "No text overlay on screen," "No scene cuts." This is the opposite of Veo, where a separate negativePrompt field takes bare noun lists and the word "no" is discouraged. Mixing the two conventions up is the most common Omni prompting mistake.
What does Omni 1.1 cost?
API billing is token-based at $17.50 per million video output tokens, which Google translates to roughly $0.10 per second of 720p video. Effective rates work out to about $0.03/sec at 360p, $0.10/sec at 720p, $0.15/sec at 1080p and $0.30/sec at 4K — a 10x spread that makes the draft-at-360p, finish-at-4K workflow worth building around. There is no free API tier, but Google Flow has a free tier with 50 daily credits, and YouTube Shorts and the YouTube Create app offer Omni at no cost with no subscription.
Where can I use the generated prompts?
Paste them into the Gemini app (Google AI Plus, Pro or Ultra), Google Flow, YouTube Shorts, the YouTube Create app, Google Vids, Google AI Studio, the Gemini API, or Gemini Enterprise Agent Platform. Note that editing or extending uploaded video is unavailable in the EEA, Switzerland and the UK — model-generated clips can be edited everywhere.
Is this tool free?
Yes! You get 1 free Gemini Omni prompt generation per day. For unlimited generations across all 30+ AI tools, sign up for a Promptslove membership.
Want Unlimited Omni Prompt Generation?
Get unlimited access to the Google Omni Prompt Generator, all 30+ AI tools, 30,000+ premium prompts, courses, and resources.
