Happy Horse 1.1 Prompt Generator
Generate optimized prompts for Alibaba's Happy Horse 1.1, the video model that writes picture and sound in a single forward pass. Covers lip-synced dialogue in seven languages, reference-to-video casting with up to 9 images, and 3 to 15 second clips at 720p or 1080p. Built around the director formula the model responds to: Subject, then Action, Environment, Style, Camera and Audio, kept deliberately short because long prompts hurt results here.
Describe the main visual scene — subjects, environment, mood, and key visual details
Generated Prompt
Fill in the form and click "Generate" to create an optimized Happy Horse 1.1 video prompt.
Tip: Describe the motion and temporal progression of your scene. Think in terms of "what happens over time" rather than a static description.
Happy Horse 1.1 Tips
- • Keep prompts tight — around 20 words per single shot. Happy Horse degrades on long, padded prompts.
- • Always write in this order: Subject → Action → Environment → Style → Camera → Audio. The model follows this hierarchy.
- • Write like a director: visible motion, concrete physical detail. Avoid literary or abstract descriptions.
- • Skip generic praise like "stunning", "masterpiece", "ultra detailed" — replace with specific cinematic facts.
- • For Image-to-Video, only describe what the image cannot show: motion, sound, expression changes, time.
- • For Reference-to-Video (new in 1.1), give each character one unmistakable detail and repeat it every beat — up to 9 reference images.
- • For Multi-Shot, use [0-3s] [3-7s] labeled blocks (or "Shot 1: / Shot 2:") so cuts land where you want.
- • Layer real-world audio that matches what is visible — sizzle, footsteps, breathing, dialogue in quotes.
- • There is no negative prompt parameter, so phrase exclusions as positive scene facts instead.
- • Prompts are capped at 2,500 characters on the fal.ai API. Seed is supported if you want repeatable runs.
Changelog
Version history for the model, drawn from Alibaba's announcements and the fal.ai API docs. Version 1.1 is the current release as of August 2026.
Happy Horse 1.1
CurrentJune 22, 2026- +Reference-to-video casting with 1 to 9 reference images, for holding characters and products consistent across a clip
- +More fluid motion in complex action scenes
- +Stronger subject consistency when fusing multiple reference sources
- +Better instruction following, with improved long-context scene planning
- +Higher visual quality, specifically facial detail and skin texture
- +Tighter audio sync and a richer range of generated sound elements
Happy Horse 1.0
April 2026- +Debuted anonymously on the Artificial Analysis Video Arena, then claimed by Alibaba on April 10, 2026
- +Held the number one spot on the text-to-video leaderboard at launch
- +Unified roughly 15B parameter transformer handling text, image, video and audio in one forward pass
- +3 to 15 second clips at 720p or 1080p with joint audio and video synthesis
- +Public gray test opened April 27, 2026
Happy Horse 1.1 specs and configs
What the model exposes, taken from the fal.ai API documentation.
| Parameters | Roughly 15B, unified single-stream self-attention transformer with no cross-attention |
|---|---|
| Resolution | 720p or 1080p. No 4K |
| Duration | 3 to 15 seconds, whole numbers, default 5 |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4, 4:5. Default 16:9 |
| Modes | Text-to-video, image-to-video, reference-to-video with 1 to 9 images |
| Audio | Generated in the same forward pass. Dialogue, ambience, Foley and music. Audio input is not supported |
| Lip-sync languages | Seven: English, Mandarin, Cantonese, Japanese, Korean, German, French |
| Prompt limit | 2,500 characters |
| Parameters exposed | prompt, aspect_ratio, resolution, duration, seed, safety checker |
| Not available | No negative prompt, no guidance scale, no step count. The model uses DMD-2 distillation at 8 sampling steps with no CFG |
| Speed | About 38 seconds to render a 5 second 1080p clip on a single H100 |
| Pricing | Around $0.14 per second at 720p and $0.18 per second at 1080p on fal.ai |
Frequently Asked Questions
What is Happy Horse 1.1?
Happy Horse is Alibaba's AI video model, built by the Future Life Lab team that sat under Taotian and has since moved into Alibaba's ATH AI Innovation unit. It debuted anonymously on the Artificial Analysis Video Arena in April 2026 and was claimed by Alibaba on April 10. Version 1.1 shipped June 22, 2026. It is a unified transformer of roughly 15B parameters that handles text, image, video and audio in one forward pass, producing 720p or 1080p clips of 3 to 15 seconds with joint audio and video synthesis and lip-sync in seven languages. It is a hosted model reached through an API, with fal.ai as the official partner, and there is no verified open weight release.
What changed from Happy Horse 1.0 to 1.1?
Alibaba describes five upgrades: more fluid motion in complex action scenes, stronger subject consistency across multiple reference sources, better instruction following with improved scene planning, higher visual quality in faces and skin texture, and tighter audio sync with richer sound. The headline addition is reference-to-video, where you supply 1 to 9 reference images to cast characters or products and hold them consistent. Core specs carry over unchanged from 1.0: 3 to 15 seconds, 720p or 1080p, and free choice of aspect ratio. Worth noting that some third-party writeups credit 1.1 with native audio and 1080p as new features, but Alibaba states those were already in 1.0.
What makes Happy Horse stand out from Sora, Veo, or Seedance?
Mainly that audio is produced in the same forward pass as the video rather than by a separate sound model, which is what makes its lip-synced dialogue land cleanly. Reference-to-video casting with up to 9 images is also strong for keeping a character or product stable. On ranking, be careful with what you read: Happy Horse held the number one spot on the Artificial Analysis text-to-video leaderboard in April 2026, and a lot of vendor pages still quote that as if it were current. As of August 2026 it sits fifth at 1148 Elo, behind Gemini Omni Flash, MiniMax H3, Seedance 2.0 and Wan2.7. Version 1.1 does score above 1.0, but the field moved faster.
What's the recommended prompt structure?
Subject → Action → Environment → Style/Composition → Camera Motion → Ambiance/Audio. This ordering is load-bearing: elements at the start anchor the visual subject, and camera + audio placed last get the most weight over motion and sound behavior. Plain-English prose beats tag lists, JSON, or weighted parentheses.
Why do shorter prompts work better?
Happy Horse uses a unified attention transformer where every token competes for rendering capacity. Long, padded prompts dilute attention and cause subject drift. The consensus guidance is roughly 20 words per single shot — a tight 20-word prompt typically beats a padded 60-word one.
How do I use multi-shot mode?
Select Multi-Shot Sequence and label the cuts explicitly — either timestamp blocks like "[0-3s] establishing wide of …", "[3-7s] cut to medium close-up of …", "[7-12s] pull-back reveals …", or "Shot 1 (wide, 0-1s): … Shot 2 (mid tracking, 1-4s): …". Keep the character description consistent across every block so identity persists across cuts.
How does reference-to-video (1.1) work?
Reference-to-video lets you attach up to 9 reference images so Happy Horse can cast specific characters or products and hold them consistent. In the prompt, give each subject one unmistakable identifying detail and repeat it in every beat (e.g., "the woman with reading glasses on a chain"), map them explicitly ("from the first reference image"), and close with a short "preserve the exact appearance of each character" clause.
How does the lip-sync work?
Add dialogue in quotation marks in the audio field and pick a lip-sync language. The generator notes the language inline (e.g., "She says in Japanese: ...") so Happy Horse's lip-sync head locks to the right phoneme set. It supports multilingual dialogue (English, Mandarin, Cantonese, Japanese, Korean, German, French, Spanish, and more).
Where can I actually run Happy Horse 1.1, and what does it cost?
Happy Horse 1.1 is available via API — fal.ai is the official partner, with text-to-video, image-to-video, and reference-to-video endpoints priced around $0.14/second at 720p and $0.18/second at 1080p. Version 1.0 is also on Replicate. There is no verified open-source weight release. This generator produces prompts you can paste into any of those; they also transfer well to other top video models.
Is there a Happy Horse 1.5?
No. As of August 2026 the latest release is 1.1, from June 22, 2026. We checked directly: fal.ai has no v1.5 endpoint and returns a 404 for one, Artificial Analysis lists only HappyHorse-1.1 and 1.0, and Chinese coverage from IT之家, TechNode and 量子位 stops at 1.1. The "Happy Horse 1.5" pages circulating online trace back to a single article that cites no sources and whose comparison table is really about subscription tiers rather than model capabilities. One site even publishes a page titled 1.5 at a URL ending in 1-1. If a genuine 1.5 ships we will update this page and the changelog above.
Does Happy Horse support negative prompts?
No. The API exposes prompt, aspect ratio, resolution, duration, seed and a safety checker, and that is all. There is no negative prompt, no guidance scale and no step count, which follows from the DMD-2 distilled design that runs 8 sampling steps with no classifier-free guidance. Phrase exclusions as positive scene facts instead, so "bare concrete walls" rather than "no decorations". Seed is available if you want to repeat a run.
Are the model weights open source?
In practice, no, though the picture is muddier than most coverage suggests. Alibaba publicly committed to releasing weights under Apache 2.0, and a Hugging Face page for 1.0 carries an Apache 2.0 label. But multiple independent checks have found no downloadable weights, no working GitHub repository and no license file, with everything marked as coming soon. Plenty of headlines call this the leading open-source video model, which the actual record does not support. Treat it as hosted and API-only until weights genuinely appear.
Is this tool free?
Yes. You get 3 free generations per day. For unlimited access, sign up for a Promptslove membership.
Want Unlimited AI Prompt Generation?
Get unlimited access to all AI tools, 20,000+ premium prompts, courses, and resources.
