Alibaba HappyHorse 1.1: Video Generation Upgraded Across Five Dimensions
Alibaba's HappyHorse 1.1 video generation model improves five dimensions including motion expressiveness and subject consistency, cuts 1080P pricing by 25%, and is now live on the Bailian platform.


Alibaba HappyHorse 1.1: Video Generation Upgraded Across Five Dimensions
Alibaba's HappyHorse 1.1 video generation model improves five dimensions including motion expressiveness and subject consistency, cuts 1080P pricing by 25%, and is now live on the Bailian platform.
Alibaba has released its latest video generation model, HappyHorse 1.1 (the domestic name literally means "Happy Little Horse 1.1"). Compared with its predecessor, it improves across five dimensions — motion expressiveness, subject consistency, instruction following, visual texture, and audio — while cutting the price of 1080P generation by 25%. It's live on the Alibaba Cloud Bailian platform and the official website; you can call the API directly or use it on the web.
What HappyHorse 1.1 is
HappyHorse is Alibaba's video generation model family, supporting both text-to-video and image-to-video. Version 1.1's technical specs match 1.0: 3-15 seconds per generation, 720p and 1080p resolutions, and free-form aspect ratios.

Core features
Motion expressiveness
The previous generation had issues with sluggish motion and weak pacing in some shots. Version 1.1 improves motion modeling and temporal consistency, making actions more coherent and forceful.
In a hands-on motorcycle riding scene: 1.1 generates footage at proper speed, obeying basic physics, and the scenery reflected in the windshield in close-up is fairly logical too. On the same task, 1.0 showed slow-motion problems, plus a motorcycle going the wrong way and a helmet reflection that didn't match the actual scene.
Subject consistency
HappyHorse 1.1 supports 9 character reference images as simultaneous input, flexibly combining product details, brand elements, characters, and scenes. For popular formats like multi-shot storyboards and N-grid image references, its understanding of reference images has also been strengthened.
In testing, three reference images depicting a particular person leaving their job were uploaded to generate a 10-second video. Version 1.1 accurately reproduced the person's face and clothing across two shots, kept scenes and character details stable and consistent, and left no flaws in the frame corners either.

Instruction following
In a surreal-scene test (a zero-gravity café, with every object required to obey real inertia and momentum), HappyHorse 1.1 reproduced the details one by one in prompt order. Still, such extreme scenes produced some slip-ups, like a character with a frozen expression and a chair appearing out of nowhere.
Visual texture
Version 1.1 largely fixes the previous generation's "greasy" look and over-sharpening. In frames involving large crowds (like a soccer match), the rendering of the main subjects improved, but background faces remain somewhat blurry, slightly lacking in realism and dynamism.
Audio
In a musical instrument performance test, version 1.1 showed no obvious improvement over 1.0 — the changes in the performance footage still don't line up with the changes in the audio.
Pricing
| Resolution | Standard price | Discounted price | vs 1.0 |
|---|---|---|---|
| 720P | ¥0.9/second | ¥0.54/second | unchanged |
| 1080P | ¥1.2/second | ¥0.72/second | down 25% |
How to use it
- Web experience: visit happyhorse.cn
- API access: through the Alibaba Cloud Bailian platform at bailian.console.aliyun.com

Use cases
- Advertising and marketing: 9 simultaneous reference images suit product detail showcases and brand element combinations
- Stylized content: stylized content like traditional Chinese painting and animation holds its style well, with no style drift
- Multi-shot storytelling: the N-grid image reference feature suits multi-shot story creation
💡 Tip: HappyHorse 1.1's upgrade is a minor-version iteration. Motion quality, character fidelity, and overall visual polish get solid optimization, but physics adherence in surreal scenes and audio sync still have room to improve.