Video · Text → video · Image → video
HappyHorse 1.1
Videos up to 15 seconds from text, an image, or several references — with consistent characters
About the model
HappyHorse 1.1 is Alibaba’s updated video model. Compared with the previous version, it follows prompts more closely, produces smoother motion and camera work, and keeps character identity and scene style noticeably more consistent.
Our service offers all three modes:
- Text-to-video — a clip built entirely from your description, with a choice of aspect ratio.
- Image-to-video — your image becomes the first frame and the model builds the motion.
- Reference-to-video — up to 9 images of characters, objects, and settings combined into one scene.
Length is adjustable from 3 to 15 seconds, at 720p or 1080p.
Strengths
- Three modes — from text, from an image, and from up to 9 reference images
- Keeps character identity, visual style, and scene details consistent
- Natural motion and confident camera work
- Realistic skin texture and stable on-screen text
- Clip length from 3 to 15 seconds, in one-second steps
- 720p or 1080p and nine aspect ratios — from 21:9 to 9:21
Best for
- Clips featuring the same character across different scenes
- Product videos made from existing photos
- Brand videos that keep the visual identity intact
- Short story-driven videos for social media
Limitations
- Maximum length is 15 seconds
- Input images up to 20 MB — JPEG, PNG, WEBP
- Reference images need a shortest side of at least 400 pixels
- For image-to-video, both sides must be at least 300 pixels, with an aspect ratio between 1:2.5 and 2.5:1
Prompting tips
- In reference mode, refer to images as [Image 1], [Image 2] — in upload order
- Say exactly what to take from a reference — for example, "the woman in a red dress in [Image 1]"
- Use clear reference images of 720p or higher, without heavy compression
- In image-to-video mode the prompt is optional, but describing motion and camera gives more precise results