Alibaba Happy Horse
Video generation family covering text-to-video, image-to-video, reference-to-video, and video editing with native audio.
Studio is in early access; availability resolves per project after sign-in.
- Endpoints
- 4
- Eligible
- 4
- Input schemas imported
- 4 / 4
- Delivery
- fal.ai
- Weights
- Not stated
- Developer
- Not stated in source
Overview
Happy Horse covers four video tasks: text-to-video, image-to-video, reference-to-video, and video editing. The image-to-video endpoint animates stills into 1080p video with synchronized native audio, Foley sounds, and multilingual lip-sync. Input images need at least 400px on the shortest side, with 720p or higher recommended. Lip-sync covers English, Mandarin, Cantonese, Japanese, Korean, German, and French.
Capabilities
- Four video tasks: text, image, and reference inputs plus editing
- 1080p motion output with synchronized native audio
- Multilingual lip-sync across seven languages
- Foley and ambient sound rendered with motion
Best for
- Stills-to-video animation
- Talking-portrait clips
- Sound-synced motion content
Use cases
- Animating stills with native sound
- Multilingual lip-synced video
- 720p-plus source enhancement into motion
Endpoints
4 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.
| Endpoint | Task | Catalog status | Inputs |
|---|---|---|---|
alibaba/happy-horse/image-to-video | image to video | Eligible | 6 fields · requires image_url |
alibaba/happy-horse/reference-to-video | reference to video | Eligible | 7 fields · requires prompt, image_urls |
alibaba/happy-horse/text-to-video | text to video | Eligible | 6 fields · requires prompt |
alibaba/happy-horse/video-edit | video editing | Eligible | 7 fields · requires video_url, prompt |
Questions
What inputs can Happy Horse animate?
Text prompts, still images, reference video, plus direct video editing.
Which languages does lip-sync support?
English, Mandarin, Cantonese, Japanese, Korean, German, and French.