authustle
3 min readUpdated Aug 2, 2026

What Is HappyHorse? Alibaba's Video Model, Explained From the Schema

HappyHorse 1.1 does three jobs and conspicuously not a fourth. The actual parameters, the two defaults that waste your first generation, and where it beats better-known models.

HappyHorse is Alibaba's video generation model, and it is the one doing most of the quiet work in vertical creator content in 2026 — well known to people building pipelines, largely unknown to people buying subscriptions. Here is what it actually is, taken from the endpoint schema rather than the marketing.

Three endpoints, and one that does not exist

HappyHorse 1.1 ships text-to-video, image-to-video and reference-to-video. There is no video-edit endpoint.

That absence is the single most important fact about it. HappyHorse cannot ingest a source video and modify it, which means it cannot copy a specific clip's choreography and cannot replace a performer in existing footage. If that is your job, this is the wrong model — use Kling Motion Control or Wan 2.7 edit-video instead.

The 1.1 release in June 2026 added native audio, multilingual lip-sync, up to nine character references and 1080p. It did not add video editing, and people waiting for that upgrade should stop.

What reference-to-video actually takes

From the endpoint schema:

ParameterValuesDefault
image_urls1–9 reference imagesrequired
promptup to 2,500 charactersrequired
duration3–15 seconds5
resolution720p, 1080p1080p
aspect_ratio16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4, 4:516:9
seed0–2147483647random
enable_safety_checkerbooleantrue

Reference images must be JPEG, PNG or WEBP, shortest side at least 400px (720p or better recommended), max 20 MB each.

The two defaults that waste your first run

aspect_ratio defaults to 16:9. For a model whose main audience is vertical creators, that is a landscape clip you cannot post. Set it to 9:16 for Reels, TikTok and Shorts, or 4:5 for feed. Nine ratios are supported, including both vertical options — the full set of verticals is genuinely rare and one of the strongest reasons to pick this model.

duration defaults to 5 seconds — which is actually the right default, because identity holds best in the 5–8 second band. Resist raising it.

The thing nobody explains: how references work

You do not just upload photos. You cite them in the prompt as character1 through character9, in the order they appear in image_urls. So the prompt is not "a woman in a kitchen" — it is a description that names character1 and says what they do.

This is why supplying one frontal selfie produces poor results and why nine references produce good ones: the model is matching a named subject across angles, and you have given it one angle to work from.

Price and positioning

$0.14 per second on both image-to-video and reference-to-video, read live from fal.ai on 2 August 2026. Flat rate — the endpoint charges the same whether you request 720p or 1080p, so quality is not where you economise.

A 5-second clip is $0.70; the 15-second maximum is $2.10.

Against alternatives at the same job: Kling 3.0 Pro image-to-video is also $0.14/s; MiniMax H3 reference-to-video is $0.26/s. HappyHorse's edge is the aspect-ratio coverage and native multilingual lip-sync in the same call.

When to use it

Yes: you have photos of a person and want them in a scene you describe; you want a talking-head clip with lip-sync in one call; you need real 9:16 or 4:5 output.

No: the movement has to match an existing clip, or the original background has to survive. Neither is what this model does, and no prompt makes it.

happyhorsealibabamodel explainerreference to videofal.ai

Keep reading