Every AI Video Model That Shipped This Autumn (September 2026)
Seedance 2.5, MiniMax H3 Max, Gemini Omni Flash 1.1, Wan 3.0, LTX 2.5, Kling 4K, PixVerse V6, FLUX.3 keyframes and Grok Imagine 1.5 all landed in three weeks. Prices read from the live endpoints, and an honest note on which ones changed anything.
Between mid-August and early September 2026, fal.ai gained about forty new video endpoints across nine model families. That sounds like a category-wide leap. It is mostly a repricing and a billing-model change, plus one genuine capability gain that we cannot use. Every price below was read from the live endpoint on 4 September 2026, one endpoint at a time.
The summer round-up covers what came before this — Kling 3.0, HappyHorse 1.1, the first MiniMax H3 and Wan 2.7.
The ones you can actually quote a price for
| Model | Endpoint | Price | Best use |
|---|---|---|---|
| Wan 3.0 Prime i2v | alibaba/wan-3.0-prime/image-to-video | $0.05 / output second | cheapest credible image-to-video; faceless motion |
| Wan 3.0 Prime t2v | alibaba/wan-3.0-prime/text-to-video | $0.05 / output second | same, from a prompt only |
| Kling 3.0 Turbo Standard i2v | fal-ai/kling-video/v3/turbo/standard/image-to-video | $0.112 / output second | 720p with native audio on a budget |
| One-to-All Animation 14B | fal-ai/one-to-all-animation/14b | $0.06 / output second | pose-driven motion transfer, 720p cap |
| FLUX.3 draft (t2v and i2v) | blackforestlabs/flux-3/*/draft | $0.03 / output second | proving a composition before the full render |
| FLUX.3 keyframes / first-last-frame | blackforestlabs/flux-3/keyframes-to-video, .../first-last-frame-to-video | $0.085 / output second | up to 10 anchor frames, 5–20s |
| Kling 3.0 Standard i2v / t2v | fal-ai/kling-video/v3/standard/* | $0.14 / output second | nothing — identical to the Pro rate |
| Kling Native 4K i2v / t2v | fal-ai/kling-video/v3/4k/* | $0.42 / output second | 4K delivery, at 3× the category rate |
| PixVerse V6 | fal-ai/pixverse/v6/* | $0.005 / output second | unverified — an order of magnitude below the class |
| sync-3 Avatar | fal-ai/sync-lipsync/v3/image-to-video | $8.00 / minute ($0.133/s) | one photo plus a voice track → talking head |
| Topaz Upscale Precision | topaz/upscale/video/precision | $0.01 / second | upscaling, with an Iris model that recovers faces |
| Seedance 2.5 (all three) | bytedance/seedance-2.5/* | $0.0214 / 1,000 video tokens (≈ $0.462 / output second at 720p24) | faceless generation only — see below |
The unpriced wave
For the larger half of this release cycle, fal's pricing API publishes no per-output price at all. Asked what these endpoints cost, it returns the platform's GPU-fallback rate — the answer documented for models "that do not have a fixed per-output price", and the same answer it gives for minimax/h3/image-to-video, which was priced per output second back in August and has not changed how it bills since. The likeliest reading is that the pricing table has not caught up with the release wave, not that seventeen endpoints simultaneously switched to charging for GPU time.
Either way we cannot quote them, so each row below carries a worst-case ceiling we would absorb instead of a price:
| Family | Endpoints | Ceiling per job |
|---|---|---|
| MiniMax H3 Max | t2v, i2v, r2v | $0.20–0.25 |
| MiniMax H3 Max Turbo | t2v, i2v | $0.15 |
| Gemini Omni Flash 1.1 | t2v, i2v, r2v, edit | $0.30–0.35 |
| Wan 3.0 (full tier) and Wan Motion | t2v, i2v, r2v, motion | $0.25–0.30 |
| LTX 2.5 | i2v Pro, i2v Fast, t2v Pro | $0.20 / $0.10 / $0.20 |
| Grok Imagine 1.5 | reference-to-video | $0.30 |
An unpriced endpoint is not a detail. Nobody building on top can show a price before the job runs, and the ceiling is a promise we make rather than a number we were quoted — it holds until one paid run measures the real bill, which is why none of those rows is on the storefront. Sixteen of them sit at candidate status waiting for that run; Grok Imagine 1.5 is the exception, rejected on architecture rather than on price, because a scene generator cannot reproduce the Reel you pasted at any rate.
One real capability hides in that table. Gemini Omni Flash 1.1 reference-to-video is the strongest new identity candidate on paper — but its reference clips cap at three seconds each, so it guides a new take rather than copying a Reel, and whether it accepts a real likeness at all is unknown. The matching edit endpoint has no reference-image slot, so it cannot be told whose face goes in.
Seedance 2.5: 30 seconds, still no real faces
The headline is real: up to 50 references and a 30-second single take, and nothing else in the category is close on duration.
It changes nothing for us, and to be precise about why: we have not re-tested the real-face filter on 2.5. It lives in ByteDance's inference layer, upstream of every reseller and every endpoint we could reach, so we treat it as unchanged until the vendor says otherwise — and we are not spending on another bypass attempt to find out. A 30-second take you cannot put a person into is a 30-second stock clip.
The one endpoint the filter cannot block — text-to-video — then loses on price: ≈ $0.462 per output second against $0.085 for the faceless model we already run, so that headline take costs about $13.90. All three rows stay in the registry marked deprecated, so nobody re-litigates this in six weeks.
What actually changed for a visitor
Three things, all in the first table. Wan 3.0 Prime replaces Wan 2.7 on the storefront at half the rate — $0.05 against $0.10, a generation newer; 2.7 stays dispatchable because it is still the only one in the family that accepts a driving audio track. Kling gained a genuinely cheaper 720p tier at $0.112/s with native audio, while the Standard tier at $0.14 costs the same as Pro for a lower tier, so there is no job where it is the pick. And the upscaler is now the precision path, with a face-recovery model behind it.
Everything else is recorded and not on sale. The catalogue holds 104 registry rows; 43 are active and 29 are visible as storefront heads. A row earns a page when its price can be quoted per output second; the rest stay in the registry until a metered run gives them one.
What we route to, and why
None of the above changed our pipeline, and it is worth saying plainly why.
Face replacement stays on HappyHorse. When you paste a Reel of a person talking, moving or living their life, we edit each scene with HappyHorse — up to five reference images (the 1.1 reference-to-video endpoint takes nine, but that is a different tool), genuine 9:16, and the original soundtrack preserved at the model layer. Its list rate is $0.14 per second, though the video-edit endpoint bills input and output seconds, so a 720p scene costs us $0.28 per delivered second. Two endpoints in this wave do edit a source clip and take reference images that say whose face belongs in it, and neither displaces that: Kling O3 4K edit, the same editor whose 1080p version we retired in April for outfit, hair and skin drift at every scene boundary, now at $0.42 per output second — triple the rate for the defect we walked away from — and Bernini-R Reference Edit, which has no published per-output-second price — the pricing API only returns a platform fallback rate for it — so it carries a $0.40 worst-case ceiling in our registry instead of a quoted price.
Dance stays on Kling 3.0 Pro Motion Control at $0.168 per output second, character_orientation set to video — the one endpoint that extracts a motion path from your reference clip and re-performs it. One-to-All Animation 14B is the first quotable per-second challenger since Scail 2 took a blind pair off Kling Pro in August, and — like Scail 2 — it gets a side-by-side before it gets a recommendation.
Faceless scenes run FLUX.3 text-to-video at $0.085 per second, and face-less shots inside a face-replace edit are not generated at all: we trim the original footage and publish it, because there is no identity to replace in a frame with nobody in it.
That is the whole routing table, and this release cycle did not move a row in it.


