authustle
6 min read

Every AI Video Model That Shipped This Autumn (September 2026)

Seedance 2.5, MiniMax H3 Max, Gemini Omni Flash 1.1, Wan 3.0, LTX 2.5, Kling 4K, PixVerse V6, FLUX.3 keyframes and Grok Imagine 1.5 all landed in three weeks. Prices read from the live endpoints, and an honest note on which ones changed anything.

Between mid-August and early September 2026, fal.ai gained about forty new video endpoints across nine model families. That sounds like a category-wide leap. It is mostly a repricing and a billing-model change, plus one genuine capability gain that we cannot use. Every price below was read from the live endpoint on 4 September 2026, one endpoint at a time.

The summer round-up covers what came before this — Kling 3.0, HappyHorse 1.1, the first MiniMax H3 and Wan 2.7.

The ones you can actually quote a price for

ModelEndpointPriceBest use
Wan 3.0 Prime i2valibaba/wan-3.0-prime/image-to-video$0.05 / output secondcheapest credible image-to-video; faceless motion
Wan 3.0 Prime t2valibaba/wan-3.0-prime/text-to-video$0.05 / output secondsame, from a prompt only
Kling 3.0 Turbo Standard i2vfal-ai/kling-video/v3/turbo/standard/image-to-video$0.112 / output second720p with native audio on a budget
One-to-All Animation 14Bfal-ai/one-to-all-animation/14b$0.06 / output secondpose-driven motion transfer, 720p cap
FLUX.3 draft (t2v and i2v)blackforestlabs/flux-3/*/draft$0.03 / output secondproving a composition before the full render
FLUX.3 keyframes / first-last-frameblackforestlabs/flux-3/keyframes-to-video, .../first-last-frame-to-video$0.085 / output secondup to 10 anchor frames, 5–20s
Kling 3.0 Standard i2v / t2vfal-ai/kling-video/v3/standard/*$0.14 / output secondnothing — identical to the Pro rate
Kling Native 4K i2v / t2vfal-ai/kling-video/v3/4k/*$0.42 / output second4K delivery, at 3× the category rate
PixVerse V6fal-ai/pixverse/v6/*$0.005 / output secondunverified — an order of magnitude below the class
sync-3 Avatarfal-ai/sync-lipsync/v3/image-to-video$8.00 / minute ($0.133/s)one photo plus a voice track → talking head
Topaz Upscale Precisiontopaz/upscale/video/precision$0.01 / secondupscaling, with an Iris model that recovers faces
Seedance 2.5 (all three)bytedance/seedance-2.5/*$0.0214 / 1,000 video tokens (≈ $0.462 / output second at 720p24)faceless generation only — see below

The unpriced wave

For the larger half of this release cycle, fal's pricing API publishes no per-output price at all. Asked what these endpoints cost, it returns the platform's GPU-fallback rate — the answer documented for models "that do not have a fixed per-output price", and the same answer it gives for minimax/h3/image-to-video, which was priced per output second back in August and has not changed how it bills since. The likeliest reading is that the pricing table has not caught up with the release wave, not that seventeen endpoints simultaneously switched to charging for GPU time.

Either way we cannot quote them, so each row below carries a worst-case ceiling we would absorb instead of a price:

FamilyEndpointsCeiling per job
MiniMax H3 Maxt2v, i2v, r2v$0.20–0.25
MiniMax H3 Max Turbot2v, i2v$0.15
Gemini Omni Flash 1.1t2v, i2v, r2v, edit$0.30–0.35
Wan 3.0 (full tier) and Wan Motiont2v, i2v, r2v, motion$0.25–0.30
LTX 2.5i2v Pro, i2v Fast, t2v Pro$0.20 / $0.10 / $0.20
Grok Imagine 1.5reference-to-video$0.30

An unpriced endpoint is not a detail. Nobody building on top can show a price before the job runs, and the ceiling is a promise we make rather than a number we were quoted — it holds until one paid run measures the real bill, which is why none of those rows is on the storefront. Sixteen of them sit at candidate status waiting for that run; Grok Imagine 1.5 is the exception, rejected on architecture rather than on price, because a scene generator cannot reproduce the Reel you pasted at any rate.

One real capability hides in that table. Gemini Omni Flash 1.1 reference-to-video is the strongest new identity candidate on paper — but its reference clips cap at three seconds each, so it guides a new take rather than copying a Reel, and whether it accepts a real likeness at all is unknown. The matching edit endpoint has no reference-image slot, so it cannot be told whose face goes in.

Seedance 2.5: 30 seconds, still no real faces

The headline is real: up to 50 references and a 30-second single take, and nothing else in the category is close on duration.

It changes nothing for us, and to be precise about why: we have not re-tested the real-face filter on 2.5. It lives in ByteDance's inference layer, upstream of every reseller and every endpoint we could reach, so we treat it as unchanged until the vendor says otherwise — and we are not spending on another bypass attempt to find out. A 30-second take you cannot put a person into is a 30-second stock clip.

The one endpoint the filter cannot block — text-to-video — then loses on price: ≈ $0.462 per output second against $0.085 for the faceless model we already run, so that headline take costs about $13.90. All three rows stay in the registry marked deprecated, so nobody re-litigates this in six weeks.

What actually changed for a visitor

Three things, all in the first table. Wan 3.0 Prime replaces Wan 2.7 on the storefront at half the rate — $0.05 against $0.10, a generation newer; 2.7 stays dispatchable because it is still the only one in the family that accepts a driving audio track. Kling gained a genuinely cheaper 720p tier at $0.112/s with native audio, while the Standard tier at $0.14 costs the same as Pro for a lower tier, so there is no job where it is the pick. And the upscaler is now the precision path, with a face-recovery model behind it.

Everything else is recorded and not on sale. The catalogue holds 104 registry rows; 43 are active and 29 are visible as storefront heads. A row earns a page when its price can be quoted per output second; the rest stay in the registry until a metered run gives them one.

What we route to, and why

None of the above changed our pipeline, and it is worth saying plainly why.

Face replacement stays on HappyHorse. When you paste a Reel of a person talking, moving or living their life, we edit each scene with HappyHorse — up to five reference images (the 1.1 reference-to-video endpoint takes nine, but that is a different tool), genuine 9:16, and the original soundtrack preserved at the model layer. Its list rate is $0.14 per second, though the video-edit endpoint bills input and output seconds, so a 720p scene costs us $0.28 per delivered second. Two endpoints in this wave do edit a source clip and take reference images that say whose face belongs in it, and neither displaces that: Kling O3 4K edit, the same editor whose 1080p version we retired in April for outfit, hair and skin drift at every scene boundary, now at $0.42 per output second — triple the rate for the defect we walked away from — and Bernini-R Reference Edit, which has no published per-output-second price — the pricing API only returns a platform fallback rate for it — so it carries a $0.40 worst-case ceiling in our registry instead of a quoted price.

Dance stays on Kling 3.0 Pro Motion Control at $0.168 per output second, character_orientation set to video — the one endpoint that extracts a motion path from your reference clip and re-performs it. One-to-All Animation 14B is the first quotable per-second challenger since Scail 2 took a blind pair off Kling Pro in August, and — like Scail 2 — it gets a side-by-side before it gets a recommendation.

Faceless scenes run FLUX.3 text-to-video at $0.085 per second, and face-less shots inside a face-replace edit are not generated at all: we trim the original footage and publish it, because there is no identity to replace in a frame with nobody in it.

That is the whole routing table, and this release cycle did not move a row in it.

seedance 2.5minimax h3 maxgemini omni flashwan 3.0ltx 2.5kling 4kmodel releases

Keep reading