authustle
2 min read

Talking Photo Apps Compared 2026: Four Models, $0.06 to $0.16 a Second

Animating a still into a speaking clip is one API call, but the four models that do it are driven differently and priced 3x apart. Which to pick for a script, an audio file, or a budget.

"Talking photo" covers two different inputs — a script the model must voice, or an audio file you already have — and the right model depends entirely on which you hold. Here are the four worth using, with prices checked against the live endpoints on 3 August 2026.

The four

Model$/second10s clipDriven by
Kling AI Avatar v2 Standard$0.0562$0.56script or audio
Kling AI Avatar v2 Pro$0.115$1.15script or audio
HappyHorse 1.1 Image to Video$0.14$1.40script, native audio
OmniHuman 1.5$0.16$1.60audio file

That is a 2.8× spread for the same nominal job, which makes the choice worth two minutes of thought.

If you have a script

Use HappyHorse 1.1 Image to Video. It generates the audio natively and lip-syncs multilingually in the same call, outputs 1080p, and covers nine aspect ratios including real 9:16 and 4:5. One request, finished vertical clip.

The default aspect ratio is 16:9. Set it, every time.

If you already have the audio

Use OmniHuman 1.5. It is audio-driven: expressions and body movement follow the waveform rather than a text prompt. If you recorded a voice-over, or you are using a cloned voice from elsewhere, this is the fit — and you are not paying a model to regenerate audio you already own.

If the budget is binding

Kling AI Avatar v2 Standard at $0.0562 per second is less than half the price of anything else here, and it is a genuine talking-avatar model rather than a cut-down demo. A ten-second clip is 56 cents.

The Pro tier at $0.115 buys visibly better facial fidelity. The rule we use: Standard for anything watched small or produced at volume, Pro when the face fills the frame.

What decides quality, and it is not the model

Motion restraint. "Speaks calmly to camera, slight head movement" holds a face where "gestures energetically" does not. This is the single highest-leverage change available and it is free.

Clip length. Five to eight seconds is where identity is reliably stable. For longer pieces, generate several short shots and cut — a cut is invisible, drift is not.

The source photo. Even frontal light, whole head in frame, no heavy side shadow, no sunglasses. The model is inferring a three-dimensional head from your still; ambiguity in the input becomes wobble in the output.

What none of them do

Put your animated head into an existing video. That is a different family — see Wan 2.7 Edit Video at $0.10/second for editing a clip you already have, or Kling 3.0 Pro Motion Control at $0.168/second for copying a clip's movement onto your character.

Asking a talking-photo model to do that produces exactly the disappointing result you would expect, at full price.

talking photolip synccomparisonomnihumanavatars

Keep reading