authustle
Step by step

How to make a photo talk

A talking photo is the same tool as animating one, with the optional audio field filled in. You upload a still, add a voice track, and the lips follow it. The interesting decision is not how — it is which model: the tool's default handles a short line well, while the catalogue's dedicated talking-avatar models are built for the case where the photo has to keep speaking.

Open the Studio

No credit card to sign up · 5 steps

Same form, one more field

  1. 1

    Open Studio and pick Animate a photo

    There is no separate talking-photo tool. The audio field is what turns this one into it, and it only appears when the model behind the tool accepts an audio input.

    studio/animate-photo

  2. 2

    Upload the portrait under Photo

    Field one, labelled Photo, hinted "A clear, front-facing photo works best". For a talking result this matters more than anywhere else: the mouth has to be visible and unobstructed, or there is nothing to animate.

    studio/animate-photo#photo

  3. 3

    Add the voice track under Audio (optional)

    The second field is labelled "Audio (optional)" and hinted "Add a voice track and the lips will follow it". Bring your own recording — this tool syncs a photo to audio, it does not synthesise a voice for you.

    studio/animate-photo#audio

  4. 4

    Leave the prompt empty, set the length, press Generate

    "What should happen? (optional)" is best left blank here — the audio already dictates the performance, and a competing instruction usually costs you the lip-sync. Set the length to cover the clip and press Generate; the credit cost is printed beside the button.

    studio/animate-photo#duration

  5. 5

    Collect the clip from your Library

    The finished file lands in your Library. Runs are processed in the background, so you can close the tab — the finished file turns up in your Library. We do not quote a fixed wall-clock time: it moves with the length of the run and with how busy the queue is.

    library

Tool default, or a dedicated avatar model

HappyHorse 1.1 Image to Video

Product default

A short line, a single shot, a photo you already have. This is what the Animate a photo tool runs on, so it is the zero-decision path.

The trade-off: It is an image-to-video model that accepts audio, not a talking-avatar model. Over a long stretch of speech the alternatives below hold a performance better.

Specs and credit price →

OmniHuman 1.5

From the model spec

The photo has to keep talking — a full read, a piece to camera, an explainer. This is the catalogue's recommended pick in the talking-avatar category, and it treats the voice track as the job rather than as an extra input.

The trade-off: A dedicated avatar model is a different price and a different run; check its catalogue page before you commit a long clip to it.

Specs and credit price →

Kling AI Avatar v2 Pro

From the model spec

A second opinion in the same category when a result comes back stiff — different families fail differently on the same portrait.

The trade-off: Another talking-avatar model with its own ceiling and its own price, both on its catalogue page.

Specs and credit price →

These picks come from our own side-by-side runs and from what the tools actually default to. The specs and credit prices live on each model's own page and are generated from the catalogue, so they cannot drift out of date here. We do not quote anyone else's review, benchmark or price.

What it costs to follow this guide

Your first video from $0.99

A one-time purchase — 35 credits land in your balance immediately. No subscription, nothing renews, and if a generation fails on our side the credits come back automatically.

What this does not do

  • It does not generate the voice. You supply the audio; the lips follow it.
  • It does not translate. Re-voicing into another language is a different tool, and it starts from a video rather than a still.
  • It does not animate a mouth it cannot see. Masks, hands over the face and extreme profiles have nothing to sync.
  • It does not post anywhere. We never ask for your Instagram, TikTok or YouTube password, never read your DMs and never publish anything for you. You download the file and post it yourself.
FAQ

Common questions

Where does the voice come from?
From you. This tool syncs a photo to an audio file you upload; it does not synthesise speech, and there is no voice picker in the form.
Why is the audio field sometimes missing?
Because the field is drawn from the model's declared inputs rather than hard-coded. If the model behind the tool does not accept audio, the form does not offer a box that would be ignored.
Should I write a prompt as well?
Usually not. The audio already determines the performance, and a competing instruction in "What should happen?" tends to cost you the lip-sync.
What does it cost to try?
Your first video from $0.99. A one-time purchase — 35 credits land in your balance immediately. No subscription, nothing renews, and if a generation fails on our side the credits come back automatically.

Keep reading

Make your photo talk

Bring a still and a voice track. The form prints the credit cost of the run before you press Generate.

Open the Studio

No credit card to sign up