Video Lip Sync · 4 min read

AI lip sync video: best inputs, limits and examples

What kind of footage works, what does not, and exactly where the current limits are — before you spend credits finding out.

Try Lip Sync

Updated 2026-09-06

ReferenceRecreated
The same take, before and after the spoken line was replaced.

The short answer

The best input is a single speaker, facing the camera, mouth visible and unobstructed for the whole clip. Lip Sync replaces the visible mouth movement only — it does not translate, dub, handle two speakers, or make a still photo sing. Free previews allow up to 8 seconds of speech, and an 8-second preview costs 240 credits. Speech can be typed, uploaded or recorded; all three produce the same kind of result.

Step by step

  1. What works

    One person on camera, front-facing or close to it, mouth visible from the first frame to the last. Talking-head footage — presenters, product explainers, pieces to camera — is the case this is built for.

  2. What does not work yet

    Two people talking in the same shot, a face that leaves frame mid-sentence, a mouth covered by a hand or microphone, or a still photo. A still photo is AI Pet Podcast or Recreate territory, not Lip Sync.

  3. The hard limits

    Free previews cap the speech at 8 seconds. Typed lines are checked against that cap before anything is generated, so an over-long script is refused rather than half-charged.

  4. What it costs

    Lip Sync is billed by the second of speech. An 8-second preview is 240 credits. Typing the line adds a small synthesis charge on top; the exact number is shown on the generate button before you commit.

Try Lip Sync

What actually makes a difference

Test with a short line first

A two-second line tells you whether the footage suits the model, at a fraction of the credits of a full take.

Prefer even, front-lit faces

Heavy side lighting and deep shadow around the mouth give the model less to work with than a flatly lit face.

Say what you mean it to say

The output is a person appearing to say words they did not say. Only do it with footage and a voice you have the right to use, and never to impersonate someone.

Keep the clip tight

Shorter speech costs fewer credits and gives the model less opportunity to drift. Trim the clip to the sentence you actually want to change.

Questions

Can it sync two speakers in the same clip?

No. One visible speaker per clip. A second face on camera is not tracked.

Can it make a photo sing?

No. Lip Sync needs video. A single photo goes through AI Pet Podcast for pets, or Recreate to cast a character into an existing performance.

How long can the speech be?

Free previews allow up to 8 seconds. Typed lines are measured before generation, so you are told the line is too long instead of being charged for a failure.

Does typing the line cost more than uploading audio?

Slightly. Typing adds a small speech-synthesis charge on top of the lip sync itself. The generate button shows the exact total before you commit.

Ready to try it?

What kind of footage works, what does not, and exactly where the current limits are — before you spend credits finding out.

Try Lip Sync