Recreate · 4 min read

How to replace a video character without changing the audio

Change who is on screen while the speech, the music and the timing of the original are carried through.

Try Recreate with original audio

Updated 2026-09-04

ReferenceRecreated
Same audio, same HOT TAKE caption, same delivery. Only the host changed.

The short answer

Choose “Keep the original sound” before you generate. The reference clip's audio track is carried into the result, and because the motion is driven by that same reference, the new subject stays in sync with it. No speech is re-synthesized, so there is no voice to match. Burned-in captions generally come through in place — worth a look on the final result.

Step by step

  1. Start from a clip whose audio you want to keep

    This works best when the audio is the part worth keeping — a good take, a licensed music bed, a line delivered with the right timing. A template, or your own clip between 3 and 15 seconds.

  2. Add the character who should be on screen instead

    One image, one subject, front-facing. The audio does not constrain the character at all — a dog can deliver a human voice line, and it stays in sync because both the mouth movement and the audio come from the same reference.

  3. Select “Keep the original sound”

    This is the setting that matters for this workflow. The alternative, “No sound”, returns a silent video — useful when you intend to dub, and exactly what you do not want here.

  4. Check the sync on the result

    Play the result against the original. Beats should land in the same places and captions should sit where they sat before. This is the check worth doing before you publish.

Try Recreate with original audio

What the reference controls

Motion

The output follows the reference action and its overall timing closely.

Camera

Framing and camera movement are driven by the reference clip, not re-invented.

Timing

The output duration follows the reference duration, and key beats are designed to stay aligned.

Audio

With “Keep the original sound” selected, the reference audio is carried into the result.

Captions

Burned-in captions and stickers are generally preserved with the reference, but generative output can vary — review the final result before publishing.

Graphics

UI, product shots and price badges are designed to stay in place. Small text and fine detail are worth checking on the result.

What actually makes a difference

Let the audio pick the reference

Work backwards: find the clip whose sound you want, then cast into it. Choosing a reference for its visuals and hoping the audio fits is the wrong order.

Captions are part of the audio story

Burned-in captions are generally preserved along with the reference, so they should still match the audio afterwards. Check them on the result rather than assuming.

Do not expect a new voice

This workflow does not synthesize speech. If you need the character to say something different, that is Video Lip Sync, not Recreate.

Keep it short to keep it tight

Sync holds best on short references. A 6-second clip has far fewer frames in which anything can drift than a 15-second one.

Questions

Does the new character lip-sync to the original speech?

The mouth movement is taken from the reference performance, which was already in sync with that audio — so yes, it lands in sync, without any speech being re-generated.

What if I want different words?

Then you want Video Lip Sync, which matches a mouth to replacement audio you provide. Recreate changes who is on screen, not what is said.

Are burned-in captions preserved?

Generally yes — on-screen text, stickers and graphics are designed to stay in place with the reference. Generative output can vary, so review small text on the final result.

Can I mute it instead?

Yes — choose “No sound” and you get a silent video to score or dub yourself.

Ready to try it?

Change who is on screen while the speech, the music and the timing of the original are carried through.

Try Recreate with original audio