The output follows the reference action and its overall timing closely.
Recreate · 4 min read
How to replace a video character without changing the audio
Change who is on screen while the speech, the music and the timing of the original are carried through.
Try Recreate with original audioUpdated 2026-09-04
The short answer
Choose “Keep the original sound” before you generate. The reference clip's audio track is carried into the result, and because the motion is driven by that same reference, the new subject stays in sync with it. No speech is re-synthesized, so there is no voice to match. Burned-in captions generally come through in place — worth a look on the final result.
Step by step
- Start from a clip whose audio you want to keep
This works best when the audio is the part worth keeping — a good take, a licensed music bed, a line delivered with the right timing. A template, or your own clip between 3 and 15 seconds.
- Add the character who should be on screen instead
One image, one subject, front-facing. The audio does not constrain the character at all — a dog can deliver a human voice line, and it stays in sync because both the mouth movement and the audio come from the same reference.
- Select “Keep the original sound”
This is the setting that matters for this workflow. The alternative, “No sound”, returns a silent video — useful when you intend to dub, and exactly what you do not want here.
- Check the sync on the result
Play the result against the original. Beats should land in the same places and captions should sit where they sat before. This is the check worth doing before you publish.
What the reference controls
Framing and camera movement are driven by the reference clip, not re-invented.
The output duration follows the reference duration, and key beats are designed to stay aligned.
With “Keep the original sound” selected, the reference audio is carried into the result.
Burned-in captions and stickers are generally preserved with the reference, but generative output can vary — review the final result before publishing.
UI, product shots and price badges are designed to stay in place. Small text and fine detail are worth checking on the result.
What actually makes a difference
Work backwards: find the clip whose sound you want, then cast into it. Choosing a reference for its visuals and hoping the audio fits is the wrong order.
Burned-in captions are generally preserved along with the reference, so they should still match the audio afterwards. Check them on the result rather than assuming.
This workflow does not synthesize speech. If you need the character to say something different, that is Video Lip Sync, not Recreate.
Sync holds best on short references. A 6-second clip has far fewer frames in which anything can drift than a 15-second one.
Questions
Does the new character lip-sync to the original speech?
The mouth movement is taken from the reference performance, which was already in sync with that audio — so yes, it lands in sync, without any speech being re-generated.
What if I want different words?
Then you want Video Lip Sync, which matches a mouth to replacement audio you provide. Recreate changes who is on screen, not what is said.
Are burned-in captions preserved?
Generally yes — on-screen text, stickers and graphics are designed to stay in place with the reference. Generative output can vary, so review small text on the final result.
Can I mute it instead?
Yes — choose “No sound” and you get a silent video to score or dub yourself.
Ready to try it?
Change who is on screen while the speech, the music and the timing of the original are carried through.
Try Recreate with original audio

