← back to blog
September 12, 2026·11 min read

A clip with my own face and voice: the whole walkthrough

The topic looks solved from the outside: upload a photo, write a prompt, get a video of yourself. In practice there are about five places between «uploaded» and «worked» where the result quietly breaks. And it does not break with an error, it just comes out bad, after you have already paid for the run.

Below is the whole chain: face, voice, swapping a face into footage you shot, generating from scratch, translation and lip-sync. With the price of each step and the places I tripped over myself.

What actually goes into a clip like this

Let me split this up front, because it causes confusion later. «A clip with my face» is two completely different jobs, handled by different models:

  • Face swap. I have footage I shot and I replace the person in it. Background, motion and audio stay mine. No prompt needed, because whatever is on screen I already filmed.
  • Generation from scratch. There is no source clip, the model builds the shot from a description. The prompt is mandatory: it sets the action, the background and the camera.

The first gives you realism, because it is real footage. The second gives you freedom, but the picture comes out cleaner and more synthetic. Choosing between them is choosing «my background or my script».

Step 1. The face

It starts with a set of reference photos. First non-obvious thing: the models need very few of them. Kling takes four references maximum, and that is a hard cap, not a suggestion. So every «upload 15 photos» instruction you have read applies to training your own model, not to ordinary generation.

I ask myself for 5-8 photos. Not because the model will consume them, but so there is something to pick the best shot from for a given scene.

Reference photo block in the persona section
The «How it looks» block. Angle tiles, requirements on the right, the right-to-use checkbox at the bottom. Generation will not start without it.

What actually moves the needle, in order:

  1. How large the face is in frame. Swap resolution caps at 720p. If the person is far away, too few pixels land on the face and the model cannot carry it. A mid shot works noticeably better than a wide one.
  2. Matching light between the photo and the target video. More on this below, it is the most interesting part.
  3. Angle. The whole face has to be visible. A hard head turn or partly hidden features hurt the result.
  4. Sharpness. Worse source, worse output. No surprises here.

And one item it took me a while to get to: include one full-length or waist-up shot. Reference sets are usually portraits, and if the scene shows more than a face, the model invents the body and the clothing. Sometimes decently, sometimes not.

If you do not have a face to use

Half the people doing this work with a character who does not exist. There is a second entry point for them: describe the appearance in text and the model renders four shots of the same non-existent person.

There was a technical trap here that I hit while building it. The obvious approach is four requests with the same description and different angles. That gives you four different people, because the model has no parameter that pins the appearance across requests. The only thing that works: generate the first shot, then derive the other three from it, passing it back in as the reference. Then it is one person.

Which is also why there is an «add angles» button: it takes your real photos and fills in the missing views. If you have one good portrait, that is already enough to start.

Creating a model from a description
The «Create a model» tab. A description of the appearance, and out come front, three-quarter, profile and smile of one face.

Step 2. The voice

Three routes, and they differ in uniqueness rather than quality:

RouteWhat it needsUniqueness
Clone your own60-90 sec of clean speechAbsolute, it is your voice
Design a voiceA written descriptionHigh, nobody has that timbre
LibraryNothingNone, it sounds the same for everyone who picks it

A minute to a minute and a half is enough for the clone. Read any text in an even voice, no music, no room echo. After that any text is read in your voice.

The designed voice had a detail that surprised me. The model has no numeric settings. It takes a description in words. All those «timbre, tempo, warmth» sliders you see in interfaces are a layer on top: they compose text, and the text is what goes to the model.

So we show that description as its own field and let you rewrite it by hand. It gives noticeably more control than three sliders. You cannot dial in «husky male voice, lazy delivery, like telling a story to a friend» with sliders at all.

Voice design tab with the description field
The sliders compose the description underneath them. The field is editable, and the text is what decides.

Step 3. Face swap. Here is the main trap

Clip on the timeline, pick the persona, hit the button. And this is where it most often comes out as garbage, with no obvious reason why.

The model reads identity from one frame. Not from the whole video. One. And by default that is the first frame of the clip.

Now think about what is in your first frame. A fade-in. Titles. A wide shot of you walking into a room. The back of your head. In any of those the model has nothing to read, and the swap either does nothing or drifts across the whole clip. And it looks like «the model is bad».

So we take the frame at the playhead. Literally: park the scrubber on a moment where your face is large and clearly visible, and run it from there. That one thing matters more than every other setting combined.

Face swap panel in the editor
The swap panel. The «read the face at the playhead» toggle and the resolution picker.

The trick that lifts quality the most

Guides for similar services contain this advice: build your reference photo from the first frame of the target video so the light and the angle match. Then follows a twenty-line prompt you are supposed to copy into a separate image tool, assemble a portrait there and come back.

The advice is completely right. Matching light is the second biggest factor after how large the face is. It is just inconvenient enough by hand that almost nobody does it.

But the clip is already on the timeline. So the frame can be grabbed for you, handed to the model together with the persona photos, and you get the persona in that frame's light and angle. And the swap runs with that.

For us that is the «match the reference to this shot» toggle, on by default. It adds $0.08, about a third of the swap itself. Worth turning off only to save on a rough draft.

Step 4. Generating from scratch

There is no source here, so the prompt becomes the only thing that defines the shot. The face is passed in as a reference you have to mention in the text, otherwise the model does not know who is on screen.

Compare two prompts. Bad:

a person dances in the rain and sings

Working:

@Image1 dances in the rain on a night street,
mid shot, black jacket and jeans, wet hair,
neon reflections in the puddles, camera slowly orbits

The difference is that the second one says who is on screen, what they are wearing, how close the shot is and how the camera moves. Anything you leave out, the model invents for you.

And about singing, because this is the most asked question

Write «sings a song» and the mouth will move and it will look like singing. But the voice will be invented by the model, not yours, and there will be no particular song.

This is not one service falling short. Every video model that generates audio invents the voice itself. None of them accept your clone. To get your actual voice you need a three-step chain: generate the video, voice it with the clone, align the lips with lip-sync. And even then a limit remains: speech synthesis does not sing, it reads. Proper singing in your own voice does not exist anywhere right now, and it is worth knowing that before you go looking.

Step 5. Translation and lip-sync

The most underrated case: you have your own clip and you want it in English in your own voice. For that you do not even need a voice clone: the dubbing model lifts the timbre straight off the track. Upload, pick the language, get the same voice in another one.

Now the price, and there is some arithmetic here that matters.

Dubbing is billed per started minute, rounded up. Which means a 19-second clip costs the same as a 59-second one: one full minute. I first priced it pro-rata by seconds and was off by a factor of three. If you plan to translate a lot of short clips, it is cheaper to stitch them together and translate in one batch.

Lip-sync is a separate operation and it is five times more expensive than the translation itself. So we keep it off by default, with that stated plainly. Without it the lips will not match, and for a talking head that is fine, which is how most dubbing services work. Turn it on when the face is close up and the mismatch is obvious.

Voice panel, translation tab
The translation tab. Languages, voice choice, and lip-sync as its own checkbox with the price.

What this actually costs

I collected these myself against the model pages as of 12 September 2026. They change, so verify, but the order of magnitude is this:

OperationPriceHow it is counted
Face swap, 720p$0.20Per clip up to 5 sec, double above
Matching the reference to a frame$0.08Per image
Generation with a face$0.14 / secWith audio. $0.112 without
Animating a photo~$0.30 for 5 secPer second
Translation$0.60 / minPer started minute, rounded up
Lip-sync$3 / min$5 for the pro variant
Text to speech$0.10 / 1000 charsA 30-sec clip is ~$0.05

For calibration: one of my 5-second clips with audio came to $0.28. That matched the formula, so the table can be trusted.

Note the spread. Text to speech costs pennies, a face swap costs cents, and lip-sync on a minute of video is already comparable to lunch. Budget around lip-sync, everything else is noise next to it.

Short checklist

  • 5-8 photos, one of them full-length.
  • Even light, no glasses or hats, the whole face visible.
  • Before a swap, park the playhead on a frame with a large face.
  • Leave the reference matching on.
  • 720p is the ceiling, nobody has HD.
  • In the prompt, describe clothing, shot size and camera.
  • Translation needs no clone.
  • Turn lip-sync on only for close-ups.
  • Translate short clips in a batch, not one by one.

What nobody can do yet

So you do not go looking where there is nothing:

  • Singing in your own voice. The video model invents a voice, speech synthesis does not sing.
  • HD face swap. The ceiling is 720p, and that is the technology, not the service.
  • Long clips in one pass. The longer the source, the less stable the result. If the video has a cut or a wardrobe change, split it, run each part and stitch back.

That last point, incidentally, is the main reason this is easier in an editor with a timeline than in a bot. Cutting by scene, running each part and stitching back is simply not something you can do there.