The topic looks solved from the outside: upload a photo, write a prompt, get a video of yourself. In practice there are about five places between «uploaded» and «worked» where the result quietly breaks. And it does not break with an error, it just comes out bad, after you have already paid for the run.
Below is the whole chain: face, voice, swapping a face into footage you shot, generating from scratch, translation and lip-sync. With the price of each step and the places I tripped over myself.
What actually goes into a clip like this
Let me split this up front, because it causes confusion later. «A clip with my face» is two completely different jobs, handled by different models:
- Face swap. I have footage I shot and I replace the person in it. Background, motion and audio stay mine. No prompt needed, because whatever is on screen I already filmed.
- Generation from scratch. There is no source clip, the model builds the shot from a description. The prompt is mandatory: it sets the action, the background and the camera.
The first gives you realism, because it is real footage. The second gives you freedom, but the picture comes out cleaner and more synthetic. Choosing between them is choosing «my background or my script».
Step 1. The face
It starts with a set of reference photos. First non-obvious thing: the models need very few of them. Kling takes four references maximum, and that is a hard cap, not a suggestion. So every «upload 15 photos» instruction you have read applies to training your own model, not to ordinary generation.
I ask myself for 5-8 photos. Not because the model will consume them, but so there is something to pick the best shot from for a given scene.

What actually moves the needle, in order:
- How large the face is in frame. Swap resolution caps at 720p. If the person is far away, too few pixels land on the face and the model cannot carry it. A mid shot works noticeably better than a wide one.
- Matching light between the photo and the target video. More on this below, it is the most interesting part.
- Angle. The whole face has to be visible. A hard head turn or partly hidden features hurt the result.
- Sharpness. Worse source, worse output. No surprises here.
And one item it took me a while to get to: include one full-length or waist-up shot. Reference sets are usually portraits, and if the scene shows more than a face, the model invents the body and the clothing. Sometimes decently, sometimes not.
If you do not have a face to use
Half the people doing this work with a character who does not exist. There is a second entry point for them: describe the appearance in text and the model renders four shots of the same non-existent person.
There was a technical trap here that I hit while building it. The obvious approach is four requests with the same description and different angles. That gives you four different people, because the model has no parameter that pins the appearance across requests. The only thing that works: generate the first shot, then derive the other three from it, passing it back in as the reference. Then it is one person.
Which is also why there is an «add angles» button: it takes your real photos and fills in the missing views. If you have one good portrait, that is already enough to start.

Step 2. The voice
Three routes, and they differ in uniqueness rather than quality:
| Route | What it needs | Uniqueness |
|---|---|---|
| Clone your own | 60-90 sec of clean speech | Absolute, it is your voice |
| Design a voice | A written description | High, nobody has that timbre |
| Library | Nothing | None, it sounds the same for everyone who picks it |
A minute to a minute and a half is enough for the clone. Read any text in an even voice, no music, no room echo. After that any text is read in your voice.
The designed voice had a detail that surprised me. The model has no numeric settings. It takes a description in words. All those «timbre, tempo, warmth» sliders you see in interfaces are a layer on top: they compose text, and the text is what goes to the model.
So we show that description as its own field and let you rewrite it by hand. It gives noticeably more control than three sliders. You cannot dial in «husky male voice, lazy delivery, like telling a story to a friend» with sliders at all.

Step 3. Face swap. Here is the main trap
Clip on the timeline, pick the persona, hit the button. And this is where it most often comes out as garbage, with no obvious reason why.
The model reads identity from one frame. Not from the whole video. One. And by default that is the first frame of the clip.
Now think about what is in your first frame. A fade-in. Titles. A wide shot of you walking into a room. The back of your head. In any of those the model has nothing to read, and the swap either does nothing or drifts across the whole clip. And it looks like «the model is bad».
So we take the frame at the playhead. Literally: park the scrubber on a moment where your face is large and clearly visible, and run it from there. That one thing matters more than every other setting combined.

The trick that lifts quality the most
Guides for similar services contain this advice: build your reference photo from the first frame of the target video so the light and the angle match. Then follows a twenty-line prompt you are supposed to copy into a separate image tool, assemble a portrait there and come back.
The advice is completely right. Matching light is the second biggest factor after how large the face is. It is just inconvenient enough by hand that almost nobody does it.
But the clip is already on the timeline. So the frame can be grabbed for you, handed to the model together with the persona photos, and you get the persona in that frame's light and angle. And the swap runs with that.
For us that is the «match the reference to this shot» toggle, on by default. It adds $0.08, about a third of the swap itself. Worth turning off only to save on a rough draft.
Step 4. Generating from scratch
There is no source here, so the prompt becomes the only thing that defines the shot. The face is passed in as a reference you have to mention in the text, otherwise the model does not know who is on screen.
Compare two prompts. Bad:
a person dances in the rain and sings
Working:
@Image1 dances in the rain on a night street, mid shot, black jacket and jeans, wet hair, neon reflections in the puddles, camera slowly orbits
The difference is that the second one says who is on screen, what they are wearing, how close the shot is and how the camera moves. Anything you leave out, the model invents for you.
And about singing, because this is the most asked question
Write «sings a song» and the mouth will move and it will look like singing. But the voice will be invented by the model, not yours, and there will be no particular song.
This is not one service falling short. Every video model that generates audio invents the voice itself. None of them accept your clone. To get your actual voice you need a three-step chain: generate the video, voice it with the clone, align the lips with lip-sync. And even then a limit remains: speech synthesis does not sing, it reads. Proper singing in your own voice does not exist anywhere right now, and it is worth knowing that before you go looking.
Step 5. Translation and lip-sync
The most underrated case: you have your own clip and you want it in English in your own voice. For that you do not even need a voice clone: the dubbing model lifts the timbre straight off the track. Upload, pick the language, get the same voice in another one.
Now the price, and there is some arithmetic here that matters.
Dubbing is billed per started minute, rounded up. Which means a 19-second clip costs the same as a 59-second one: one full minute. I first priced it pro-rata by seconds and was off by a factor of three. If you plan to translate a lot of short clips, it is cheaper to stitch them together and translate in one batch.
Lip-sync is a separate operation and it is five times more expensive than the translation itself. So we keep it off by default, with that stated plainly. Without it the lips will not match, and for a talking head that is fine, which is how most dubbing services work. Turn it on when the face is close up and the mismatch is obvious.

What this actually costs
I collected these myself against the model pages as of 12 September 2026. They change, so verify, but the order of magnitude is this:
| Operation | Price | How it is counted |
|---|---|---|
| Face swap, 720p | $0.20 | Per clip up to 5 sec, double above |
| Matching the reference to a frame | $0.08 | Per image |
| Generation with a face | $0.14 / sec | With audio. $0.112 without |
| Animating a photo | ~$0.30 for 5 sec | Per second |
| Translation | $0.60 / min | Per started minute, rounded up |
| Lip-sync | $3 / min | $5 for the pro variant |
| Text to speech | $0.10 / 1000 chars | A 30-sec clip is ~$0.05 |
For calibration: one of my 5-second clips with audio came to $0.28. That matched the formula, so the table can be trusted.
Note the spread. Text to speech costs pennies, a face swap costs cents, and lip-sync on a minute of video is already comparable to lunch. Budget around lip-sync, everything else is noise next to it.
Short checklist
- 5-8 photos, one of them full-length.
- Even light, no glasses or hats, the whole face visible.
- Before a swap, park the playhead on a frame with a large face.
- Leave the reference matching on.
- 720p is the ceiling, nobody has HD.
- In the prompt, describe clothing, shot size and camera.
- Translation needs no clone.
- Turn lip-sync on only for close-ups.
- Translate short clips in a batch, not one by one.
What nobody can do yet
So you do not go looking where there is nothing:
- Singing in your own voice. The video model invents a voice, speech synthesis does not sing.
- HD face swap. The ceiling is 720p, and that is the technology, not the service.
- Long clips in one pass. The longer the source, the less stable the result. If the video has a cut or a wardrobe change, split it, run each part and stitch back.
That last point, incidentally, is the main reason this is easier in an editor with a timeline than in a bot. Cutting by scene, running each part and stitching back is simply not something you can do there.