Kling AI Avatar
A speaking or singing portrait from one photo and a finished audio track.
When to choose it
Kling AI Avatar is for scripts you have already recorded and now need a face on screen. The model animates the portrait to your audio, so approve the wording, voice and pauses first. Test a short excerpt with a clearly visible face before processing the full recording.
- Uses recorded speech or vocals rather than inventing a new line from a prompt.
- Build explainers, greetings and singing portraits without filming a presenter.
Modes and settings
- One photo and one audio file are required.
- Standard outputs 720p; Pro outputs 1080p. Duration follows the audio, up to 300 seconds.
Price in coins
| Version | Quality / size | Input | Coins / second |
|---|---|---|---|
| Standard | 720p | Portrait + audio | ≈ 5.6 |
| Pro | 1080p | Portrait + audio | ≈ 10.92 |
Billing is per second, not per fixed-length clip. The rate is multiplied by the billable duration; the complete request total is rounded up to a whole coin once.
Duration follows the uploaded audio track; the rate depends on the selected quality.
≈ marks a fractional rate rounded to two decimal places for display; billing uses full precision.
What you can make
- A presenter for a short tutorial.
- A singing character for a prepared vocal excerpt.
How to direct the result
- Prepare a clean single-speaker recording with natural pauses.
- Use a front-facing or near-front portrait with nothing covering the mouth.
Preparation example: a presenter greeting
Photo: a chest-up portrait facing the camera, with mouth and chin clearly visible. Audio: a brief single-speaker greeting without music. Start with Standard and a ten-second excerpt; check lip movement, pauses and head position.
What to keep in mind
- Typed text cannot replace the audio file.
- A strong profile, hidden face or multiple voices can weaken lip synchronization.
Sources
- Prices and availability — Mixer AI catalog, updated 2026-09-18.
- Official maker site: Kuaishou
Model facts verified: 2026-09-20.
Guides and comparisons using this model
FAQ
Can I just type a script?
No. Record or synthesize the speech first, then upload the audio and portrait.
Can it animate singing?
Yes. Use a track with clear vocals. Loud accompaniment makes the task harder.
What determines video length?
The audio duration, with a maximum of 300 seconds per result.
Which portrait should I choose?
One clearly visible face, near-frontal framing and even light.
Why test a short excerpt?
It checks the particular face-and-voice combination before processing the entire recording.
Can I make two speakers?
Prepare one speaker for this tool. Process different characters separately and combine their lines in editing.
What changes the price?
The coin rate per second depends on quality, mode and, for some models, sound. Multiply it by the billable duration. Where input video is charged, its duration also counts. The complete request total is rounded once.
Do I need a separate model subscription?
No. In Mixer AI, you top up one coin balance and use it for the tasks you need. You do not need a separate subscription to this model.