Mixer AIMixer AI
Start generating
Mixer AI › Knowledge Base › Video › Kling AI Avatar

Kling AI Avatar

A speaking or singing portrait from one photo and a finished audio track.

When to choose it

Kling AI Avatar is for scripts you have already recorded and now need a face on screen. The model animates the portrait to your audio, so approve the wording, voice and pauses first. Test a short excerpt with a clearly visible face before processing the full recording.

Modes and settings

Price in coins

VersionQuality / sizeInputCoins / second
Standard720pPortrait + audio≈ 5.6
Pro1080pPortrait + audio≈ 10.92

Billing is per second, not per fixed-length clip. The rate is multiplied by the billable duration; the complete request total is rounded up to a whole coin once.

Duration follows the uploaded audio track; the rate depends on the selected quality.

≈ marks a fractional rate rounded to two decimal places for display; billing uses full precision.

What you can make

How to direct the result

  1. Prepare a clean single-speaker recording with natural pauses.
  2. Use a front-facing or near-front portrait with nothing covering the mouth.

Preparation example: a presenter greeting

Photo: a chest-up portrait facing the camera, with mouth and chin clearly visible. Audio: a brief single-speaker greeting without music. Start with Standard and a ten-second excerpt; check lip movement, pauses and head position.

What to keep in mind

Sources

Model facts verified: 2026-09-20.

Guides and comparisons using this model

FAQ

Can I just type a script?

No. Record or synthesize the speech first, then upload the audio and portrait.

Can it animate singing?

Yes. Use a track with clear vocals. Loud accompaniment makes the task harder.

What determines video length?

The audio duration, with a maximum of 300 seconds per result.

Which portrait should I choose?

One clearly visible face, near-frontal framing and even light.

Why test a short excerpt?

It checks the particular face-and-voice combination before processing the entire recording.

Can I make two speakers?

Prepare one speaker for this tool. Process different characters separately and combine their lines in editing.

What changes the price?

The coin rate per second depends on quality, mode and, for some models, sound. Multiply it by the billable duration. Where input video is charged, its duration also counts. The complete request total is rounded once.

Do I need a separate model subscription?

No. In Mixer AI, you top up one coin balance and use it for the tasks you need. You do not need a separate subscription to this model.