Kling 2.6
A short scene you can hear: a line of dialogue, footsteps, rainfall or room ambience.
When to choose it
Kling 2.6 suits shots where sound matters alongside the action. A portrait can say a short line; a street can have rain and footsteps. Start from text or one photo, then describe the visual action and the soundscape separately.
- Picture and sound are created in one request.
- Compare a scene with and without sound while keeping the same starting image.
Modes and settings
- Text or one opening image, with a duration of 5 or 10 seconds.
- Sound is a separate setting and changes the price.
Price in coins
| Input | Sound | Duration | Coins / generation |
|---|---|---|---|
| First frame | Without sound | 5 s | 35 |
| First frame | Without sound | 10 s | 68 |
| First frame | With sound | 5 s | 68 |
| First frame | With sound | 10 s | 133 |
| Text | Without sound | 5 s | 35 |
| Text | Without sound | 10 s | 68 |
| Text | With sound | 5 s | 68 |
| Text | With sound | 10 s | 133 |
Coins are Mixer AI balance units. The table shows a base request without extra variations. Multiple results and new attempts can increase the total; check it before submitting.
What you can make
- A short presenter line or character reaction.
- An atmospheric insert with a cup, a door or footsteps.
How to direct the result
- Name the speaker and put the exact short line in quotation marks.
- List background sounds separately and say whether music belongs in the scene.
Example: a line in a workshop
5 seconds. The person in the photo stands at a workbench, lifts a finished wooden toy and says, 'Ready to play.' Locked waist-up shot. Preserve face and clothing. Voice, a light wooden knock and quiet room ambience, no music.
What to keep in mind
- One photo does not lock the final frame.
- Pronunciation and lip movements can disagree. Listen before publishing.
Sources
- Prices and availability — Mixer AI catalog, updated 2026-09-18.
- Official maker site: Kuaishou
Model facts verified: 2026-09-20.
Guides and comparisons using this model
FAQ
Do I need a separate voice-over?
Not necessarily. Enable sound and describe it in the prompt. Check the finished track regardless.
Can I supply recorded speech?
This page generates sound from a description. For a portrait driven by a recording, use the separate avatar tool.
Can I disable sound?
Yes. The table lists sound-on and sound-off configurations separately.
Why does the spoken line get cut short?
Shorten it, remove competing actions or choose ten seconds.
Can I add several photos?
This option uses one opening image, not separate character files.
Why did it add music?
Specify the desired soundscape, such as voice and room tone only, without music. Generation may still depart from the brief.
What changes the price?
It depends on version, quality and mode. The table prices one generation at the listed settings. Additional results, new attempts and chargeable input materials can increase the total.
Do I need a separate model subscription?
No. In Mixer AI, you top up one coin balance and use it for the tasks you need. You do not need a separate subscription to this model.