Both Motion Control versions use the same basic workflow in Mixer AI: one image defines the character and one driving video supplies the motion. A direct example shows the same motion rendered in 2.6 and 3.0 side by side.
| Spec | Kling 2.6 Motion Control | Kling 3.0 Motion Control |
|---|---|---|
| Generation modes | Motion Control | Motion Control |
| Input photos | 1, required | 1, required |
| Input videos | 1, required | 1, required |
| Input video length | 3–30 s | 3–30 s |
| Background source | No | from photo or video |
| Minimum duration | 3 s | 3 s |
| Maximum duration | 30 s | 30 s |
| Resolutions | 720p / 1080p | 720p / 1080p |
| Aspect ratios | Follows the input | Follows the input |
| Prompt limit | 2,500 characters | 2,500 characters |
| Price, 1 s · 1080p | ≈ 12.21 coins | ≈ 17.91 coins |
Because the Mixer AI input form is nearly identical for 2.6 and 3.0, a controlled A/B is straightforward: reuse the same image and driving video. Check more than pose matching—inspect the face through turns, fingers, background movement and body edges.
In the direct side-by-side, the reviewer primarily called out more accurate movement and better expression/face consistency.
The basic input pattern is the same: one image plus one driving video. Both list 720p/1080p.