Seedance 2.5
Build a short story, not just a moving picture: up to 30 seconds, recurring characters, your own references and sound.
When to choose it
Choose Seedance 2.5 when a shot needs a sequence of events: someone enters, notices an object, picks it up and reacts. Alongside a written brief, you can supply photographs, motion footage and audio. Give it the ingredients and a clear progression instead of relying on a vague request for a cinematic result.
- An opening, a development and an ending can fit in one generated clip rather than a sequence you must assemble first.
- References can have different jobs: a photo for appearance, a video for movement and audio for rhythm or voice.
Modes and settings
- Text mode starts without files. Frame mode uses a starting image and, optionally, an ending image.
- Reference mode accepts up to 30 images, 10 videos and 10 audio files. Video and audio each have a separate 30-second total budget.
- Video mode transforms one 4–30-second clip; the source determines the result's duration and aspect ratio.
- Mixer AI offers 480p, 720p and 1080p. For regular generation, choose a duration between 4 and 30 seconds.
Price in coins
| Quality / size | Input | Coins / second |
|---|---|---|
| 480p | With video reference | ≈ 11.02 |
| 480p | Text or images | ≈ 18.08 |
| 720p | With video reference | ≈ 24.2 |
| 720p | Text or images | ≈ 40.1 |
| 1080p | With video reference | ≈ 60.43 |
| 1080p | Text or images | ≈ 100.51 |
Billing is per second, not per fixed-length clip. The rate is multiplied by the billable duration; the complete request total is rounded up to a whole coin once.
Without video input, billing uses the output duration. With a video reference, it uses input video time plus output time at the video-reference rate. Photos and audio do not add video seconds.
≈ marks a fractional rate rounded to two decimal places for display; billing uses full precision.
What you can make
- A product spot built around an action, rather than another orbit around a package.
- A short-film moment with identifiable characters and a clear sense of where everyone stands.
- A new treatment of existing footage, changing the lighting, season or setting while retaining the scene's foundation.
How to direct the result
- Work out who stands where and which direction they move before adding lighting, lenses and texture.
- Use consecutive time ranges with no overlaps. Give each range one main event and a clear ending state.
- Assign every reference a role. When transforming footage, separate what must stay from what should change.
Example: a small moment in a coffee shop
10 seconds, one continuous shot. The person in the first photo stands on the left at a counter, with the cup from the second photo in front of them. 0–3s: a barista sets the cup on its saucer. 3–7s: the person wraps both hands around it to warm their palms. 7–10s: they look up and give a small smile. The camera slowly moves closer at chest height. Rain outside, soft warm light inside. Rainfall and a quiet clink of crockery, no speech.
What to keep in mind
- Timestamps guide pacing; they are not frame-accurate editing commands.
- A large, conflicting reference set can work against you. Start with the few files that actually matter.
- Check faces after occlusions, who holds each prop, movement direction and spoken lines.
Sources
- Prices and availability — Mixer AI catalog, updated 2026-09-18.
- Official maker site: ByteDance
Model facts verified: 2026-09-10.
Guides and comparisons using this model
FAQ
What changes from Seedance 2.0?
The practical differences are a 30-second ceiling instead of 15 and a larger reference budget. Mixer AI also offers a dedicated whole-video transformation mode for 2.5. That does not guarantee a better result for every prompt.
Should I use all 50 reference slots?
No. That is a limit, not a target. A portrait and a short motion clip may be enough; extra material can make the brief less clear.
Is a video reference the same as Video mode?
No. A reference supplies guidance for motion or the scene. Video mode transforms the uploaded clip itself and follows its duration and aspect ratio.
Can I specify the final frame?
Yes. In Frame mode, add a first and a last image. Choose compatible poses and framing, then describe the action connecting them.
Why does uploading video change the total?
Video references use a separate per-second rate, multiplied by input video time plus output time. Without video input, only output seconds are billed. The table separates both rates at each quality setting.
How do I avoid background music?
Name the sounds you need, such as footsteps and street ambience, and explicitly ask for no music or speech. Listen to the finished track separately; unwanted sound can still appear.
What changes the price?
The coin rate per second depends on quality, mode and, for some models, sound. Multiply it by the billable duration. Where input video is charged, its duration also counts. The complete request total is rounded once.
Do I need a separate model subscription?
No. In Mixer AI, you top up one coin balance and use it for the tasks you need. You do not need a separate subscription to this model.