Grok Imagine Video 1.5 and Seedance 2.0 were compared on matching tests: physics, text, fighting, dialogue and a complex storyboard. In Mixer AI, Grok reaches 30 seconds and accepts up to 7 images, while Seedance 2.0 reaches 15 seconds but additionally accepts video references.
| Spec | Grok Imagine Video | Seedance 2.0 |
|---|---|---|
| Model variants | Grok Imagine Video | Mini / Fast / Pro |
| Generation modes | from text / from references | from text / from first and last frames / from references |
| First frame | No | Yes |
| Last frame | No | Yes |
| Input photos | up to 7 | up to 9 |
| Input videos | No | up to 3 |
| Input audio files | No | up to 3 |
| Input video length | No | 2–15 s |
| Input audio length | No | 2–15 s |
| Extra modes | normal / fun / spicy | No |
| Output audio | No | Can be enabled or disabled |
| Minimum duration | 6 s | 4 s |
| Maximum duration | 30 s | 15 s |
| Resolutions | 480p / 720p / 1080p | 480p / 720p / 1080p / 4K |
| Aspect ratios | 9:16 / 16:9 / 1:1 / 2:3 / 3:2 | 9:16 / 16:9 / 1:1 / 21:9 / 4:3 / 3:4 / adaptive |
| Prompt limit | 5,000 characters | 20,000 characters |
| Price, 1 s · 1080p | ≈ 5.12 coins | ≈ 61.32 coins · Pro |
Grok's strong text result and one physics result did not repeat across every round. Seedance more often preserved long causal sequences inside the scene. Grok still has a concrete product-level distinction in Mixer AI: a 30-second maximum, versus 15 seconds for Seedance 2.0.
Seedance 2.0 followed the requested scene sequence much more accurately in this direct test.
Grok preserved the lettering better in that specific A/B. It is one observed test, not a guarantee for every text-heavy scene.