Video from text, first and last frames, and image, video or audio references.
MiniMax H3 makes 4–15 second clips at 768p or 2K with native stereo sound. Start from text, set first and last frames, or add references: up to 9 images, 3 videos and 3 audio files. Audio needs a visual input.
| Type | Video generation (text-to-video, image-to-video) |
|---|---|
| Animate a photo | Yes |
| Input frames | first and last frames |
| References | up to 9 images, 3 videos and 3 audio files; audio needs a visual input |
| Audio | yes, native |
| Clip length | 4–15s |
| Resolution | 768p, 2K |
| Prompt length | 7000 characters |
| Provider model | MiniMax H3 |
| Released | 2026-07-31 |
Video from text, first and last frames, and image, video or audio references. It is a video model by MiniMax (MiniMax H3), available on Mixer AI pay-as-you-go — the price is shown before generation.
Pay as you go, no plans — the price is shown before generation. The exact price is shown before you run it.
Yes — upload a photo as a frame or reference and the model turns it into video. Text-to-video also works.
No. Mixer AI is pay-as-you-go: you top up a balance in coins and spend it only on the generations you want. Available on the site and in the Telegram bot @addbeer_bot.