Veo 3.1
A short scene with both picture and sound in mind, from rustling fabric to a character's line.
When to choose it
Veo 3.1 is useful for a self-contained moment: a product reveal, a reaction or an atmospheric insert. Sound can be part of the idea from the start rather than something added to rescue the footage later. A simple action and a specific sound brief are more useful than a long string of quality adjectives.
- One brief can cover action, camera movement, speech and the surrounding sounds.
- First and last frames help define a transition; references can show a particular person or product separately.
Modes and settings
- Lite, Fast and Quality are separately priced versions. Test the idea with an affordable option, then compare another version using the same brief.
- Text mode needs no images. Frame mode takes a first image and an optional last one. Reference mode accepts 1–3 photos for subjects, products and the scene.
- New clips in Mixer AI use 720p or 1080p and run for 4, 6 or 8 seconds. Reference mode uses 8 seconds.
- Extension and 4K retrieval are separate actions after generation, not free upgrades to the original request.
Price in coins
| Version | Quality / size | Input | Duration | Coins / generation |
|---|---|---|---|---|
| Lite | 720p | Text or images | 4 s | 20 |
| Lite | 720p | Text or images | 6 s | 20 |
| Lite | 720p | Text or images | 8 s | 20 |
| Lite | 1080p | Text or images | 4 s | 23 |
| Lite | 1080p | Text or images | 6 s | 23 |
| Lite | 1080p | Text or images | 8 s | 23 |
| Fast | 720p | Text or images | 4 s | 38 |
| Fast | 720p | Text or images | 6 s | 38 |
| Fast | 720p | Text or images | 8 s | 38 |
| Fast | 1080p | Text or images | 4 s | 41 |
| Fast | 1080p | Text or images | 6 s | 41 |
| Fast | 1080p | Text or images | 8 s | 41 |
| Quality | 720p | Text or images | 4 s | 151 |
| Quality | 720p | Text or images | 6 s | 151 |
| Quality | 720p | Text or images | 8 s | 151 |
| Quality | 1080p | Text or images | 4 s | 154 |
| Quality | 1080p | Text or images | 6 s | 154 |
| Quality | 1080p | Text or images | 8 s | 154 |
Coins are Mixer AI balance units. The table shows a base request without extra variations. Multiple results and new attempts can increase the total; check it before submitting.
This table covers new clips. Extending a result or obtaining a 4K version afterwards are separate operations priced in the interface. Reference mode uses 8 seconds.
What you can make
- Product footage with tactile sound: a cap clicking, ice in a glass or paper being unwrapped.
- A brief exchange between characters or an atmospheric shot to cut into a larger edit.
How to direct the result
- Give the shot one main change: an initial state and a visible ending. Choose one clear camera movement.
- Describe sound in its own sentence. Quote dialogue exactly and name the speaker; a few seconds will not fit a long monologue.
Example: a product shot with sound
6 seconds. Close-up of a cold glass tea bottle on a wooden table. A hand slowly unscrews the cap; condensation beads remain visible on the glass. Locked-off camera, soft window light from the right. A click from the cap and quiet street ambience outside. No speech or music.
What to keep in mind
- Review written signs and pronunciation separately from the visual quality.
- Very different first and last frames can produce an awkward transformation rather than a believable transition.
Sources
- Prices and availability — Mixer AI catalog, updated 2026-09-18.
- Official maker site: Google DeepMind
Model facts verified: 2026-09-20.
Guides and comparisons using this model
FAQ
Which version should I try first?
Compare Lite and Fast pricing when testing composition. Try Quality once the scene is clear; spending more will not resolve contradictory directions.
How are frames different from references?
First and last frames define the ends of a clip. References show what a character, product or setting should look like and do not necessarily become the opening frame.
Why can't I choose 4 seconds with references?
Reference mode uses 8 seconds in this integration. The 4- and 6-second choices belong to other modes and should not be assumed to work with references.
Can I start a generation directly in 4K?
Veo's initial generation in Mixer AI uses 720p or 1080p. Obtaining 4K is a separate action on a finished result with its own price calculation.
How do I keep only scene sounds?
List the sounds you need and explicitly exclude music and speech: rain against glass and footsteps, for example. Listen to the result even when the brief is unambiguous.
Can I make a longer story?
Break it into short, complete scenes with shared references. Extension can continue a successful shot, but it does not replace editing and continuity checks across the whole story.
What changes the price?
It depends on version, quality and mode. The table prices one generation at the listed settings. Additional results, new attempts and chargeable input materials can increase the total.
Do I need a separate model subscription?
No. In Mixer AI, you top up one coin balance and use it for the tasks you need. You do not need a separate subscription to this model.