Gemini

Google

Available via API

Veo 3.1 Fast

Generate 4, 6, or 8-second 720p videos with native audio from text, images, or reference-guided input.

Video Generation

Model Guide

About Veo 3.1 Fast

Generate 4, 6, or 8-second 720p videos with native audio from text, images, or reference-guided input.

Why choose Veo 3.1 Fast?

Veo 3.1 is Google DeepMind’s leading cinematic video model, built around stronger prompt adherence, realistic motion and physics, visual consistency, and native audio. Wizzx currently exposes the Veo Fast tier represented by the backend model key veo3, with 720p output and 4, 6, or 8-second durations.

Core strengthWhat it changes
Native audioDescribe dialogue, ambience, sound effects, or musical intent in the same brief as the visuals
Prompt adherenceDirect the subject, action, camera, lighting, and sound with a more structured shot description
Motion and physical coherenceBetter fit for scenes where object interaction and believable movement carry the idea
Multiple starting modesGenerate from text, animate a starting image, or guide the clip with visual references

Choose the right generation mode

ModeBest when
Text to videoYou want the model to design the complete shot from a written brief
Image to videoYou already have the first visual and need to define motion, camera movement, and sound
Reference to videoUp to three visual references should guide the subject, style, or scene; Wizzx requires an 8-second duration for this mode

Current Wizzx delivery scope

  • Variant: veo3_fast
  • Resolution: 720p
  • Durations: 4, 6, or 8 seconds
  • Aspect ratios: 16:9, 9:16, or Auto
  • Visual input: first/last frames or up to three references, depending on generation mode

These are the controls currently exposed by Wizzx, not the full feature surface of every Veo product. The pricing cards above come directly from the Wizzx backend catalog.

Build the prompt like a shot brief

  1. Subject and setting: Who or what is in the shot, and where is it happening?
  2. Action: Describe one clear movement or event that can develop within the selected duration.
  3. Camera: Specify framing and camera behavior, such as a locked close-up, slow dolly-in, handheld tracking shot, or aerial reveal.
  4. Light and visual language: Add time of day, light direction, palette, texture, and genre only when they support the idea.
  5. Audio: Write dialogue in quotation marks and describe ambience or sound effects separately.

Text-to-video example: A quiet neighborhood bakery before sunrise. The baker slides a tray of croissants into the oven as the camera slowly dollies past the flour-dusted counter. Warm tungsten light, cool blue street light through the window. Soft oven fan, metal tray scrape, distant delivery truck; no music.

Image-to-video example: Keep the person’s face, wardrobe, and the neon storefront unchanged. Add a gentle handheld push-in as rain falls and reflections move across the pavement. The subject looks toward camera, then smiles slightly. Light city ambience and rain, no dialogue.

Where Veo 3.1 Fast fits best

  • Short campaign shots and social video concepts
  • Product-motion studies and cinematic mood tests
  • Image animation where the starting composition must remain recognizable
  • Storyboards and pre-visualization with synchronized audiovisual intent

Native audio is a defining capability, but Google notes that natural, consistent spoken audio—especially longer speech—continues to improve. Keep dialogue short and evaluate speech separately from the visual result.

Sources

Backend pricing

Rates published by the API.

Model identity, billing unit and every price tier below are read from the backend catalog.

Fast · 720p

120

credits / second

= $0.12 USD / second

Official: $0.1520%

duration
4, 6, 8
model
veo3_fast
resolution
720p

1 credit = $0.001 USD. The final deduction returned by the API is authoritative.

当前仅开放 veo3_fast 720p;duration 支持 4/6/8 秒,默认 8 秒;reference-to-video 仅支持 8 秒

Back to all models