Fast · 720p
120
credits / second
= $0.12 USD / second
Official: $0.15−20%
- duration
- 4, 6, 8
- model
- veo3_fast
- resolution
- 720p
Available via API
Generate 4, 6, or 8-second 720p videos with native audio from text, images, or reference-guided input.
Video Generation
Model Guide
Generate 4, 6, or 8-second 720p videos with native audio from text, images, or reference-guided input.
Veo 3.1 is Google DeepMind’s leading cinematic video model, built around stronger prompt adherence, realistic motion and physics, visual consistency, and native audio. Wizzx currently exposes the Veo Fast tier represented by the backend model key veo3, with 720p output and 4, 6, or 8-second durations.
| Core strength | What it changes |
|---|---|
| Native audio | Describe dialogue, ambience, sound effects, or musical intent in the same brief as the visuals |
| Prompt adherence | Direct the subject, action, camera, lighting, and sound with a more structured shot description |
| Motion and physical coherence | Better fit for scenes where object interaction and believable movement carry the idea |
| Multiple starting modes | Generate from text, animate a starting image, or guide the clip with visual references |
| Mode | Best when |
|---|---|
| Text to video | You want the model to design the complete shot from a written brief |
| Image to video | You already have the first visual and need to define motion, camera movement, and sound |
| Reference to video | Up to three visual references should guide the subject, style, or scene; Wizzx requires an 8-second duration for this mode |
veo3_fastThese are the controls currently exposed by Wizzx, not the full feature surface of every Veo product. The pricing cards above come directly from the Wizzx backend catalog.
Text-to-video example: A quiet neighborhood bakery before sunrise. The baker slides a tray of croissants into the oven as the camera slowly dollies past the flour-dusted counter. Warm tungsten light, cool blue street light through the window. Soft oven fan, metal tray scrape, distant delivery truck; no music.
Image-to-video example: Keep the person’s face, wardrobe, and the neon storefront unchanged. Add a gentle handheld push-in as rain falls and reflections move across the pavement. The subject looks toward camera, then smiles slightly. Light city ambience and rain, no dialogue.
Native audio is a defining capability, but Google notes that natural, consistent spoken audio—especially longer speech—continues to improve. Keep dialogue short and evaluate speech separately from the visual result.
Backend pricing
Model identity, billing unit and every price tier below are read from the backend catalog.
120
credits / second
= $0.12 USD / second
Official: $0.15−20%
1 credit = $0.001 USD. The final deduction returned by the API is authoritative.
当前仅开放 veo3_fast 720p;duration 支持 4/6/8 秒,默认 8 秒;reference-to-video 仅支持 8 秒