Veo 3.1 is Google DeepMind’s leading cinematic video model, built around stronger prompt adherence, realistic motion and physics, visual consistency, and native audio. Wizzx currently exposes the Veo Fast tier represented by the backend model key veo3, with 720p output and 4, 6, or 8-second durations.
These are the controls currently exposed by Wizzx, not the full feature surface of every Veo product. The pricing cards above come directly from the Wizzx backend catalog.
Text-to-video example: A quiet neighborhood bakery before sunrise. The baker slides a tray of croissants into the oven as the camera slowly dollies past the flour-dusted counter. Warm tungsten light, cool blue street light through the window. Soft oven fan, metal tray scrape, distant delivery truck; no music.
Image-to-video example: Keep the person’s face, wardrobe, and the neon storefront unchanged. Add a gentle handheld push-in as rain falls and reflections move across the pavement. The subject looks toward camera, then smiles slightly. Light city ambience and rain, no dialogue.
Native audio is a defining capability, but Google notes that natural, consistent spoken audio—especially longer speech—continues to improve. Keep dialogue short and evaluate speech separately from the visual result.