A video model for rapid iteration of ads and product shorts
veo3-fast is a video generation model in the Google Veo 3 series designed for speed and creative iteration, suitable for turning advertising concepts, product images, and animation scripts into dynamic clips. On this platform, you can start creating from text, or use first-frame or first-and-last-frame images to structure shots, combined with landscape and portrait formats, 1080p retrieval, and asynchronous task processing to form a workflow from creative testing to asset delivery.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Model positioning
Google Veo 3 Fast, prioritizing speed and rapid iteration
720p by default; optional 1080p, gif; 4k not supported
HD retrieval
get1080p with the video ID from the generation result
Task delivery
Supports asynchronous processing and callbacks; returns video link and status upon completion
Fast's speed-first positioning is a native model characteristic; the number of images, aspect ratios, and HD retrieval method are calling specifications of this platform.
Core capabilities
Turn creative comparisons into dynamic samples
The core value of veo3-fast is rapidly exploring different approaches. For the same product selling point, you can separately describe the scene, subject actions, camera movement, and visual style to generate comparable clips. It is ideal for selecting advertising directions before formal production, rather than packing multiple complete stories into a single request.
Structure shots with opening and closing images
Image-to-video can start from one image, or use two images to define the opening and ending. Product stills, character concept art, or scene design images can all become starting points for creation; prompts describe the action and camera changes in between, moving assets from static displays toward purposeful dynamic expression.
Integrate generation tasks into the asset workflow
Video creation can use asynchronous submission, so your application does not need to keep waiting for a connection. After submission, save the task_id, receive results through queries or callbacks, and then retrieve the video_url based on status. When an HD version is needed, use the video ID from the generation result to retrieve 1080p, making it easier to integrate with asset libraries and review workflows.
Applicable Scenarios
Ad Concept and Opening Tests
Enter product selling points, audience, and shot descriptions, then try lifestyle scenes, studio presentations, or animated expressions respectively, delivering a set of dynamic samples for team review. It is recommended to adjust only one creative variable at a time, such as the background, action, or camera distance, to make it easier to compare which opening is more suitable for subsequent production.
Dynamic Product Image Presentations
Use a product image as the first frame and describe the desired presentation action, such as the camera slowly moving closer, lighting changes, or the product entering a usage scenario. If a specific final composition already exists, you can also provide an end frame to deliver clips suitable for product introductions and social content editing, rather than using multiple product images as source material for blending.
Animated Explanations and Training Previews
Break an explanatory script into independent visual segments, clearly specify the subject, environment, and action for each segment, then generate animation mockups or operational scenarios. The deliverables are suitable for discussing pacing, supplementing training materials, or validating storyboards; the final course still requires organized editing, subtitles, and reviewed explanatory content.
How to Choose This Model
Choose Fast for Exploration, Compare the Standard Version for Final Review
When the task is to test multiple ad concepts, prepare dynamic samples for a product catalog, or repeatedly revise shot expressions, prioritize veo3-fast. If the main goal shifts to image quality, compare actual results with the veo3 standard version from the same series: the former emphasizes iteration speed, while the latter is a quality-first mode; do not substitute fixed processing times or uniform image-quality differences for footage review.
Choose a Version Based on Reference Method and Clarity Requirements
For text creation, a single-image first frame, or two-image first and last frames, veo3-fast already provides corresponding workflows. If the task is to blend multiple independent reference images, choose veo31-fast-ingredients; if 4k output is explicitly required, consider veo31, which supports that option. Two first and last frames define temporal endpoints and are not equivalent to multi-image content blending.
Get Started
Choose Text or Start and End Frames
For text only, use text2video; for image-driven generation, use image2video. One image in image_urls is the first frame, and two images are the first and last frames; describe the action that occurs in between.
Choose the Dedicated Creation Endpoint
Specify model=veo3-fast at /veo/videos, choose an action based on text or images, and fill in prompt. Start with aspect_ratio=16:9 and resolution=720p, then increase the output tier according to this model's supported range.
Save Task and Video IDs Separately
After setting async=true, save the task_id, and retrieve the completed video through /veo/tasks or a callback; verify the visuals and audio, and save the video ID and downloaded file for subsequent production.
Trial suggestion: short advertising creative test shoot
Input and goal
A glass of lemon soda on the table slowly rotates in the light, bubbles rise upward, the camera gently pushes in, and the only sounds are ice cubes and bubbles.
Acceptance and next steps
First compare visuals and sound effects to choose a direction, using text or a single first frame; Fast does not promise a fixed speed improvement for every task.
Usage boundaries
veo3-fast does not support 4k output, and writing “4k” in the prompt should not be expected to change the actual resolution. When 1080p is needed, choose the corresponding output option, or use get1080p on an already generated video; subsequent operations use the video ID, not the task_id returned by asynchronous submission.
A maximum of two image inputs can be used, respectively for first-frame or first-and-last-frame creation, and they cannot be treated as arbitrary multi-image fusion. The two images should have an explainable action or scene relationship; if the starting and ending compositions differ greatly, it is recommended to clearly state the transition intent and check changes to the subject in the intermediate frames.
Fast represents a speed-prioritized creative trade-off, not a promise of a fixed completion time. It is suitable for gradually validating ideas with independent clips; continuous narratives, strict brand details, or training steps should be assembled only after review, and a single generation cannot be directly considered a completed advertisement or course.
Frequently Asked Questions
What is the difference between veo3-fast and veo3?
Both belong to the Veo 3 series. veo3-fast is better suited for rapid iteration, ad variations, and product samples; veo3 is the quality-first mode. You can first use Fast to screen ideas, then compare standard-version results for selected shots, rather than assuming there is a fixed difference in speed or image quality between the two.
How do I use one versus two reference images?
Use image2video and pass image links in image_urls. One image serves as the first frame; two images serve as the first and last frames respectively. The prompt describes the actions and camera changes in between. This method is not multi-image blending, so it is recommended to prepare start and end images around the same subject or a clearly defined transition relationship.
Can it generate vertical or 4k videos?
For vertical video, set aspect_ratio to 9:16; use 16:9 for horizontal video. veo3-fast does not support 4k; the default output is 720p, and you can also choose 1080p or gif. Aspect ratio determines composition proportions, while resolution determines output clarity, so they should be set separately.
How do I get a 1080p version of an existing video?
Submit the get1080p operation to /veo/videos and use video_id to specify the generated video. This value comes from the id in the completed result data, not task_id. When saving generation results, it is recommended to retain the task ID, video ID, and video link at the same time for easier subsequent processing.
Do I have to wait for the request to return during generation?
No. After setting async to true, you can first obtain task_id and then query the task result; you can also provide callback_url to receive a result push upon completion. After obtaining the completed result, check state and read video_url to download the video or incorporate it into your own review and asset-management workflow.