A short-film generation model balancing first-and-last-frame control with synchronized audio
kling-v3 is the standard video generation model in Kuaishou Kling 3.0, suitable for creating short films from text concepts or static images. It supports flexible whole-second durations, first-and-last-frame guidance, synchronized audio, and multiple aspect ratios, making it useful for advertising shots and social content as well as for previewing motion in a scene. When you need to directly edit existing videos or combine multiple reference assets, choose a different Omni model.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Text-to-video, image-to-video; outputs a video link
Video duration
3–15 seconds, whole seconds
Generation modes
std / pro / 4k; 4k and camera motion control are mutually exclusive
Aspect ratios
16:9, 9:16, 1:1
Image guidance
First-frame image link; can be paired with a last-frame image link
Audio
Can be generated in sync; generate_audio defaults to false
Creation controls
std/pro support camera_control; cfg_scale is 0–1, with optional negative prompts
The above are the calling specifications for kling-v3 on this platform; its native creative capabilities and the available parameters for each mode should be understood separately.
Core capabilities
Turn text concepts into dynamic shots
Text-to-video is suitable for designing scenes from scratch: use prompts to specify the subject, action, environment, lighting, and camera intent, then choose the duration and aspect ratio. Relevance strength and negative prompts can help express creative constraints, making them suitable for iteratively adjusting the same shot rather than cramming an entire complex plot into one short film.
Guide changes with first and last images
Image-to-video uses a first-frame image to establish the visual starting point, and can add a last frame to specify the ending image. For product displays, character actions, or scene transitions, this approach can turn static designs into dynamic assets. Prompts should focus on how motion occurs between the two images, avoiding conflicts with the subjects and composition in the images.
Balance audio and camera motion by task
When a short film needs sound, synchronized audio generation can be enabled; when more explicit camera movement is needed, camera motion control can be used in std or pro mode. The 4k mode is suitable for tasks that prioritize output clarity, but camera_control cannot be added at the same time, so image quality settings and camera control should be weighed before creation.
Use Cases
Product Showcases and Advertising Assets
Input a static product image, describe the display action, background atmosphere, and camera changes, and generate product shots for subsequent editing. If you already have a clear final composition, you can also provide an end frame. Clearly specify the main subject and core action, complete individual shots first, then assemble them into a complete advertisement.
Vertical Content and Event Shorts
Use the event theme, scene, and subject actions as text input, select a 9:16 aspect ratio and a suitable whole-second duration, and create vertical short-form video assets. Enable generate_audio when sound is needed; if you already have a brand poster or visual design, use image-to-video to bring the design into dynamic expression.
Storyboard Previsualization and Creative Comparison
Turn storyboard descriptions or keyframes into short videos for discussing action pacing, scene atmosphere, and camera direction. Keep subject and scene descriptions consistent, adjust only one key factor per round, and make it easier to compare different options. The deliverable is video footage suitable for further editing, rather than completing all post-production at once.
How to Choose This Model
Choose the Standard Version for Regular Generation, Omni for Asset Editing
If the task is “generate a shot from a piece of text” or “make an image come alive,” the kling-v3 workflow is more direct and also provides first-and-last-frame and camera movement options. If you need to use multiple assets as independent references, or add or remove elements from existing videos or change their style, choose kling-v3-omni or kling-o1; the standard version's first-and-last-frame guidance is not equivalent to video editing.
Better Suited Than Older Versions for Flexible Durations and Sound-Enabled Creation
Compared with most older models that accept only 5-second or 10-second durations, kling-v3 can schedule shots between 3–15 seconds in whole-second increments, making it easier to match short-video pacing. It supports synchronized audio, while kling-v2-6 limits audio to pro. Existing older-version workflows can be retained; when you need flexible duration, 4K, or a combination of camera movement and audio, consider this model first.
Get Started
First determine the shots and start/end frames
For text-only, choose text2video; for image animation, choose image2video and provide start_image_url, optionally with end_image_url. Clearly describe the action for a single shot; for multiple shots, configure them separately according to this model's storyboard fields.
Choose resolution and audio separately
Specify model=kling-v3 to /kling/videos, and start with a whole-number duration of 5 seconds and std/pro. Set generate_audio=true when audio is needed; do not pass camera_control when 4k is needed.
Review the complete sequence and selected controls
Save the asynchronous task and video ID, and obtain the finished video through queries or callbacks; check the actions, shot transitions, and final frame, and listen to the audio after enabling it. For multi-shot tasks, also verify the action order and duration of each segment.
Trial suggestion: a multi-shot short film with audio
Input and goal
A craftsperson first takes out a piece of wood, then shows the carving details, and finally places the finished product on the table, in a cinematic style, accompanied by subtle crafting sounds.
Acceptance and next steps
Organize actions using whole-number durations of 3–15 seconds; use only combinations supported by the guide for multi-shot, audio, and camera position control, and do not pair 4k with camera_control.
Usage limits
4k mode cannot be used together with camera_control. If the task focuses on explicitly controlling camera movements such as dolly moves and turns, choose std or pro; if the focus is 4K output, omit this camera movement parameter and do not submit both configurations together.
Image-to-video must provide start_image_url; end_image_url can only be used as a corresponding final frame and cannot be used alone. Start and end frames guide the beginning and end of the image, rather than specifying the intermediate process frame by frame; complex changes should be split into clearer shot tasks.
kling-v3 should not use Omni's multi-image reference or reference-video editing workflow. image_list and video_list are creation methods for other models; enabling accompanying audio also does not mean that dubbing, mixing, and subtitle production are complete. Videos with audio still need to be reviewed and post-processed according to delivery requirements.
Frequently Asked Questions
How do I start generating videos with kling-v3?
Submit a request to POST /kling/videos and explicitly set model to kling-v3. For text creation, use action=text2video and fill in prompt; for image creation, use image2video and provide start_image_url, then choose the duration, aspect ratio, and audio as needed.
What durations and aspect ratios can be generated?
Integer durations from 3–15 seconds are supported, along with three aspect ratios: 16:9, 9:16, and 1:1. You can arrange the composition for advertising shots, vertical short videos, or square content; the duration should match the subject's actions, and too many continuous events should not be planned in a short clip.
How do I use start and end frames?
Select image2video and first provide the starting frame with start_image_url; add end_image_url when you need to specify the ending image. The end frame cannot be submitted without the start frame. The prompt can describe how the subject transitions from the initial state to the final state, keeping the visual guidance consistent with the action requirements.
Can 4K videos control camera movement and generate audio at the same time?
4k mode supports generate_audio but does not support camera_control. Explicitly enable audio generation if you need sound; use std or pro if you need parameterized camera movement. Audio is disabled by default, so selecting 4k alone will not automatically produce a video with sound.
How do I retrieve generated results, and can they be processed asynchronously?
You can set async=true to first obtain task_id, then query task progress and results; you can also configure callback_url to receive completion notifications. Completed results include information such as video_url, video_id, and state. Retrieve the video after confirming success, rather than treating successful task submission as completed video generation.