Image-and-Text Generation and Video Extension Model for Short-Form Shot Creation
kling-v1-6 is a video generation model in Kuaishou Kling's V1 series, suitable for turning text concepts or static images into short shots, then continuing to create from existing clips. It provides text-to-video, image-to-video, and video extension. In pro mode, you can also use first and last frames to constrain the beginning and end of the visuals. When characters need to speak, you can combine photos with existing audio to animate them and synchronize lip movements.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Video creation methods
Text-to-video, image-to-video, and existing video extension
Generation duration
This platform's video entry: 5 or 10 seconds
Generation modes
std / pro; 4k mode is not supported
Video aspect ratios
16:9, 9:16, 1:1
First and last frame control
Image-to-video requires a first frame; pro can use a last frame, while std does not support last frames
Audio method
Native audio is not supported; talking photos use external audio for driving
Task delivery
Video link, video ID, task ID, and status; supports asynchronous processing and callbacks
The durations, modes, and operational scope above correspond to this platform's kling-v1-6 entry. Photo lip-sync is a combined creation feature and is not equivalent to the model's native audio.
Core Capabilities
Learn what kling-v1-6 can bring to your work.
From Text Concepts to Animated Images
When no existing image is available, use prompts to describe the subject, environment, and actions to generate a shot; when product images, character images, or scene images are already available, use the first frame to fix the starting point, then describe the changes you want to happen. These two approaches are suitable for concept exploration and asset recreation respectively, allowing short videos to begin with clear visual goals.
Use First and Last Frames to Set Shot Start and End
Image-to-video in pro mode can specify both a first frame and a last frame to arrange changes in character poses, before-and-after product presentation states, or scene transitions. The last frame must work together with the first frame and cannot be used independently. It constrains the starting and ending images of a shot, not a frame-by-frame trajectory, and cannot replace dedicated camera movement control parameters.
Continue Creating from Existing Clips
Video extension can be used in std and pro modes: retain the generated result's video_id, then submit prompts for the next actions or plot points. This lets you first establish a short shot and then continue with subsequent content, without rebuilding the scene from text each time. After the task is complete, retrieve the result through the video link for further editing.
Use Cases
Start with specific tasks to find where the model can be effective.
Turn Static Product Assets into Short Videos
Provide a product first-frame image, describe the presentation actions, background atmosphere, and visual changes, and generate short footage suitable for editing. Portrait format can be used for mobile display, while landscape format can be used for presentation visuals; if a clear ending composition already exists, add a last frame in pro, then supplement the finished video with captions, music, and brand information.
Storyboard Previsualization and Plot Continuation
Write a storyboard as a prompt with a clear subject, actions, and environment, first generate a single shot, then use video extension to try subsequent plot developments. The deliverable is viewable, editable video clips, suitable for discussing pacing and visual direction before formal shooting or production, rather than directly replacing a complete multi-shot editing project.
Photo Talking Heads and Character Explanations
Prepare a clear front-facing photo of a single person and recorded audio, select kling-v1-6 in the talking photo feature, and add requirements for movements or expressions. A single task completes photo animation and audio lip-syncing, returning the final talking-head video and intermediate animation clips, suitable for short explanations, character greetings, or content previews.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
When to Choose V1.6 and pro
If the task involves image animation, short-shot generation, or continuing an existing clip, V1.6 offers more direct operation combinations; choose pro when you need to control the ending frame. Compared with kling-v1, you cannot carry over the older version's camera movement parameters to V1.6: kling-v1 camera movement has a 5-second condition, while V1.6 does not support camera_control, so choose according to actual control needs.
When to Switch to Later Models
If you need to generate visuals and sound simultaneously, consider kling-v2-6 pro, which supports accompanying audio; if you need flexible durations of 3–15 seconds or 4k mode, consider kling-v3. V1.6 is better suited to short-video workflows that clearly use 5 or 10 seconds and produce sound separately. External audio lip-sync and native accompanying audio are different tasks and should not be mixed.
Getting Started
From a small-scale task to formal integration.
01
Prepare Tasks and Materials
Clarify the objective, required inputs, and output requirements, and use real business examples as a starting point.
02
Try It in the API Debugging Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limitations
Before formal use, understand output quality and capability scope.
kling-v1-6 does not support generate_audio, camera_control, or 4k mode. You can describe camera intentions in prompts, but this is not equivalent to parameterized camera movement; if the finished video must include sound, add voiceover and music separately, or use the photo-with-audio lip-sync feature.
std image-to-video cannot specify an end frame; pro end frames must also be submitted together with the start frame. Start and end frames are suitable for expressing beginning and ending compositions, but do not guarantee that all intermediate actions will follow a preset path. Before production, clearly specify subject changes and action goals, then review the generated clip.
For talking photos, use a clear front-facing image of one person. The driving audio must be an accessible link and supports mp3, wav, m4a, and aac, with a maximum size of 5MB. Audio length is recommended not to exceed the selected video duration, to avoid placing long spoken segments directly into short-video tasks.
Frequently Asked Questions
Answers to common questions about using kling-v1-6.
How do I ensure kling-v1-6 is used when making a request?
Explicitly specify model=kling-v1-6 in video generation or talking photo requests rather than relying on the default selection when omitted. For standard videos, also choose text2video, image2video, or extend, and provide a prompt, first-frame image, or existing video ID according to the task.
How long a video can V1.6 generate?
The video generation endpoint supports selecting 5 or 10 seconds, and talking photos also offer these two durations. To continue creating, you can use extend to submit an existing video_id and a follow-up prompt; extension continues an existing clip and does not mean generating a video of any length in a single request.
Why can't the std mode use an end frame?
The std version of kling-v1-6 does not support end-frame constraints. When you need a specific ending image, use pro image2video and provide both start_image_url and end_image_url; providing only an end frame cannot complete this type of task, as a first frame is still required input.
Does talking photo mean V1.6 can generate sound?
No. V1.6 does not support native audio; talking photos use the audio you provide: the photo is first animated, then lip-synced to the audio. You need to submit image_url and audio_url. Prompts are mainly used for movements or expressions during the animation stage, not for generating speech content.
How should generation tasks be integrated into an application?
You can use async=true to obtain a task_id and then query the task status, or configure callback_url to receive the completed result. Applications should distinguish between a submitted task and a completed video, and read video_url after completion; if you later need to extend it, you should also save video_id and not use task_id as a substitute.
Model information · Updated: 2026-10-01. See the API and pricing sections for request parameters and billing rules.