All models

kling-o1

KuaishouVideo
Get your API key
kling-o1

Kling Creative Model Combining Multi-Image References and Video Editing

Kling O1 is a multimodal video model from Kuaishou Kling for generation and editing, suitable for combining images of people, products, and scenes with existing clips into clear creative instructions. Its value goes beyond animating static images, including adjusting video style, visual elements, and composition with reference materials. This platform's kling-o1 entry provides 5-second short video creation, suitable for refining ads and serialized content shot by shot.

KuaishouModel brand
VideoModel type
Multi-image · Video editingCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
modelkling-o1

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Creation methods
Text-to-video, image-to-video, multi-image reference, reference video, and video editing
Generation duration and modes
Platform entry: 5 seconds; std / pro
Aspect ratios
16:9、9:16、1:1
Image references
Up to 7 images without a reference video; up to 4 images with a reference video
Video reference
Up to 1 MP4/MOV clip; 3–10 seconds; ≤200MB; 24–60fps; width and height each 700–2160px
Audio handling
Does not generate audio; original audio can be retained or removed for video references
Result delivery
Video link, video ID, task ID, and task status; supports asynchronous queries and callbacks

O1's multimodal generation and editing positioning corresponds to its creative capabilities. The above duration, material requirements, and modes are the invocation specifications for this platform entry.

Core Capabilities

Put characters, scenes, and styles into the same shot

Multiple image references can separately provide character appearance, product form, scenes, and visual style, then use text to describe their relationship within the shot. Compared with describing appearance using only long prompts, this approach is better suited to creating a series of short videos around the same character or product; reference materials still need to be explicitly cited by their corresponding numbers in the prompt.

Make targeted edits around existing clips

Set a reference video as the base to add, remove, or modify elements in existing footage, adjust composition, or change style, color, and weather. Instructions should describe both what you want to change and what you want to preserve, such as changing the overall visual style while retaining the original motion and composition, making the editing goal clearer than a vague request to “optimize the video.”

Use video features to guide new shots

Set a video as a feature to reference its style, camera movement, or direction for the next shot, rather than directly treating it as a clip to be modified. This is suitable for continuing creation based on an existing visual approach. Image-to-video can also specify first and last frames, providing anchors for the opening and ending images, but this cannot be used together with base video editing.

Use Cases

Short-shot variations for product advertising

Input product reference images, set images, and clear action descriptions to create short shots suitable for landscape, portrait, or square placements. You can keep the product and set references unchanged while trying different actions or atmospheres, then use the resulting videos for ad editing; packaging text and small marks should be checked frame by frame before delivery.

Visual revisions of existing footage

Input a compliant short video, set it as the base, and describe the colors, weather, composition, or style that need to change, as well as the subjects and motion that should be retained. The deliverable is a regenerated video clip, suitable for comparing creative options and revising visuals; when the original sound needs to be retained, you can separately choose to preserve the original audio.

Shot continuity for serialized content

Prepare fixed character and scene reference images, or use the previous video as a feature reference, and describe the action and visual direction of the next shot. Produce one short video each time, then organize them into serialized content during editing. Consistent reference materials and instruction structure help maintain recognizability, but continuity between shots still needs to be checked.

How to choose this model

Prioritize O1 when you already have reference materials

If the core task is to combine character, product, and scene references, or make targeted edits to existing clips, O1's multi-image and video reference workflow is more worth trying first. If there is only a first image, use image2video; for combining multiple materials and video editing, use text2video with a reference list, without categorizing all material-based tasks as image-to-video.

Choose V3 Omni when you need longer clips or audio

kling-o1 and kling-v3-omni are independent models, not mode switches for the same model. O1 is suitable for 5-second reference creation and editing; if you need 3–15-second generation, synchronized audio, or native 4K, consider the corresponding V3 Omni workflow. The choice should be based on material conditions and delivery requirements, rather than judging only by the version name.

Get started

First clarify the relationship between references and editing

Use action=text2video, organize reference materials through image_list/video_list, and reference them by number in the prompt. Video refer_type=base is used for original clip editing, while feature is used for feature references.

Set up a five-second reference task

For /kling/videos, explicitly specify model=kling-o1, choose std/pro and duration=5. Arrange the number of images according to whether a video is included; base editing no longer specifies first and last frames.

Verify edits against the original material

Retrieve results through async or callback_url and save the task and video IDs; check the character, scene, and areas required to be modified. This model does not generate new audio; if the reference video is set to retain original audio, also verify the original audio track.

Trial recommendation: style editing for existing footage

Input and goal

Use the original video as the base, preserve the character's movements and camera position, change the indoor scene to a warm-toned vintage photography style, and keep the clothing and subject consistent.

Verification and next steps

Reference the video by material number in the prompt, using std/pro; O1 only generates 5 seconds, and does not transfer native audio or 4k from other models into editing.

Usage Limits

  • The O1 endpoint only generates 5-second clips and does not support 4k mode, audio generation, or camera_control. Video feature references can guide motion direction, but are not equivalent to numerical camera control; longer works should be organized through storyboard generation and post-production editing.
  • Omni-reference requests cannot use negative_prompt, cfg_scale, or camera_control. Reference images and videos must also be cited by index in the prompt; simply uploading assets will not apply them automatically, and mixing first/last-frame fields with image lists may also change asset order.
  • Reference videos must meet format, duration, frame rate, and dimension requirements. Base editing cannot specify first and last frames again; the last frame for image-to-video must be used together with the first frame. Images must be JPG/JPEG/PNG, ≤10MB, with the shortest side ≥300px; arbitrary assets cannot be submitted directly.

Frequently Asked Questions

Which action does Kling O1 use for video editing?

Reference video editing uses action=text2video, explicitly specifying model=kling-o1, and setting the asset as base in video_list. The prompt must cite that asset and describe the intended modifications; feature is used for feature reference, not for directly editing the base clip.

Why did my reference image not take effect after uploading it?

Assets in the reference list must be cited by their corresponding index in the prompt; uploading only image_list or video_list will not apply them automatically. You should also avoid mixing separate first/last-frame fields with image lists, so that changes in asset order do not cause instruction references to differ from expectations.

Can O1 generate sound or retain the original video's sound?

O1 does not support synchronized audio generation, and generate_audio should remain false. When using reference videos, you can choose to retain or remove the original audio through keep_original_sound; this is original asset audio processing and does not mean the model will generate new dialogue, music, or sound effects.

If the reference video can be up to 10 seconds, can the output also be 10 seconds?

You cannot treat the reference asset duration as the output duration. O1 accepts reference videos from 3–10 seconds, but the generation duration for this endpoint is 5 seconds. When preparing assets, also check video dimensions, frame rate, format, and file size to avoid failure at submission due to unmet asset requirements.

How can I obtain generation results in a business system?

Submit a task to POST /kling/videos and explicitly specify kling-o1. Setting async=true lets you obtain task_id first and then query the task; you can also use callback_url to receive the completed result. Determine whether it is complete based on state and success, then use video_url to obtain the video.