Start with a first-frame image and turn static ideas into dynamic videos
minimax-i2v is MiniMax's image-to-video creation model, suitable for tasks where visuals are already prepared and you want to further design dynamic effects. Enter a first-frame reference image link and describe the actions and scene changes in text to start video generation. It uses existing images as the starting point for creation, making it suitable for product visuals, animated illustrations, and storyboard experiments. It also supports asynchronous task processing and result callbacks for easy integration into content production workflows.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Creation method
Image-to-video, using a first-frame reference image as input
Image submission
first_image_url image link; Base64 is not supported
Text guidance
prompt describes actions, scene changes, and creative intent
Generation request
POST /hailuo/videos;model=minimax-i2v;action=generate
Task processing
Supports async asynchronous submission and callback_url result callbacks
Result delivery
Returns the video_url video link, task_id, and task status information
The above lists this model's creation method and this platform's API capabilities; native parameters from other Hailuo versions are not treated as specifications for this model.
Core Capabilities
Create from existing visuals
The core workflow of minimax-i2v is generating videos from reference images. You can first define the subject, background, and composition, then submit a first-frame image to begin dynamic creation instead of relying only on text to describe the entire scene from scratch. It is suitable for workflows with existing design drafts, product images, or illustration assets that seek to explore dynamic expression.
Use text to supplement dynamic intent
The first-frame image provides the visual starting point, while the prompt explains the desired actions and scene changes. In practice, descriptions can be organized around subject movement, environmental dynamics, and camera intent—for example, having a person turn their head or water in the background ripple. These are creative instructions, not animation programs executed frame by frame.
Integrate generation tasks into application workflows
Video generation can be handled asynchronously: after submission, save the task_id first, then query the task result; you can also set callback_url to receive a JSON callback upon completion. Applications can use this to handle submission, waiting, and result display separately, and ultimately retrieve the video through video_url, avoiding treating successful submission as completed generation.
Use Cases
Dynamic Prototyping for Static Product Images
Prepare a product key visual image, describe the desired motion effects or camera intent in the prompt, and generate visual clips for review. Suitable for exploring presentation directions before formal shooting or editing; before delivery, product shape, packaging text, and labels should still be checked to confirm that the visuals meet display requirements.
Animating Illustrations and Character Scenes
Use a completed character illustration or scene image as the first frame, add descriptions of character actions, background changes, and atmosphere, and try turning static works into animated assets. It is recommended to first focus on one main action, observe character details and visual changes, then decide whether to continue refining the prompt or replace the reference image.
Dynamic Previews for Storyboard Concepts
When storyboard images are available, select each image as the first frame and write motion descriptions for each shot to generate preview assets for discussion. Manage results through task identifiers, then pass them to the editing stage to arrange the shot order; this approach is for validating ideas and does not mean automatically completing an entire film.
How to Choose This Model
Choose Image-to-Video When You Already Have Images
The main basis for choosing between minimax-i2v and minimax-t2v is the creative starting point: the former requires a reference image, while the latter generates video from text. When product images, character images, or clear storyboards are already available, choosing minimax-i2v better matches the available assets; if the visual image has not yet been determined and you want to start directly from a text concept, consider a text-to-video model.
Distinguish Standard Creation from Director Mode
minimax-i2v and minimax-i2v-director are different invocation options. The former is suitable for image-to-video creation using a first frame and prompt; if the task places greater emphasis on director-style creative control, further consider the director model. Do not directly apply camera settings from director mode to minimax-i2v, and do not assume their control features are the same simply because both belong to Hailuo.
Get Started
Prepare the Creative Starting Point for This Model
Prepare a publicly accessible image for first_image_url, and describe the subject action, camera movement, and visual elements to retain separately in the prompt.
Use the Hailuo Video Endpoint
Specify model=minimax-i2v, action=generate, and the prompt for /hailuo/videos; for image-to-video, also provide the first-frame URL, without using the MiniMax H3 content structure.
Retrieve the Video Link
Set async=true to first obtain the task_id, query it through /hailuo/tasks or receive results via callback_url; read the successful video_url, and proceed to editing only after a complete review.
Trial suggestion: subtle motion from a product still
Input and goal
Based on the first-frame product image, keep the bottle and label unchanged, slowly move the camera closer, slightly vary the lighting from left to right, and do not switch the background.
Acceptance and next steps
Provide a public URL for first_image_url, and check the subject and reflections; do not directly submit Base64 or Director-specific controls.
Usage boundaries
The image input for minimax-i2v uses a first-frame link and does not accept Base64. Before submission, confirm that the link can be accessed by the service; a local file path alone cannot replace an image URL. First prepare the asset as an accessible link, then initiate generation together with the prompt.
The reference image is the starting point for generation, not an exact constraint for every frame. When product text, character details, or complex movements are involved, check whether the finished video meets the requirements; for visual elements that must remain strictly unchanged, human review and necessary post-production adjustments are recommended.
This model should be used for image-to-video tasks and should not be assumed to include director mode presets, end-frame control, multiple reference images, or audio creation features. If the project requires precise editing, voice-over, or organization across shots, include these stages in a separate production workflow.
Frequently Asked Questions
Can minimax-i2v accept text input only?
You should provide a first-frame reference image link and create based on a text description. minimax-i2v is an image-to-video model and is not suitable for use in a text-only generation workflow. If you do not have a reference image and want to generate visuals directly from text, you can choose minimax-t2v.
How should reference images be submitted? Can Base64 be used?
Submit the image link through first_image_url; Base64 encoding is not supported. First prepare the image as a URL that the service can access, then send the request; local paths or encoded image strings cannot directly replace this link.
What should prompts focus on describing?
It is recommended to focus on the desired actions, environmental changes, and camera intent, rather than repeatedly listing every detail in the image. Start by trying one main dynamic subject, then adjust the description based on the results; prompts guide generation and are not equivalent to precise frame-by-frame control instructions.
Is it the same model as minimax-i2v-director?
No. Both are image-to-video options, but minimax-i2v-director is a separate director-mode model with a stronger focus on creative control. When using minimax-i2v, work around the first-frame image and prompt, and do not treat director-mode control settings as default features.
How do I obtain the video after asynchronous submission?
After setting async=true, first save the returned task_id, then query the task to obtain its status and results; you can also set callback_url to receive a completion callback. Once generation is complete, read video_url in the results to obtain the video; do not treat receiving a task ID as meaning the video is already complete.