All models

minimax-i2v

MiniMaxVideo
Get your API key
minimax-i2v

Start with a first-frame image and turn static ideas into dynamic videos

minimax-i2v is MiniMax's image-to-video creation model, suitable for tasks where visuals are already prepared and you want to further design dynamic effects. Enter a first-frame reference image link and describe the actions and scene changes in text to start video generation. It uses existing images as the starting point for creation and is suitable for product visuals, animated illustrations, and storyboard prototyping. It also supports asynchronous task processing and result callbacks, making it easy to integrate into content production workflows.

MiniMaxModel brand
VideoModel type
VideoTask capability

Specifications and API features

Clarify capacity, inputs and outputs, and invocation methods before choosing a model.

Creation method
Image-to-video, using a first-frame reference image as input
Image submission
first_image_url image link; Base64 is not supported
Text guidance
prompt describes actions, scene changes, and creative intent
Generation request
POST /hailuo/videos;model=minimax-i2v;action=generate
Task processing
Supports async asynchronous submission and callback_url result callbacks
Result delivery
Returns video_url video link, task_id, and task status information

The above lists this model's creation method and this platform's invocation features; native parameters from other Hailuo versions are not treated as specifications for this model.

Core Capabilities

Learn what minimax-i2v can bring to your work.

Create from Existing Images

The core workflow of minimax-i2v is generating video based on reference images. You can first determine the subject, background, and composition, then submit a first-frame image to begin dynamic creation, without having to describe the entire scene from scratch using only text. It is suitable for workflows that already have design drafts, product images, or illustration assets and want to explore their dynamic expression.

Use Text to Add Motion Intent

The first-frame image provides a visual starting point, while the prompt is used to describe the desired actions and scene changes. In actual creation, descriptions can be organized around subject actions, environmental movement, and camera intent, such as having a person turn their head or making the water in the background ripple; these are creative instructions, not animation programs executed frame by frame.

Integrate Generation Tasks into Application Workflows

Video generation can be handled asynchronously: after submission, first save the task_id, then query the task result; you can also set callback_url to receive a JSON callback upon completion. Applications can use this to handle submission, waiting, and result display separately, and ultimately obtain the video through video_url, avoiding treating a successful submission as completed generation.

Use Cases

Start with specific tasks to find where the model can be effective.

Dynamic Prototyping for Product Still Images

Prepare a product key visual image, describe the desired dynamic effect or camera intent in the prompt, and generate visual clips for review. This is suitable for exploring presentation directions before formal shooting or editing; when delivering, you should still check the product shape, packaging text, and markings to confirm that the image meets display requirements.

Animating Illustrations and Character Images

Use a completed character illustration or scene image as the first frame, add descriptions of character actions, background changes, and atmosphere, and try turning static work into dynamic assets. It is recommended to first focus on one main action, observe character details and changes in the image, then decide whether to continue adjusting the prompt or replace the reference image.

Dynamic Preview of Storyboard Concepts

When storyboard images are already available, you can select each image as a first frame and write motion descriptions for each shot to generate preview assets for discussion. Manage results through task identifiers, then pass them to the editing stage to organize the shot sequence; this approach is used to validate creative ideas and is not equivalent to automatically completing an entire film.

How to Choose This Model

Choose based on task complexity, input materials, and expected results.

Choose Image-to-Video When You Have an Image

The primary basis for choosing between minimax-i2v and minimax-t2v is the creative starting point: the former requires a reference image, while the latter generates video from text. If you already have product images, character images, or clear storyboards, choosing minimax-i2v is better suited to your materials; if you have not yet determined the visual scene and want to start directly from a text concept, consider a text-to-video model.

Distinguish Between Standard Creation and Director Mode

minimax-i2v and minimax-i2v-director are different calling options. The former is suitable for image-to-video generation using a first frame and prompt; if the task places greater emphasis on director-style creative control, you can further consider the director model. Do not directly apply director mode camera settings to minimax-i2v, and do not assume their control functions are the same merely because both belong to Hailuo.

Getting Started

From a small-scale task to full integration.

01

Prepare the Task and Materials

Define the objective, required inputs, and output requirements, using real business examples as a starting point.

02

Try It in the API Debugging Area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate According to the API Documentation

Retain the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage Limits

Before formal use, understand the output quality and scope of capabilities.

  • Image input for minimax-i2v uses a first-frame link and does not accept Base64. Before submission, confirm that the service can access the link; a local file path alone cannot replace an image URL. First prepare your materials as accessible links, then initiate generation together with a prompt.
  • The reference image is the generation starting point, not a precise constraint for every frame. When product text, character details, or complex actions are involved, check whether the finished video meets requirements; for visual elements that must be strictly preserved, human review and necessary post-production adjustments are recommended.
  • This model should be used for image-to-video tasks and should not be assumed to include director mode presets, last-frame control, multiple reference images, or audio creation capabilities. If a project requires precise editing, voice-over, or cross-shot organization, incorporate these stages into a separate production workflow.

Frequently Asked Questions

Answers to common questions about using minimax-i2v.

Can minimax-i2v accept text only?

You should provide a first-frame reference image link and create with a text description. minimax-i2v is an image-to-video model and is not suitable for use in a text-only generation workflow. If you do not have a reference image and want to generate visuals directly from text, you can choose minimax-t2v.

How should reference images be submitted? Can I use Base64?

Submit the image link through first_image_url; Base64 encoding is not supported. First prepare the image as a URL accessible to the service, then make the request; local paths or encoded image strings cannot directly replace this link.

What should the prompt focus on describing?

It is recommended to focus on the desired actions, environmental changes, and camera intent, rather than repeatedly listing every detail in the image. Start by trying one main dynamic subject, then adjust the description based on the results; prompts guide generation and are not equivalent to precise frame-by-frame control instructions.

Is it the same model as minimax-i2v-director?

No. Both are image-to-video options, but minimax-i2v-director is a separate director-mode model with a stronger focus on creative control. When using minimax-i2v, work around the first-frame image and prompt, and do not treat director-mode control settings as default features.

How do I get the video after submitting asynchronously?

After setting async=true, first save the returned task_id, then query the task to obtain its status and result; you can also set callback_url to receive a completion callback. Once generation is complete, retrieve the video from video_url in the result; do not regard receiving a task ID as meaning the video is already complete.

Model information · Updated: 2026-10-01. For request parameters and billing rules, see the API and pricing sections.

Use minimax-i2v for your next task

Start with a clear goal, and judge from actual results whether it suits your work.