Can 1.5:official generate videos using only text?
No. This model uses image-to-video generation and requires an image_url. Text prompts are used to supplement actions, camera movements, or scene changes in the image. If you do not yet have image assets, you can use grok-imagine-video:official or fast:reverse, which support text-to-video generation.
Do I still need to write a prompt besides the image?
You can generate image-to-video without a prompt, but adding one is recommended when you have clear motion goals. Descriptions should focus on how the subject moves, how the camera changes, and what happens in the scene. Avoid requesting too many conflicting actions at once so the short clip's goal is easier to determine.
How long can the videos be, and what resolutions are available?
This model supports 1–15 seconds, with 6 seconds as the default; available resolutions are 480p, 720p, or 1080p, with 480p as the default. It is best to first create short drafts to check motion, then generate higher-resolution candidates for ideas you are satisfied with, while still reviewing the specific footage at the end.
How do I retrieve asynchronously generated videos?
After setting async to true, save the returned task_id, then query the result through /grok/tasks; you can also provide a callback_url to receive completion notifications. Only when the status reaches succeeded and you obtain video_url should the result be considered a downloadable video asset.
How do I ensure I am calling this model?
When submitting a request to /grok/videos, explicitly set model to grok-imagine-video-1.5:official and provide image_url at the same time. Do not omit the model or use a name containing preview; retain task_id and trace_id to help associate assets, task status, and generated results.