Model Capabilities

Video Overview

Grok Imagine video models turn a set of references, a still image, an existing clip, or a prompt into video with generated audio. grok-imagine-video-1.5 generates lip-synced speech, renders text-to-video and image-to-video at native 1080p, takes up to 14 reference images and 3 voice references, and pins first, last, and mid-video frames, in clips up to 15 seconds. grok-imagine-video-1.5-lite does text-to-video and image-to-video at the lowest price per second, reaching 1080p by upscaling 720p.

Requests are asynchronous: submit, poll with the returned request_id, then download the clip. The xAI SDK and the Vercel AI SDK poll for you (how it works).

Several generations from one character image and one voice reference made from an audio file, edited together. Voice references from your own audio files are available to trusted partners on request.
Capability
Highest quality, references, voices, and pinned frames
Lowest price, text and image to video
Video editing and extension
Modes
Text-to-videoUp to 1080p, nativeUp to 1080p, upscaled from 720pUp to 720p
Image-to-videoUp to 1080p, nativeUp to 1080p, upscaled from 720pUp to 720p
Reference imagesUp to 14 images, 15 s, 720p—Up to 7 images, 10 s, 720p
Voice referencesUp to 3 voices——
First & last frame——
KeyframesUp to 4 mid-video frames——
Video editing——Source up to 8.7 s
Video extension——Adds 2–10 s
Output
Duration1–15 s1–15 s1–15 s
Resolution480p, 720p, 1080p480p, 720p, 1080p480p, 720p
Aspect ratio1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 21:9, 5:21:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 21:9, 5:21:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 21:9, 5:2
Image-to-video aspect ratioAutomatic: matches the input imageAutomatic: matches the input imageAutomatic: matches the input image
AudioNative audio with lip-synced speechGenerated audioGenerated audio
Price per second480p $0.08720p $0.141080p $0.25480p $0.02720p $0.031080p $0.14480p $0.05720p $0.07

Reference-to-video

Reference images supply the characters, products, and locations without fixing the first frame: up to 14 per request on grok-imagine-video-1.5, 7 on grok-imagine-video. On grok-imagine-video-1.5, reference_audios gives up to three characters a preset voice, with lip-synced speech. Voice references from your own audio files are available to trusted partners on request.

  • Reference portrait of Lena
    Lena
  • Reference portrait of Sam
    Sam
  • Reference photo of a café table by a rainy window with two coffees and two croissants
    Café
References: two characters and a café
Dialogue and storytellingWide shot, then a cut over his shoulder

Sam: “I took the job in Berlin. We could leave in March.” Lena: “We? You didn’t even ask me.” Sam: “I thought you’d be happy for me.” Lena: “I am. I just thought we would… chat through it first then decide.”

grok-imagine-video-1.5 · 720p · 15 s · 16:9

14 reference images and a preset voice: a model, a room, and 12 catalog products

  1. Model0
    Modelvoice ursa
  2. Room1
    Room
  3. Sofa2
    Sofa
  4. Armchair3
    Armchair
  5. Coffee table4
    Coffee table
  6. Side table5
    Side table
  7. Rug6
    Rug
  8. Floor lamp7
    Floor lamp
  9. Bookshelf8
    Bookshelf
  10. Plant9
    Plant
  11. Wall art10
    Wall art
  12. Throw11
    Throw
  13. Cushions12
    Cushions
  14. Vase13
    Vase
Interior and e-commerce staging14 reference images and a voice

She says: “Welcome to my dream living room.”

grok-imagine-video-1.5 · 720p · 15 s · 16:9 · 1 preset voice

First, key, and last frames

On grok-imagine-video-1.5, image sets the first frame, last_frame the last, and up to four keyframes pin frames at chosen timestamps; the model fills in the motion. Pass the same image as image and last_frame to make a loop.

  • Oak tree over young green wheat
    image + last_frame
  • Oak tree over ripe golden wheat
    keyframe3 s
  • Oak tree over a harvested field
    keyframe6 s
  • Oak tree in a snowy field
    keyframe9 s

grok-imagine-video-1.5 · 720p · 12 s · 16:9

Image-to-video

The still you pass as image becomes the first frame, and the prompt describes what happens next. The output always matches the still's aspect ratio; aspect_ratio is ignored.

  • A frosted green Lunaria perfume bottle among white roses, moss, glass marbles, and a thistle
First frame
Product photo to videoLabel text stays exact

grok-imagine-video-1.5 · 1080p · 8 s

  • A woman on a balcony overlooking the sea
First frame
Portrait animation

grok-imagine-video-1.5 · 720p · 8 s

  • A sports car parked on a city street
First frame
Automotive

grok-imagine-video-1.5 · 720p · 12 s

  • A figure under the Milky Way at night
First frame
Landscape time-lapse

grok-imagine-video-1.5 · 720p · 12 s

Text-to-video

Describe the shot, including camera movement, lighting, and sound. On grok-imagine-video-1.5 and grok-imagine-video-1.5-lite, the model renders a first frame from the prompt, then animates it.

Establishing shotCrowds and legible signage

grok-imagine-video-1.5 · 720p · 10 s · 16:9

AnimationStop-motion style

grok-imagine-video-1.5 · 720p · 10 s · 16:9

Action POVExtreme camera motion

grok-imagine-video-1.5 · 1080p · 16:9

Food and product adViscous fluid

grok-imagine-video-1.5 · 720p · 8 s · 16:9

Video editing

Send a clip and an instruction to grok-imagine-video. It applies the change and keeps everything else as close to the source as it can; say what must stay the same, as these prompts do.

Source video
Edited result
Wardrobe and propsChange one thing, keep the rest

grok-imagine-video


Last updated: October 7, 2026