Video · Text → video · Image → video

Grok Video 1.5

Brings images to life with realistic motion and built-in sound

About the model

Grok Video 1.5 is the next generation of xAI’s video model. The key difference from the first version is that sound is created together with the picture: dialogue, sound effects, ambience, and music stay in sync with the action, with no separate voiceover needed.

The model is strongest in image-to-video mode: at launch it took first place on the Image-to-Video Arena leaderboard. Motion looks realistic, expressions and lighting are vivid, and detailed directions for action and camera are followed precisely.

For longer clips up to 30 seconds, see Wan 3.0; for work with audio and video references, see Seedance 2.0.

Strengths

  • Strongest at image-to-video — it topped the Image-to-Video Arena leaderboard at launch
  • Sound generated with the video — speech, effects, ambience, and music
  • Realistic facial expressions, lighting, and motion physics
  • Follows detailed prompts closely — actions, camera moves, character behavior
  • Up to 7 reference images
  • Clips from 2 to 15 seconds, 480p or 720p

Best for

  • Animating illustrations, covers, and product photos
  • Short social media clips with sound
  • Ad and promo videos
  • Character scenes where lively expressions matter

Limitations

  • Resolution up to 720p
  • With a single uploaded image, the aspect ratio follows that image
  • Images — JPEG, PNG, or WEBP up to 20 MB

Prompting tips

  • When animating an image, describe motion and camera rather than what's in the frame
  • Direct the sound right in the prompt — lines, noises, music
  • Use 480p for drafts and 720p for the final version