Video · Text → video · Image → video
Grok Video 1.5
Brings images to life with realistic motion and built-in sound
About the model
Grok Video 1.5 is the next generation of xAI’s video model. The key difference from the first version is that sound is created together with the picture: dialogue, sound effects, ambience, and music stay in sync with the action, with no separate voiceover needed.
The model is strongest in image-to-video mode: at launch it took first place on the Image-to-Video Arena leaderboard. Motion looks realistic, expressions and lighting are vivid, and detailed directions for action and camera are followed precisely.
For longer clips up to 30 seconds, see Wan 3.0; for work with audio and video references, see Seedance 2.0.
Strengths
- Strongest at image-to-video — it topped the Image-to-Video Arena leaderboard at launch
- Sound generated with the video — speech, effects, ambience, and music
- Realistic facial expressions, lighting, and motion physics
- Follows detailed prompts closely — actions, camera moves, character behavior
- Up to 7 reference images
- Clips from 2 to 15 seconds, 480p or 720p
Best for
- Animating illustrations, covers, and product photos
- Short social media clips with sound
- Ad and promo videos
- Character scenes where lively expressions matter
Limitations
- Resolution up to 720p
- With a single uploaded image, the aspect ratio follows that image
- Images — JPEG, PNG, or WEBP up to 20 MB
Prompting tips
- When animating an image, describe motion and camera rather than what's in the frame
- Direct the sound right in the prompt — lines, noises, music
- Use 480p for drafts and 720p for the final version