Video · Text → video · Image → video
Wan 3.0
Clips up to 30 seconds with sound — from text, first and last frames, or references
About the model
Wan 3.0 is Alibaba’s all-in-one video model designed for longer clips. One generation produces up to 30 seconds of video — enough for a scene to have a beginning, development, and payoff instead of a single short shot.
The model works in several modes: describe the scene in text, set a first frame or a first-and-last-frame pair, or upload up to 10 reference images so the model keeps a character’s appearance, clothing, props, and setting as motion and angles change. Audio is generated together with the picture: speech, music, and ambience follow the rhythm of the action.
Our service offers both variants:
- Wan 3.0 — the standard model at the best price.
- Wan 3.0 Prime — a faster version with the same capabilities: noticeably quicker generation at a higher cost.
Strengths
- Clip lengths from 2 to 30 seconds in one generation
- Native audio track — speech, music, and ambience generated with the video
- First-frame mode and first + last frame mode
- Up to 10 reference images for characters, clothing, props, and settings
- 480p, 720p, or 1080p output
- Aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, and adaptive
- Long prompts — up to 20,000 characters
Best for
- Complete short stories with developing action
- Product videos that keep the product's look intact
- Scenes with a recurring character from references
- Smooth transitions between a set start and end frame
Limitations
- First and last frames can't be combined with reference images
- Images — JPEG, PNG without transparency, BMP, or WEBP, up to 20 MB, sides 240–8000 px
- Cost grows with length and resolution
Prompting tips
- Describe the scene in beats — setup, action, payoff; the length leaves room for a story
- In reference mode, refer to images by order — Image1, Image2, and so on
- To move between two states, set both a first and a last frame
- Draft in 480p or with Prime, render the final version in 1080p