Video · Text → video · Image → video

Wan 2.7 Video

720p and 1080p video from text, frames, or multiple references

About the model

Wan 2.7 is a line of video models from Alibaba’s Tongyi Lab, each covering its own step of the video workflow. Compared with the previous version, it adds first- and last-frame control and generation from several references at once, along with better overall quality and temporal consistency.

Our service offers three modes:

  • Text to video — a 2–15 second clip in 720p or 1080p from a description alone, with a choice of aspect ratio.
  • Image to video — animate a single frame or a transition between a first and last frame; the model fills in the motion and keeps the subject’s look.
  • Reference to video — up to 5 images the model uses to keep characters and objects looking the same in a new scene.

Strengths

  • Three modes — from text, from an image, and from references
  • First- and last-frame control — the model fills in the motion between them
  • Up to 5 references for consistent character appearance
  • 720p or 1080p output
  • Clips up to 15 seconds (up to 10 seconds in reference mode)
  • Intelligent prompt expansion on by default

Best for

  • Animating finished illustrations and product photos
  • Transitions between two set frames
  • Clips featuring the same character across scenes
  • Quick social media videos from a text description

Limitations

  • Reference mode — up to 10 seconds and up to 5 references
  • Image mode doesn't let you choose the aspect ratio
  • Minimum resolution is 720p

Prompting tips

  • In reference mode, name references by order — "image 1", "image 2"
  • In image mode, describe the motion rather than what is already in the frame
  • The model expands short prompts on its own; detailed prompts are followed more precisely