Video · Text → video · Image → video

PixVerse V6

Cinematic video up to 15 seconds in 1080p — with native audio and consistent characters

About the model

PixVerse V6 is a video model built for cinematic results: realistic motion, expressive camera work, and detailed environments. It understands rich prompts with several actions and camera changes, and characters keep their appearance and clothing throughout the clip.

Its key difference is native audio: the model can generate a soundtrack synced to the picture in the same pass.

Our service offers two modes:

  • Text-to-video — a clip from your description, with a choice of aspect ratio.
  • Image-to-video — bring an existing image to life.

Strengths

  • Native audio generation synced to the video
  • Cinematic visuals, realistic motion, and advanced camera moves
  • Understands prompts with multiple actions, camera changes, and scene progression
  • Keeps a character's appearance, clothing, and identity consistent
  • 360p, 720p, or 1080p and eight aspect ratios in text-to-video

Best for

  • Ad clips and product videos
  • Short story-driven videos and animated scenes
  • Vertical clips for social media
  • Bringing product and character images to life

Limitations

  • Clip length is 4 to 15 seconds
  • Input images up to 20 MB — JPG, PNG, WebP
  • Prompts must be at least 3 characters long

Prompting tips

  • Describe the scene as a sequence of actions — the model handles story progression well
  • Specify camera movement — push-in, tracking shot, orbit
  • Turn on audio when you need a finished clip without separate sound work
  • Use 360p for tests and 1080p for the final clip