Four input modalities as references
Text, image, audio, and video can all guide composition, camera language, motion rhythm, and sound cues, so generation is not limited to a single prompt or first frame.
Supports image, video, and audio references with flexible motion.
Built around text, image, audio, and video inputs in one audio-video generation flow, useful for synchronized sound, visual control, and director-style video creation.
Text, image, audio, and video can all guide composition, camera language, motion rhythm, and sound cues, so generation is not limited to a single prompt or first frame.
Stable motion, physical behavior, lighting detail, and native audio can be generated together for scenes where action, ambience, dialogue, and effects need to align.
Prompts can describe camera motion, shot size, character interaction, performance rhythm, and lighting direction so clips follow storyboard intent more closely.
It can generate from references, revise specific clips, characters, actions, or story beats, and continue shots for iterative video polishing.
Pick the subscription or credit pack that fits how often you create.
A simple way to try image and video generation
For steady image, video, and editing workflows
For teams, heavy creators, and commercial work