Reference to Video
Reference to Video is a multimodal generation workspace for creating coherent videos from image, video, and audio references. Instead of asking one text prompt to define every production decision, each reference can anchor a specific role such as character identity, product appearance, motion, camera behavior, style, sound, or timing.
Core capabilities
- Upload multiple reference images, videos, and audio files in one generation brief
- Preserve faces, hair, clothing, products, packaging, and visual worlds across scenes
- Guide gestures, performance timing, framing, and camera movement from video
- Use audio references for voice, music, rhythm, and synchronization on compatible models
- Start from templates for characters, UGC ads, fashion, anime, products, music, and mascots
- Choose among compatible generation models and output settings
Reference inputs can contain copyrighted media, voices, and identifiable people. Obtain permission, keep each reference's role clear, and review generated output for unwanted likeness or brand changes.

