One Model, Unified Tasks
Role-aware conditioning shares knowledge across diverse generation tasks.
Human video generation is advancing rapidly, yet existing systems often solve audio-centered tasks in isolation, limiting cross-task transfer and long-form stability. We present a unified human-centric framework that supports six post-training tasks with a single joint audio-video model. It combines role-aware multi-task conditioning, an automatic pipeline for audio, timbre, identity, and multi-person binding annotations, and a lightweight context-forcing strategy for stable generation up to five minutes. Our model preserves identity and lip synchronization over long sequences while delivering high visual fidelity and strong performance against leading closed-source systems.
Role-aware conditioning shares knowledge across diverse generation tasks.
Automatic audio, timbre, identity, and multi-person binding annotation.
Lightweight context forcing improves long-range identity and motion stability.
Expressive short-form generation across audio-driven animation, image and reference-audio conditioning, and multiple references.
Extend human-centric video generation to multi-minute sequences while maintaining identity and audio-visual synchronization across extended sequences.
Side-by-side comparisons with recent state-of-the-art audio-driven and R2AV methods.
Five-method audio-driven animation comparison.
Reference-to-audio-video results with expandable prompts.