Vorch-Human Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Yang Ding1*, Haoran Yu2*, Xin Ma1*, Yulei Lu1, Menglin Han3, Yaole Wang1, Siqian Yang1, Gang Yue1, Kaihao Zhang2,
Yaohui Wang1†, Lin Ma

1 Vorch Team    2 Harbin Institute of Technology, Shenzhen    3 Tongji University

* Equal contribution † Corresponding author

Paper Code (coming soon)
Overview

Human video generation is advancing rapidly, yet existing systems often solve audio-centered tasks in isolation, limiting cross-task transfer and long-form stability. We present a unified human-centric framework that supports six post-training tasks with a single joint audio-video model. It combines role-aware multi-task conditioning, an automatic pipeline for audio, timbre, identity, and multi-person binding annotations, and a lightweight context-forcing strategy for stable generation up to five minutes. Our model preserves identity and lip synchronization over long sequences while delivering high visual fidelity and strong performance against leading closed-source systems.

01

One Model, Unified Tasks

Role-aware conditioning shares knowledge across diverse generation tasks.

02

Unified Data Pipeline

Automatic audio, timbre, identity, and multi-person binding annotation.

03

Five-minute Generation

Lightweight context forcing improves long-range identity and motion stability.

Generation 01

Short-video

Expressive short-form generation across audio-driven animation, image and reference-audio conditioning, and multiple references.

Generation 02

Long-video

Extend human-centric video generation to multi-minute sequences while maintaining identity and audio-visual synchronization across extended sequences.

Related Works

Compare

Side-by-side comparisons with recent state-of-the-art audio-driven and R2AV methods.

Compare 01

Audio Driven

Five-method audio-driven animation comparison.

Compare 02

R2AV

Reference-to-audio-video results with expandable prompts.