Unified audio-visual intelligence
Vorch-Omni
Multi-Task Orchestration of Sight and Sound
Vorch Team
Text to video✦Joint A/V generation✦Reference synthesis✦Audio-driven animation✦Video editing✦Temporal extension✦
✦✦✦✦✦✦
01 / Method
Any Multimodal Reference In, Consistent Video Out.
Our unified architecture achieves this by formulating each task as a configurable quadruple of video, audio, image, and text conditions. By simply reconfiguring these multimodal references, whether a portrait image, a voice timbre, or a scene description, the same model adapts to 10+ diverse tasks on the fly, without any architectural changes.
Conditions & targets
TText instruction
VVideo / image
AAudio signal
Role-aware sequence
TaskConfig
(xv, cv, xa, ca)
- Condition mask
- Task ID
- Position type
Shared flow transformer × 48
Video streamSelf-attn → Text → A2V
Audio streamSelf-attn → Text → V2A
Selected outputs
Video frames V
Audio waveform A
02 / Qualitative showcase
