A Real-Time Audio-Visual Generation Framework

Vorch-Streamer Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

Menglin Han1,2*, Yang Ding1*, Yulei Lu1, Haoran Yu1,3, Xin Ma1, Junyi Chen1,4,
Zhangkai Ni2†, Lin Ma, Yaohui Wang1†

1Vorch Team 2Tongji University    3Harbin Institute of Technology, Shenzhen    4Shanghai Jiao Tong University

*Equal contribution †Corresponding authors

Overview

Vorch-Streamer post-trains LTX2.3 for causal, block-wise generation and adds implicit LLM-based speech and audio planning. The model emits synchronized audio and video before future content exists, then remains stable while its own predictions become context for every next block.

01

Causal Audio-Video Forcing

Predict synchronized blocks under a bounded context window before future frames are available.

02

Long-Horizon Self Forcing

Adapts the model to its own generated context for stable identity and motion over long rollouts.

03

27.12 FPS Streaming

Native text-to-audio-video generation exceeds the 24 FPS real-time playback threshold.

04

LLM-Based Implicit Speech and Audio Planning

Plans speech and audio progression implicitly so long-form generation stays coherent and synchronized.

Vorch-Streamer framework diagram
Related Works

Compare

Flat, same-condition comparisons across native T2AV pipelines.

Compare 01

Four methods video compare

Every method receives the same benchmark case.

Quantitative Comparison

Real-time speed without giving up long-horizon fidelity

Real-time throughput27.12 FPS

1.13× real-time playback at 24 FPS, with strong synchronization, speech accuracy, and identity preservation.

MethodParamsTaskStreamFPS Sync-C WER Identity
LTX2.322BT2AV-1.836.807.59%0.8733
JoyAI-Echo22BT2AV-3.082.449.34%0.9789
OmniForcing19BT2AVYes12.120.7598.79%0.8012
Hallo-Live11.7BT2AVYes11.510.5894.13%0.8670
Vorch-Streamer Ours22.8BT2AVYes27.126.627.92%0.9996