Causal Audio-Video Forcing
Predict synchronized blocks under a bounded context window before future frames are available.
A Real-Time Audio-Visual Generation Framework
Vorch-Streamer post-trains LTX2.3 for causal, block-wise generation and adds implicit LLM-based speech and audio planning. The model emits synchronized audio and video before future content exists, then remains stable while its own predictions become context for every next block.
Predict synchronized blocks under a bounded context window before future frames are available.
Adapts the model to its own generated context for stable identity and motion over long rollouts.
Native text-to-audio-video generation exceeds the 24 FPS real-time playback threshold.
Plans speech and audio progression implicitly so long-form generation stays coherent and synchronized.
Interviews, cinematic dialogue, product presentation, and everyday conversation across a growing set of real-time rollouts.
Flat, same-condition comparisons across native T2AV pipelines.
Every method receives the same benchmark case.
| Method | Params | Task | Stream | FPS ↑ | Sync-C ↑ | WER ↓ | Identity ↑ |
|---|---|---|---|---|---|---|---|
| LTX2.3 | 22B | T2AV | - | 1.83 | 6.80 | 7.59% | 0.8733 |
| JoyAI-Echo | 22B | T2AV | - | 3.08 | 2.44 | 9.34% | 0.9789 |
| OmniForcing | 19B | T2AV | Yes | 12.12 | 0.75 | 98.79% | 0.8012 |
| Hallo-Live | 11.7B | T2AV | Yes | 11.51 | 0.58 | 94.13% | 0.8670 |
| Vorch-Streamer Ours | 22.8B | T2AV | Yes | 27.12 | 6.62 | 7.92% | 0.9996 |