Hallo-Live: Real-Time Text-Driven Audiovisual Digital Humans

·Toolin Editorial Team

An open-source real-time digital human generation system: text input synchronously generates talking video and speech, with 20.38 FPS throughput and 0.94-second end-to-end latency

Hallo-Live: Real-Time Text-Driven Audiovisual Digital Humans

Text-driven audiovisual digital humans are moving from "can generate" to "can interact in real time." Hallo-Live, a project from Shanghai Innovation Institute, Fudan University, and other institutions, achieves 20.38 FPS throughput and 0.94-second end-to-end latency on two NVIDIA H200 GPUs, while keeping visual quality and audio-visual sync close to the teacher model's level.

Paper: arxiv.org/abs/2604.23632 Code: github.com/fudan-generative-vision/Hallo-Live

What Is Hallo-Live

Traditional audio-driven digital humans only need lip sync, but a text-driven approach has to do two things at once: first "understand" the character, scene, and tone in the text, then synchronously generate the matching talking video and speech. Lip shapes, pronunciation, expressions, and even upper-body movements all have to land on the same timeline.

Against the teacher model Ovi, Hallo-Live's key metrics:

MetricHallo-LiveImprovement
Throughput20.38 FPS16.0x increase
End-to-end latency0.94 secondsdown 99.3%

It supports anime styles, photorealistic characters, and multi-person scenes, and is already open-sourced on GitHub.

Core Technology

Asynchronous Dual-Stream Diffusion

Hallo-Live's training has two stages:

  • Stage 1 -- Dual-Stream ODE Init: the model takes in audio-video blocks at different noise levels simultaneously, training a dual-stream DiT with unimodal and cross-modal Block-Causal Masks
  • Stage 2 -- Self-Rollout + Dual-Stream DMD: the student model autoregressively generates full audio-video from the audio-video KV Cache, then audio-video-sync-related reward distillation is introduced to distill it into a few-step model

Overall architecture

Causal Fusion Block

This is the core unit of the dual-stream DiT. The video stream and audio stream each do unimodal Block-Causal Self-Attention first, then text conditioning is injected, and they subsequently exchange information through cross-modal Block-Causal Cross-Attention — completing audio-video fusion under streaming generation.

Causal Fusion Block architecture

Future-Expanding Attention

In real speech, lip movements often arrive ahead of the sound (the coarticulation phenomenon). Strictly causal block-level attention can't see the "near-future" speech information, which leads to unnatural lip shapes.

Hallo-Live makes the video-to-audio cross-modal attention "asymmetric": video focuses on the current block, but the audio key-value range extends a small look-ahead window further forward — effectively giving the video stream a short-term "preview area."

Future-Expanding Attention mechanism

This future audio block isn't a final output but a temporary transition block, so audio quality isn't degraded.

Human-Preference-Guided Distillation

Few-step distillation speeds things up but tends to bring "averaged" degradation -- blurrier video texture, more mechanical speech, drifting audio-visual alignment. Hallo-Live introduces audio, video, and audio-video-sync-related rewards to weight the dual-stream DMD loss, maintaining generation quality while accelerating.

Application Scenarios

  • Virtual streamers / digital human livestreams: real-time text-driven, ideal for interactive livestream scenarios
  • Customer service digital humans: generate voice and video replies in real time from text input
  • Content creation: batch-generate digital human videos with speech
  • Education/training: generate digital human instructors to deliver course explanations

Technical Requirements

  • Hardware: two NVIDIA H200 GPUs (the paper's experimental environment)
  • Input: text + a reference character image
  • Output: synchronized talking video + speech
  • Latency: 0.94 seconds end to end

The project is open-sourced on GitHub, well suited for researchers and developers working on digital human interaction and real-time streaming generation.

References: