
MultiTalk
Officially listedAn open-source AI framework generating audio-driven multi-person conversation videos with precise lip synchronization.
MultiTalk
MultiTalk is an open-source AI framework jointly developed by Sun Yat-sen University's Shenzhen campus, Meituan, and The Hong Kong University of Science and Technology, focused on audio-driven multi-person conversation video generation. Built on the Wan2.1-I2V-14b diffusion model, it generates multi-person conversation videos with precise lip synchronization from multi-stream audio, reference images, and text prompts, and has been accepted by NeurIPS 2025.
Key Features
- Multi-person conversation generation: Supports both single-person and multi-person scenes, binding different audio streams precisely to corresponding characters for natural multi-person interactive video
- Interactive character control: Directly control virtual characters' behavior and scene settings through natural-language prompts
- Multi-scene support: Covers conversation, singing, cartoon character animation, and other application scenarios
- High-resolution output: Delivers 480p and 720p output at any aspect ratio, with videos up to 15 seconds
- Performance optimization: Integrates TeaCache acceleration (2-3x speedup), INT8 quantization, and multi-GPU inference — a single RTX 4090 can generate 480p video
Use Cases
In content creation, creators can use MultiTalk to quickly generate dialogue videos with precise lip sync from static photos, suitable for video dubbing and character animation production. In film and game pre-production, the tool can quickly visualize dialogue scenes and multi-character interaction prototypes. In education and training, virtual instructors can be created for multilingual teaching content.
Unique Advantages
MultiTalk's proposed Label Rotary Position Embedding (L-RoPE) method effectively solves the technical challenge of binding multi-stream audio to people — an important breakthrough in this field. The project uses the Apache 2.0 open-source license, providing complete code, weights, and documentation, and supports ComfyUI integration, greatly lowering the barrier to use. Compared with similar methods, MultiTalk shows superior performance across multiple datasets (talking head, talking body, multi-person).
Pricing
完全免费使用
FAQ
What is MultiTalk?
MultiTalk is an AI tool. An open-source AI framework generating audio-driven multi-person conversation videos with precise lip synchronization.
Is MultiTalk free?
Yes, MultiTalk offers a free version.
How do I use MultiTalk?
You can use MultiTalk by visiting the official website. Click "Visit website" above to get started.