Vidu S1: The Open-Class Interactive Video Model That Lets Digital Humans Talk in Real Time
Shengshu Technology's Vidu S1 achieves real-time voice-driven video generation at 540P/25FPS, runs on consumer GPUs, and turns a single image into a conversable digital human.


Vidu S1: The Open-Class Interactive Video Model That Lets Digital Humans Talk in Real Time
Shengshu Technology's Vidu S1 achieves real-time voice-driven video generation at 540P/25FPS, runs on consumer GPUs, and turns a single image into a conversable digital human.
The video generation race is shifting from "who generates better-looking video" to "who can interact in real time." You used to type a prompt, wait tens of seconds, and receive a fixed-length clip; now Shengshu Technology's Vidu S1 pushes video generation into "video call" mode — you say a line, it generates the matching footage while listening, you can interrupt and change instructions at any time, and it can even keep the conversation going for hours without dropping. If you're working on digital humans, virtual companionship, interactive livestreaming, or AI characters, Vidu S1 is worth a spin.
What Vidu S1 is
Vidu S1 is a real-time interactive video generation model unveiled by Shengshu Technology (founded by Professor Jun Zhu's team at Tsinghua University) at the 2026 Global Digital Economy Conference. Think of it as "an AI version of a video call" — except the one sitting across from you isn't a human but a character generated frame by frame by the model.
Unlike traditional digital human solutions (model first, then train, then drive the lips), Vidu S1 takes a purely generative route: you upload a single first-frame image, the model immediately understands the character's identity, appearance, and style, then generates expressions, lip movements, gestures, and body motion in real time throughout the interaction — no offline modeling or character training stage required.
Core capabilities
Voice drives behavior, not just lip movement
Most digital human products are still stuck at the "audio drives the lips" stage — a limited set of motions with obvious splicing artifacts. Vidu S1 uses a real-time video generation architecture: it not only hears the words but understands the semantics and emotion, generating matching expressions, gestures, and complete body motion in real time.
Tell it to "give a thumbs up," "touch your nose," or "blink," and it performs the corresponding action on screen in real time. Voice here extends into a control signal for the character's behavior.

From "voice-driven lips" to "voice-driven behavior": the character understands what it hears, moves precisely, and gives more natural feedback.
Unlimited-duration real-time generation
Vidu S1 is the first to achieve unlimited-duration real-time video generation — even after hours of continuous generation, the footage stays stable without rapid drift or collapse. Over long runs, the model keeps the character's identity consistent and the motion natural and coherent, while continuing to take user instructions and respond in real time.
540P + 25FPS, runnable on consumer GPUs
In real-time interactive scenarios, resolution and frame rate directly determine whether the experience feels smooth. Vidu S1's numbers:
- Resolution: 540P (960×540)
- Frame rate: 25FPS (up to 42FPS supported)
This capability set runs on consumer-grade graphics cards. Behind it are Shengshu Technology's in-house TurboDiffusion inference acceleration framework (few-step generation, SageAttention low-bit attention, SLA/SpargeAttention sparse attention) and the TurboServe inference deployment engine (dynamic resource scheduling).

540P + 25FPS (up to 42FPS) clears the baseline for bringing real-time video generation into video calls, livestreaming, real-time companionship, interactive games, and even XR scenarios.
Custom characters: any image + any voice
Users upload an image to create a character — real people, anime, pets, game characters, and brand IPs all work as starting characters. On the voice side, you can choose a system voice or record your own for a custom one.
In the official demo, uploading an image of the Mona Lisa and entering a call made the Mona Lisa speak in response to voice input, generating lip movements, expressions, and gestures as feedback; the raised hand and the expression and tone when angry both came across fairly naturally.
How to try it
Vidu S1 is fully open to try, with three ways in:
- Web experience (mainland China): vidu.cn/vidu-stream
- API platform: platform.vidu.cn/live/landing
- Client app: search "Vidu AI Pro" in your phone's app store, then tap "Vidu S1" inside the app
Once you're on the experience page, you can start a call with a preset character right away, or upload an image to create your own character. After picking a character, issue voice commands through the microphone and the character responds on screen in real time.
💡 Tip: after uploading an image and finishing basic setup, a new character is nearly ready to enter the conversation immediately — no training wait.
The technical foundation
Vidu S1 follows the autoregressive diffusion (AR + Diffusion) route. Rather than producing a complete clip in one shot, the model predicts and generates the next frame in real time based on the footage already generated, combined with context such as the user's current voice and instructions. This frame-by-frame generation is naturally interruptible and rewriteable in real time.
Tech report: jt-zhang.github.io/files/Vidu_S1.pdf
Use cases
- Content creators: a historical portrait, an illustration, or a brand IP can quickly become a digital character that can converse and perform
- Enterprises: wire brand IPs, virtual customer service, digital anchors, game NPCs, and education companions into the business via API
- Developers: build AI characters, interactive content, and real-time video infrastructure
Limitations
The team is candid that the model's understanding of the physical world is still being refined. Over long generations, character identity and style can slowly drift, and the physical consistency of complex actions (such as how convincingly hands make contact with objects) still has room to improve. These are hurdles the whole real-time interactive video field needs to clear together.