ShotStream: the Open-Source Framework for Directing Multi-Shot Long Videos in Real Time (ECCV 2026)
CUHK MMLab and Kuaishou Kling jointly open-sourced ShotStream, the first real-time streaming multi-shot long-video generation framework — roughly 25x faster, with plot adjustments mid-generation.


ShotStream: the Open-Source Framework for Directing Multi-Shot Long Videos in Real Time (ECCV 2026)
CUHK MMLab and Kuaishou Kling jointly open-sourced ShotStream, the first real-time streaming multi-shot long-video generation framework — roughly 25x faster, with plot adjustments mid-generation.
Getting AI to generate a coherent multi-shot long video used to work like this: stuff every shot's prompt into the model at once, wait about half an hour for the output, and have zero ability to adjust along the way. ShotStream breaks that deadlock — it is the first real-time streaming multi-shot long-video generation framework, released jointly by the Chinese University of Hong Kong's MMLab and Kuaishou's Kling team, accepted by ECCV 2026, with the model, training code, and test code all open sourced, and roughly a 25x speedup over traditional methods.
What ShotStream Is
The core problem ShotStream solves is the "high latency + zero interactivity" trap of multi-shot long-video generation.
The traditional workflow (based on bidirectional diffusion architectures):
- You must write all the shot prompts up front
- The model processes every shot together to keep them coherent
- You wait about half an hour for the full video
- Want to change the plot midway? Run it all again
ShotStream takes a completely different approach — it models multi-shot synthesis as a history-based autoregressive process, generating each shot in a stream conditioned on the shots already generated, which enables:
- Adjusting the plot while generation is in progress
- Real-time interactive creation
- Roughly a 25x speedup

Autoregressive streaming = every new shot is generated "looking at" the shots before it, which both preserves coherence and lets you inject new plot directions at any moment.
Core Capabilities
Real-time streaming generation
No more waiting half an hour. ShotStream works like a live stream, outputting shot by shot:
- The moment shot 1 finishes generating, you can watch it
- After you've seen shot 1, you can rewrite the prompt for shot 2
- Shot 2 is generated based on shot 1 and your new instructions
For creators, this means you can genuinely "direct" an AI video instead of writing all the prompts and praying for a one-shot miracle.
Multi-shot consistency
The biggest technical challenge in multi-shot long video is cross-shot consistency — a character wearing red in shot 1 can't turn up in blue in shot 2, and the light direction in a scene has to stay consistent too.
ShotStream solves this through autoregressive history modeling: every new shot is conditioned on the previous shots, so it naturally inherits character appearance, scene setup, and camera language.
25x speedup
Compared with traditional one-shot multi-shot video models, ShotStream is roughly 25x faster. The core reasons:
- Autoregressive streaming doesn't process all shots at once
- Single-shot generation can be fully parallelized and optimized
- History is compressed into a compact representation that doesn't need recomputing

How It Works (Briefly)
ShotStream's key innovations:
- Streaming autoregressive modeling: changing multi-shot synthesis from "one-shot bidirectional" to "history-based autoregressive"
- History compression: compressing information from already-generated shots into a compact context that conditions the next shot
- Real-time instruction injection: each shot can accept new user instructions before it is generated, enabling "adjust as you generate"
This architecture is a natural fit for interactive storytelling — every viewer choice can instantly shape the next shot.
How to Use It
Preparation
- Open-source repository: search
ShotStreamor visit the official GitHub of CUHK MMLab / Kuaishou Kling - Paper: accepted at ECCV 2026; search the paper title
- Hardware: GPU inference recommended (see the repository README for exact specs)
Typical Flow
Step 1: Define the opening shot
Enter a description of the first shot, for example:
"Shot 1: overhead view, a city rooftop at night, the protagonist standing at the edge in a trench coat, the wind catching the hem."
ShotStream generates shot 1 and displays it.
Step 2: Append shots in a stream
After watching shot 1, enter the next instruction:
"Shot 2: cut to a front close-up, the protagonist turns toward the camera with a determined expression."
ShotStream generates shot 2 based on shot 1, keeping the character's appearance and the night lighting consistent.
Step 3: Adjust the plot in real time
If shot 2 isn't right, or you want the story to go another way:
"Change shot 2 to: the protagonist looks up at the sky as a drone flies over."
The model regenerates shot 2 based on shot 1 and the new instruction.
Step 4: Export the long video
Once all shots are generated, ShotStream automatically stitches them into a coherent multi-shot long video and exports it as an mp4.
Tip: while streaming, preview each shot as soon as it lands and change it immediately if needed — that's faster than waiting until everything is generated.
Best-Fit Scenarios
| Scenario | Recommended? | Why |
|---|---|---|
| Interactive storytelling / interactive video | Strongly recommended | Real-time adjustment is the headline feature |
| Short-video / ad multi-shot scripts | Strongly recommended | Fast + coherent across shots |
| Film pre-production storyboard previews | Recommended | Directors can adjust shots in real time |
| Single-shot short videos | Average | Little advantage for single shots |
| Live real-time generation | With caution | Still some latency after the 25x speedup; needs further optimization |
Comparison With Peer Approaches
| Approach | Real-time interaction | Multi-shot consistency | Open source | Best for |
|---|---|---|---|---|
| ShotStream | Yes | Strong (autoregressive) | Yes | Interactive / streaming creation |
| Traditional multi-shot models | No | Strong (bidirectional) | Partial | One-shot final renders |
| Single-shot generation (e.g. Kling, Sora) | N/A | Requires manual stitching | No | Single-shot scenarios |
Common Issues
- What is the real-time generation latency: check the repository's benchmarks for single-shot latency; overall it is about 25x faster than traditional methods.
- Can it be used in commercial projects: the model is open source; specific commercial terms are governed by the repository LICENSE.
- How long can the video be: autoregression can in theory continue indefinitely; in practice VRAM and consistency drift set the limit, so keep it under 10 shots per run.
- Character consistency occasionally drifts: characters may shift slightly late in a long video; consider inserting a "character anchor" prompt every 5 shots.
Related Links
- Sohu coverage: https://m.sohu.com/a/1049298810_129720
- ECCV 2026 paper (search ShotStream)
- CUHK MMLab / Kuaishou Kling GitHub