RhymeFlow: Open-Sourcing a 1.8x Speedup for Video Generation
Tsinghua University has open-sourced RhymeFlow, a video generation acceleration framework that speeds inference on DiT models like Wan 2.1 and CogVideoX by 1.5x-1.8x with no retraining required, at near-lossless quality — 62.5% of users couldn't tell the difference.


RhymeFlow: Open-Sourcing a 1.8x Speedup for Video Generation
Tsinghua University has open-sourced RhymeFlow, a video generation acceleration framework that speeds inference on DiT models like Wan 2.1 and CogVideoX by 1.5x-1.8x with no retraining required, at near-lossless quality — 62.5% of users couldn't tell the difference.
Tsinghua University and GigaAI have jointly open-sourced RhymeFlow, a completely training-free video generation acceleration framework. It requires no model retraining and applies "inter-frame asynchronous scheduling" directly at inference time, boosting inference speed on DiT video models like Wan 2.1 and CogVideoX by 1.5x to 1.8x.
In a double-blind user study with 82 participants, 62.5% could not distinguish RhymeFlow's output from the original model's.
The Problem It Solves
Today's mainstream DiT video models (Wan 2.1, CogVideoX, Sora) share one pain point: generating a single 81-frame 720p clip takes nearly 17 minutes on a single A800 GPU.
Existing acceleration methods (sparse attention, KV caching, quantization) optimize the compute within each step. But nobody has touched a more fundamental issue — every frame is treated equally, so even adjacent frames with nearly identical content still run through the full 50-step denoising process.
RhymeFlow's core insight: a video's semantics and motion are continuous, keyframes determine the global structure, and non-keyframe trajectories are highly predictable. Given that, why not let different frames take different paths?

Three Core Modules
1. Content-aware keyframe selection
Rather than naive uniform sampling, it uses latent-space semantic similarity to automatically identify keyframes containing scene cuts or abrupt object motion. These frames get full compute resources, ensuring the video's structural integrity and semantic accuracy.
2. Progressive asynchronous denoising schedule
Keyframes update at every step, while non-keyframes skip steps on a schedule that varies by noise stage:
- Warm-up stage (first 15 steps): all frames denoise in sync to lay the foundation of the global composition
- High-noise stage (structure-sensitive): non-keyframes update every 2 steps
- Low-noise stage (detail refinement): non-keyframes update every 3 steps
- Sync points: all frames periodically reconverge to calibrate non-keyframe trajectories and prevent error accumulation
3. Latent trajectory projection
Once non-keyframes skip steps, the missing intermediate states break the temporal consistency of 3D attention. RhymeFlow uses a linear projection module with negligible compute cost to precisely predict the intermediate latents from the two known states before and after — effectively drawing a smooth motion trajectory for each non-keyframe.

Benchmark Results
Results on mainstream open-source models:
Against SOTA methods:
- On Wan 2.1: RhymeFlow's PSNR is 1.84 higher than SAP's and its SSIM 0.053 higher, at comparable speed
- On CogVideoX: a 1.78x speedup while retaining 98.6% subject consistency
- Stacked with SAP: the speedup rises further to 1.93x, with better quality than SAP alone

82-person double-blind user study:
- 53.7% of users rated RhymeFlow's temporal coherence above SVG
- 74.4% of users preferred RhymeFlow over SAP
- Against the original model, 62.5% of users could not tell the difference — no statistically significant gap
Who It's For
RhymeFlow fits these scenarios:
- Batch video generation with open-source DiT video models like Wan 2.1 and CogVideoX
- Cutting video generation inference costs (GPU time roughly halved)
- Quality-sensitive workloads that can tolerate an extremely slight quality loss
Where to Get It
- Paper: arxiv.org/abs/2604.08370
- GitHub: github.com/Simon-Dcs/RhymeFlow
- Project page: simon-dcs.github.io/Website-of-RhymeFlow
The framework is fully open source, requires no model retraining, and plugs directly into existing DiT inference pipelines.