JoyAI-Echo: An Open-Source Framework for 5-Minute Long-Video Generation
JD's open-source AI long-video framework generates cross-shot audio-video up to 5 minutes in a single pass, supports local edits, and says goodbye to gacha-style rerolling.


JoyAI-Echo: An Open-Source Framework for 5-Minute Long-Video Generation
JD's open-source AI long-video framework generates cross-shot audio-video up to 5 minutes in a single pass, supports local edits, and says goodbye to gacha-style rerolling.
AI video generation has long been stuck at the "short clip" barrier. Most models on the market can only produce segments under 20 seconds — stretch to minute-level and characters change faces across shots, voices drift, and fixing a single shot means regenerating everything, keeping AI long video perpetually at the demo stage.
JoyAI-Echo, a framework JD recently open-sourced, sets out to break that bottleneck. It generates cross-shot audio-video up to 5 minutes long in a single pass, keeps character faces and speaking voices consistent throughout, supports local edits via natural language, and has published both code and weights.
What Is JoyAI-Echo
JoyAI-Echo is JD's open-source long-form audio-video generation framework. Unlike the short-video generation models common on the market, it focuses on the core problem of "long-duration consistency" — keeping the same character on the same face and the same voice across five minutes and a dozen-plus shot changes.
Both the code and the weight files are now public on GitHub, free to download and use.
- GitHub: https://github.com/jd-opensource/JoyAI-Echo
- Project homepage: https://echo-team-joy-future-academy-jd.github.io/Echo-LongVideo-Page/
Core Features
Dual Consistency Across Shots, for Both Face and Voice
The biggest pain point of traditional AI video is the "face swap." JoyAI-Echo binds each character's facial features to their voice through a "slot-paired" audio-visual memory interaction mechanism. When generating a new shot, the system retrieves the corresponding character's visual and audio tokens from the memory bank, ensuring consistency across shots.

The slot-paired audio-visual memory interaction mechanism: each historical event holds aligned visual and audio memory tokens, with paired visual and audio memory slots interacting one to one, preventing faces and voices from getting mixed up across events.
Non-Linear Editing and Local Repainting
In the past, changing one shot meant regenerating the entire video. JoyAI-Echo introduces a Director Agent, supporting local modifications directed in natural language. Unhappy with a shot? Just tell it "change the background of this chase scene to a rainy day," and the system locates that shot and repaints it, leaving the others untouched.
The Director Agent divides long-video generation into three stages — planning, generation, and review — supporting non-linear modification driven by local feedback.
High-Resolution Real-Time Super-Resolution
Through a Unified One-Step SR architecture, JoyAI-Echo supports two tiers of real-time super-resolution under streaming latency constraints, outputting HD video at up to 1472x2560 resolution directly. A single diffusion forward step scales 720p up to 2K quality.
Technical Highlights
A Million-Scale Identity-Centric Corpus
Traditional AI video training relies on flat datasets optimized for single-shot quality — the model learns how to draw a frame over short spans but never grasps the visual continuity of the same character across different times and spaces. JoyAI-Echo built an all-new Identity-Centric Video Corpus, extracting over 1 million character identity prototypes from movies, TV series, and long-form video to ensure the consistency of generated content.
Evolving Memory Bank
Rather than end-to-end generation, it uses an iterative storyboard synthesis mechanism based on an Evolving Memory Bank. During generation, target video and audio tokens are processed by two diffusion branches, while memory tokens serve only as conditional context and take no part in the loss computation.
Post-Training Pipeline
- Lip sync: long-context loss redirection with gradient amplification, scaled up 6x in the second stage, strengthening control of dialogue over mouth movements
- Visual quality: progressive resolution scheduling from 480p to 720p, balancing single-shot texture with multi-shot consistency
- RLHF alignment: the OmniNFT framework resolves bottlenecks like "audio-video reward inconsistency"
- Inference acceleration: DMD distillation compresses multi-step generation down to 8 steps, with memory-input degradation simulation added for robustness
Real-World Results
On an extra-long generation benchmark of 100 scripted stories and 3000 sequential shots, JoyAI-Echo reached a dialogue accuracy of 0.8646, with audio-visual consistency metrics ranking near the top across the board.
Best-Fit Scenarios
- Film pre-visualization: rapidly generate storyboard videos to validate narrative pacing and visual style
- Digital human content creation: conversational videos that keep characters consistent over long durations
- Story-driven short videos: generate a complete multi-shot narrative in a single pass
- Academic research: open code and weights make secondary development and algorithm improvements easy
Tip: This project requires a certain amount of GPU compute. For 2K-resolution video generation, a graphics card with 24GB or more of VRAM is recommended.
Toolin Editorial Team
Categories
Related articles

Kimi Work Beta: From an Agent That Writes Code to an Agent That Does Work
Moonshot AI launches Kimi Work Beta, a general-purpose local agent for knowledge workers supporting 300 sub-agents in parallel, 13-hour long-running tasks, browser control, and skill installation.

OpenSquilla Meta Skill: Packing an Entire Workflow into a Single Skill
OpenSquilla's new Meta Skill feature nests multiple sub-skills inside one skill, running long-horizon workflows end to end while cutting token costs by 60-80%.

Mashangfei: From One Sentence to a Complete, Running Business
Mashangfei packages the three core links of a one-person company into a closed loop: generate a complete app with an admin backend from one sentence, an AI business assistant that auto-produces posters and copy, and 7x24 AI customer service plugged into WeChat in one click.

Fully Automated AI Video Editing: A Three-Tool Stack for 100 Videos a Day
A HyperFrames + Remotion + Git three-tool stack for fully automated AI video editing, from HTML to React components, with complete install commands and pitfall notes.

A Hands-On Guide to Claude Code /workflows
A detailed look at when and how to use the Claude Code /workflows feature, using multi-agent parallelism for codebase sweeps and hard-problem research.

Replacing CleanMyMac with an Open-Source Skill
An open-source computer-cleanup skill has an agent run a read-only analysis of your Mac/Windows machine, produce a visual HTML report, and clean in tiers — it freed 120G of space in real testing.