JoyAI-Echo: JD Open-Sources a 5-Minute Long-Video Generation Framework
JD open-sources JoyAI-Echo, its first long audio-video generation framework, attacking the three big problems of character consistency, voice stability, and generation speed head-on, leading on multiple metrics.


JoyAI-Echo: JD Open-Sources a 5-Minute Long-Video Generation Framework
JD open-sources JoyAI-Echo, its first long audio-video generation framework, attacking the three big problems of character consistency, voice stability, and generation speed head-on, leading on multiple metrics.
AI long-video generation has always faced an "impossible triangle": long duration, high consistency, and fast speed — you seemingly can't have all three. The same character looks different from one shot to the next, the speaker's voice wobbles up and down, and rendering still takes half a day to wait out. JD's newly open-sourced JoyAI-Echo is here to break these pain points one by one.

What Is JoyAI-Echo
JoyAI-Echo is JD's first open-sourced long audio-video generation framework, supporting minute-level narrative video generation with character appearance and voice tone staying consistent across multiple shots. Code and model weights are fully open, so developers can build on top of them and fine-tune.
- GitHub: https://github.com/jd-opensource/JoyAI-Echo
- Hugging Face: https://huggingface.co/jdopensource/JoyAI-Echo
- Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-LongVideo-Page/
Core Technical Breakthroughs
1. Cross-Modal Audio-Video Memory Bank: Solving the "Face-Changing" Problem
Traditional models lack memory of prior content when generating shot by shot, starting over each time as if stricken with amnesia. JoyAI-Echo ships with a dedicated memory bank that continuously stores and precisely recalls characters' visual and auditory features. Across 5 minutes of multi-shot generation, this memory bank acts like the "character file" in a director's hands, guaranteeing consistent output on every call.

2. Memory-Driven Post-Training: 7.5x Speedup
JoyAI-Echo designed a three-stage post-training pipeline: SFT -> cross-modal RLHF -> distribution matching distillation (DMD). DMD compresses the multi-step diffusion teacher-student distillation into 8-step fast inference, delivering roughly a 7.5x inference speedup and turning long videos from "waiting half a day" into "footage in seconds."
3. Director Agent: Conversational Editing
You can tell it what you want in natural language — say, "swap the cafe background in scene three for a library." It automatically breaks down the request, generates the video, and checks the results. Only the unsatisfactory parts get regenerated shot by shot; the whole video doesn't have to be redone.
4. Lightweight Real-Time Super-Resolution: From 720p to HD
The bundled real-time super-resolution module upscales the native 720p video to as high as 1472x2560 resolution with almost no added latency.
Benchmark Data
In a rigorous evaluation across 100 independent story scripts totaling 3000 storyboard shots:
| Metric | Result |
|---|---|
| Speech accuracy | 0.8646 (industry-leading) |
| Audio quality preference | 81.7% |
| Prompt adherence preference | 80.6% |
| IP character consistency preference | 59.4% |
Use Cases
- Virtual anime and story creation: direct AI in natural language to generate coherent anime episodes
- Digital-human livestreams and short dramas: keep voice, lip movements, and expressions consistent over long stretches
- Brand marketing content: change a line of dialogue or a few local shots to generate multiple video versions
- Film storyboard pre-visualization: quickly generate preview videos to validate the shot language
- Educational courseware and game animation: dynamically generate coherent story animations
How to Get It
Code and model weights are fully open source; head to the GitHub repo jd-opensource/JoyAI-Echo to get them.
Toolin Editorial Team
Categories
Related articles

Gemma 4 12B: Run a Multimodal AI Model on a 16GB Laptop
Google releases a 12-billion-parameter open-source multimodal model with text, image, and audio input; it runs locally on a laptop with just 9GB of VRAM, under the Apache 2.0 license.

Hermes Desktop: The Open-Source Agent Moves onto the Desktop
Nous Research launches Hermes Desktop, an open-source desktop agent covering macOS/Windows/Linux that reuses the CLI agent's full skills and memory — usable with just mouse clicks.

Kimi Work: A Local AI Agent for Knowledge Workers
Moonshot AI launches a desktop general-purpose Agent with Agent clusters, browser control, and financial data sources, letting office workers handle their daily work with AI.

OpenClaw 2026.6.1: Native Windows Support at Last
The world's largest open-source AI Agent project ships a major update: native Windows support, a Skill Workshop that lets Agents self-evolve, and multi-Agent Workboard collaboration — 1.6 billion PCs become compute nodes.

Step 3.7 Flash: The Agent Efficiency Model at 409 tok/s
StepFun's new model hits 409 tokens/s output speed, costs 1/9 of Claude Opus 4.6 per task while matching 97% of its coding ability, and is designed for high-frequency Agent call scenarios.

OpenSquilla 3.0: Auto-Orchestrating AI Skill Flows with MetaSkill
The open-source AI Agent framework OpenSquilla 3.0 introduces MetaSkill, which synthesizes multi-step workflows from natural language, paired with intelligent model routing to cut usage costs.