JoyAI-Echo: JD Open-Sources a 5-Minute Long-Video Generation Framework

·Toolin Editorial Team

JD open-sources JoyAI-Echo, its first long audio-video generation framework, attacking the three big problems of character consistency, voice stability, and generation speed head-on, leading on multiple metrics.

JoyAI-Echo: JD Open-Sources a 5-Minute Long-Video Generation Framework

AI long-video generation has always faced an "impossible triangle": long duration, high consistency, and fast speed — you seemingly can't have all three. The same character looks different from one shot to the next, the speaker's voice wobbles up and down, and rendering still takes half a day to wait out. JD's newly open-sourced JoyAI-Echo is here to break these pain points one by one.

Screenshot of JoyAI-Echo's Hugging Face page

What Is JoyAI-Echo

JoyAI-Echo is JD's first open-sourced long audio-video generation framework, supporting minute-level narrative video generation with character appearance and voice tone staying consistent across multiple shots. Code and model weights are fully open, so developers can build on top of them and fine-tune.

Core Technical Breakthroughs

1. Cross-Modal Audio-Video Memory Bank: Solving the "Face-Changing" Problem

Traditional models lack memory of prior content when generating shot by shot, starting over each time as if stricken with amnesia. JoyAI-Echo ships with a dedicated memory bank that continuously stores and precisely recalls characters' visual and auditory features. Across 5 minutes of multi-shot generation, this memory bank acts like the "character file" in a director's hands, guaranteeing consistent output on every call.

Cross-modal audio-video memory bank mechanism

2. Memory-Driven Post-Training: 7.5x Speedup

JoyAI-Echo designed a three-stage post-training pipeline: SFT -> cross-modal RLHF -> distribution matching distillation (DMD). DMD compresses the multi-step diffusion teacher-student distillation into 8-step fast inference, delivering roughly a 7.5x inference speedup and turning long videos from "waiting half a day" into "footage in seconds."

3. Director Agent: Conversational Editing

You can tell it what you want in natural language — say, "swap the cafe background in scene three for a library." It automatically breaks down the request, generates the video, and checks the results. Only the unsatisfactory parts get regenerated shot by shot; the whole video doesn't have to be redone.

4. Lightweight Real-Time Super-Resolution: From 720p to HD

The bundled real-time super-resolution module upscales the native 720p video to as high as 1472x2560 resolution with almost no added latency.

Benchmark Data

In a rigorous evaluation across 100 independent story scripts totaling 3000 storyboard shots:

MetricResult
Speech accuracy0.8646 (industry-leading)
Audio quality preference81.7%
Prompt adherence preference80.6%
IP character consistency preference59.4%

Use Cases

  • Virtual anime and story creation: direct AI in natural language to generate coherent anime episodes
  • Digital-human livestreams and short dramas: keep voice, lip movements, and expressions consistent over long stretches
  • Brand marketing content: change a line of dialogue or a few local shots to generate multiple video versions
  • Film storyboard pre-visualization: quickly generate preview videos to validate the shot language
  • Educational courseware and game animation: dynamically generate coherent story animations

How to Get It

Code and model weights are fully open source; head to the GitHub repo jd-opensource/JoyAI-Echo to get them.

Related articles

Gemma 4 12B: Run a Multimodal AI Model on a 16GB Laptop
AI Products

Gemma 4 12B: Run a Multimodal AI Model on a 16GB Laptop

Google releases a 12-billion-parameter open-source multimodal model with text, image, and audio input; it runs locally on a laptop with just 9GB of VRAM, under the Apache 2.0 license.

Toolin Editorial Team
Hermes Desktop: The Open-Source Agent Moves onto the Desktop
AI Products

Hermes Desktop: The Open-Source Agent Moves onto the Desktop

Nous Research launches Hermes Desktop, an open-source desktop agent covering macOS/Windows/Linux that reuses the CLI agent's full skills and memory — usable with just mouse clicks.

Toolin Editorial Team
Kimi Work: A Local AI Agent for Knowledge Workers
AI Products

Kimi Work: A Local AI Agent for Knowledge Workers

Moonshot AI launches a desktop general-purpose Agent with Agent clusters, browser control, and financial data sources, letting office workers handle their daily work with AI.

Toolin Editorial Team
OpenClaw 2026.6.1: Native Windows Support at Last
AI Products

OpenClaw 2026.6.1: Native Windows Support at Last

The world's largest open-source AI Agent project ships a major update: native Windows support, a Skill Workshop that lets Agents self-evolve, and multi-Agent Workboard collaboration — 1.6 billion PCs become compute nodes.

Toolin Editorial Team
Step 3.7 Flash: The Agent Efficiency Model at 409 tok/s
AI Products

Step 3.7 Flash: The Agent Efficiency Model at 409 tok/s

StepFun's new model hits 409 tokens/s output speed, costs 1/9 of Claude Opus 4.6 per task while matching 97% of its coding ability, and is designed for high-frequency Agent call scenarios.

Toolin Editorial Team
OpenSquilla 3.0: Auto-Orchestrating AI Skill Flows with MetaSkill
AI Products

OpenSquilla 3.0: Auto-Orchestrating AI Skill Flows with MetaSkill

The open-source AI Agent framework OpenSquilla 3.0 introduces MetaSkill, which synthesizes multi-step workflows from natural language, paired with intelligent model routing to cut usage costs.

Toolin Editorial Team