JoyAI-Echo: An Open-Source Framework for 5-Minute Long-Video Generation

·Toolin Editorial Team

JD's open-source AI long-video framework generates cross-shot audio-video up to 5 minutes in a single pass, supports local edits, and says goodbye to gacha-style rerolling.

JoyAI-Echo: An Open-Source Framework for 5-Minute Long-Video Generation

AI video generation has long been stuck at the "short clip" barrier. Most models on the market can only produce segments under 20 seconds — stretch to minute-level and characters change faces across shots, voices drift, and fixing a single shot means regenerating everything, keeping AI long video perpetually at the demo stage.

JoyAI-Echo, a framework JD recently open-sourced, sets out to break that bottleneck. It generates cross-shot audio-video up to 5 minutes long in a single pass, keeps character faces and speaking voices consistent throughout, supports local edits via natural language, and has published both code and weights.

What Is JoyAI-Echo

JoyAI-Echo is JD's open-source long-form audio-video generation framework. Unlike the short-video generation models common on the market, it focuses on the core problem of "long-duration consistency" — keeping the same character on the same face and the same voice across five minutes and a dozen-plus shot changes.

Both the code and the weight files are now public on GitHub, free to download and use.

Core Features

Dual Consistency Across Shots, for Both Face and Voice

The biggest pain point of traditional AI video is the "face swap." JoyAI-Echo binds each character's facial features to their voice through a "slot-paired" audio-visual memory interaction mechanism. When generating a new shot, the system retrieves the corresponding character's visual and audio tokens from the memory bank, ensuring consistency across shots.

Image

The slot-paired audio-visual memory interaction mechanism: each historical event holds aligned visual and audio memory tokens, with paired visual and audio memory slots interacting one to one, preventing faces and voices from getting mixed up across events.

Non-Linear Editing and Local Repainting

In the past, changing one shot meant regenerating the entire video. JoyAI-Echo introduces a Director Agent, supporting local modifications directed in natural language. Unhappy with a shot? Just tell it "change the background of this chase scene to a rainy day," and the system locates that shot and repaints it, leaving the others untouched.

The Director Agent divides long-video generation into three stages — planning, generation, and review — supporting non-linear modification driven by local feedback.

High-Resolution Real-Time Super-Resolution

Through a Unified One-Step SR architecture, JoyAI-Echo supports two tiers of real-time super-resolution under streaming latency constraints, outputting HD video at up to 1472x2560 resolution directly. A single diffusion forward step scales 720p up to 2K quality.

Technical Highlights

A Million-Scale Identity-Centric Corpus

Traditional AI video training relies on flat datasets optimized for single-shot quality — the model learns how to draw a frame over short spans but never grasps the visual continuity of the same character across different times and spaces. JoyAI-Echo built an all-new Identity-Centric Video Corpus, extracting over 1 million character identity prototypes from movies, TV series, and long-form video to ensure the consistency of generated content.

Evolving Memory Bank

Rather than end-to-end generation, it uses an iterative storyboard synthesis mechanism based on an Evolving Memory Bank. During generation, target video and audio tokens are processed by two diffusion branches, while memory tokens serve only as conditional context and take no part in the loss computation.

Post-Training Pipeline

  • Lip sync: long-context loss redirection with gradient amplification, scaled up 6x in the second stage, strengthening control of dialogue over mouth movements
  • Visual quality: progressive resolution scheduling from 480p to 720p, balancing single-shot texture with multi-shot consistency
  • RLHF alignment: the OmniNFT framework resolves bottlenecks like "audio-video reward inconsistency"
  • Inference acceleration: DMD distillation compresses multi-step generation down to 8 steps, with memory-input degradation simulation added for robustness

Real-World Results

On an extra-long generation benchmark of 100 scripted stories and 3000 sequential shots, JoyAI-Echo reached a dialogue accuracy of 0.8646, with audio-visual consistency metrics ranking near the top across the board.

Best-Fit Scenarios

  • Film pre-visualization: rapidly generate storyboard videos to validate narrative pacing and visual style
  • Digital human content creation: conversational videos that keep characters consistent over long durations
  • Story-driven short videos: generate a complete multi-shot narrative in a single pass
  • Academic research: open code and weights make secondary development and algorithm improvements easy

Tip: This project requires a certain amount of GPU compute. For 2K-resolution video generation, a graphics card with 24GB or more of VRAM is recommended.

Related articles

Kimi Work Beta: From an Agent That Writes Code to an Agent That Does Work
AI Products

Kimi Work Beta: From an Agent That Writes Code to an Agent That Does Work

Moonshot AI launches Kimi Work Beta, a general-purpose local agent for knowledge workers supporting 300 sub-agents in parallel, 13-hour long-running tasks, browser control, and skill installation.

Toolin Editorial Team
OpenSquilla Meta Skill: Packing an Entire Workflow into a Single Skill
AI Products

OpenSquilla Meta Skill: Packing an Entire Workflow into a Single Skill

OpenSquilla's new Meta Skill feature nests multiple sub-skills inside one skill, running long-horizon workflows end to end while cutting token costs by 60-80%.

Toolin Editorial Team
Mashangfei: From One Sentence to a Complete, Running Business
AI Products

Mashangfei: From One Sentence to a Complete, Running Business

Mashangfei packages the three core links of a one-person company into a closed loop: generate a complete app with an admin backend from one sentence, an AI business assistant that auto-produces posters and copy, and 7x24 AI customer service plugged into WeChat in one click.

Toolin Editorial Team
Fully Automated AI Video Editing: A Three-Tool Stack for 100 Videos a Day
AI Tutorials

Fully Automated AI Video Editing: A Three-Tool Stack for 100 Videos a Day

A HyperFrames + Remotion + Git three-tool stack for fully automated AI video editing, from HTML to React components, with complete install commands and pitfall notes.

Toolin Editorial Team
A Hands-On Guide to Claude Code /workflows
AI Tutorials

A Hands-On Guide to Claude Code /workflows

A detailed look at when and how to use the Claude Code /workflows feature, using multi-agent parallelism for codebase sweeps and hard-problem research.

Toolin Editorial Team
Replacing CleanMyMac with an Open-Source Skill
AI Products

Replacing CleanMyMac with an Open-Source Skill

An open-source computer-cleanup skill has an agent run a read-only analysis of your Mac/Windows machine, produce a visual HTML report, and clean in tiers — it freed 120G of space in real testing.

Toolin Editorial Team