ByteDance's Lance: One 3B Model That Sees, Draws, and Edits Images and Videos End to End

·Toolin Editorial Team

ByteDance has open-sourced Lance, a natively unified multimodal model with only 3B activated parameters that covers image and video understanding, generation, and editing all at once — hitting #1 on Hugging Face Trending the moment it shipped.

ByteDance's Lance: One 3B Model That Sees, Draws, and Edits Images and Videos End to End

ByteDance's Intelligent Creation Lab has open-sourced Lance — a natively unified multimodal model with only 3B activated parameters. It packs understanding, generation, and editing for both images and videos into a single model, and hit #1 on Hugging Face Trending the moment it shipped.

In a field dominated by multimodal models that routinely run at tens or even hundreds of billions of parameters, the 3B Lance is a breath of fresh air. But it isn't grinding leaderboards on any single skill — instead, it puts "seeing, drawing, and editing" on the same exam paper and takes the whole test at once.

Core Capabilities

Lance covers 6 task types:

Task TypeCapabilities
Image/video understandingOCR, knowledge QA, multi-image understanding, video QA
Text-to-imageImage generation from complex text instructions
Text-to-videoVideo generation with natural motion and temporal consistency
Image/video editingAdding or removing subjects, local replacement, style transfer

Video Editing: Three Consecutive Rounds of Changes

Lance doesn't just tweak a single keyframe. For example, you can chain operations together: first turn short straight hair into French curls, then add a red-and-white flower headband, and finally swap the background to a fairy-tale castle by a lakeside. Crucially, the person is still the same person, the motion stays coherent, and adjacent frames don't flicker.

Image Editing: It Understands Natural Language

Coverage includes background changes, material modifications, pose changes, portrait retouching, subject removal, replacement, and tone transfer. The core requirement is understanding natural-language instructions while preserving subject identity and visual consistency.

Lance multimodal capabilities showcase

Technical Architecture

Lance's core approach comes down to two things:

1. Unified context Text, images, and videos all live in the same interleaved multimodal context.

2. Dual-stream decoupling The pathways for understanding and generation are split apart so they don't fight each other:

  • Understanding pathway: processes text tokens and semantic visual tokens, handling QA and reasoning
  • Generation pathway: processes VAE latent tokens, handling image/video generation and editing

Lance architecture

There's also a key design called MaPE (Modality-Aware Rotary Positional Encoding), which injects modality/function-group information into the temporal dimension so the model can tell which tokens are for understanding, which are generation conditions, and which are generation targets.

Benchmark Results

BenchmarkScoreNotes
VBench (video generation)85.11Leading among unified models
MVBench (video understanding)62.0Best among unified models, 11.3% above second place
GenEval (image generation)0.90Ties for the best overall score
GEdit-Bench (image editing)7.30Best average performance among unified models

One interesting finding: after video generation and editing were added, video understanding didn't degrade — if anything, the multi-task data may have helped the model learn stronger cross-task transfer.

Open-Source Resources

Who It's For

  • Researchers and developers who need a lightweight multimodal model
  • Projects that want a unified solution across image/video understanding, generation, and editing
  • Scenarios that are deployment-cost sensitive and need a 3B-class model

Tip: A 3B parameter count means it can run on consumer-grade GPUs — no A100- or H100-class hardware required. That said, its video generation and editing quality still trails top closed-source models, so it works best as a baseline and for prototype validation.