ByteDance's Lance: One 3B Model That Sees, Draws, and Edits Images and Videos End to End
ByteDance has open-sourced Lance, a natively unified multimodal model with only 3B activated parameters that covers image and video understanding, generation, and editing all at once — hitting #1 on Hugging Face Trending the moment it shipped.


ByteDance's Lance: One 3B Model That Sees, Draws, and Edits Images and Videos End to End
ByteDance has open-sourced Lance, a natively unified multimodal model with only 3B activated parameters that covers image and video understanding, generation, and editing all at once — hitting #1 on Hugging Face Trending the moment it shipped.
ByteDance's Intelligent Creation Lab has open-sourced Lance — a natively unified multimodal model with only 3B activated parameters. It packs understanding, generation, and editing for both images and videos into a single model, and hit #1 on Hugging Face Trending the moment it shipped.
In a field dominated by multimodal models that routinely run at tens or even hundreds of billions of parameters, the 3B Lance is a breath of fresh air. But it isn't grinding leaderboards on any single skill — instead, it puts "seeing, drawing, and editing" on the same exam paper and takes the whole test at once.
Core Capabilities
Lance covers 6 task types:
| Task Type | Capabilities |
|---|---|
| Image/video understanding | OCR, knowledge QA, multi-image understanding, video QA |
| Text-to-image | Image generation from complex text instructions |
| Text-to-video | Video generation with natural motion and temporal consistency |
| Image/video editing | Adding or removing subjects, local replacement, style transfer |
Video Editing: Three Consecutive Rounds of Changes
Lance doesn't just tweak a single keyframe. For example, you can chain operations together: first turn short straight hair into French curls, then add a red-and-white flower headband, and finally swap the background to a fairy-tale castle by a lakeside. Crucially, the person is still the same person, the motion stays coherent, and adjacent frames don't flicker.
Image Editing: It Understands Natural Language
Coverage includes background changes, material modifications, pose changes, portrait retouching, subject removal, replacement, and tone transfer. The core requirement is understanding natural-language instructions while preserving subject identity and visual consistency.

Technical Architecture
Lance's core approach comes down to two things:
1. Unified context Text, images, and videos all live in the same interleaved multimodal context.
2. Dual-stream decoupling The pathways for understanding and generation are split apart so they don't fight each other:
- Understanding pathway: processes text tokens and semantic visual tokens, handling QA and reasoning
- Generation pathway: processes VAE latent tokens, handling image/video generation and editing

There's also a key design called MaPE (Modality-Aware Rotary Positional Encoding), which injects modality/function-group information into the temporal dimension so the model can tell which tokens are for understanding, which are generation conditions, and which are generation targets.
Benchmark Results
| Benchmark | Score | Notes |
|---|---|---|
| VBench (video generation) | 85.11 | Leading among unified models |
| MVBench (video understanding) | 62.0 | Best among unified models, 11.3% above second place |
| GenEval (image generation) | 0.90 | Ties for the best overall score |
| GEdit-Bench (image editing) | 7.30 | Best average performance among unified models |
One interesting finding: after video generation and editing were added, video understanding didn't degrade — if anything, the multi-task data may have helped the model learn stronger cross-task transfer.
Open-Source Resources
- Paper: https://arxiv.org/abs/2605.18678
- Project homepage: https://lance-project.github.io
- GitHub code: https://github.com/bytedance/Lance
- HuggingFace model: https://huggingface.co/bytedance-research/Lance
Who It's For
- Researchers and developers who need a lightweight multimodal model
- Projects that want a unified solution across image/video understanding, generation, and editing
- Scenarios that are deployment-cost sensitive and need a 3B-class model
Tip: A 3B parameter count means it can run on consumer-grade GPUs — no A100- or H100-class hardware required. That said, its video generation and editing quality still trails top closed-source models, so it works best as a baseline and for prototype validation.