MaineCoon: 22B Parameters at 47.5 FPS, the Fastest Streaming Audio-Video Social Model Ever
The Catnip team unveils MaineCoon, a streaming audio-video social model hitting 47.5 FPS inference on 22B parameters, generating 30+ minutes of synchronized audio and video at 1/2000 the cost of Veo 3.


MaineCoon: 22B Parameters at 47.5 FPS, the Fastest Streaming Audio-Video Social Model Ever
The Catnip team unveils MaineCoon, a streaming audio-video social model hitting 47.5 FPS inference on 22B parameters, generating 30+ minutes of synchronized audio and video at 1/2000 the cost of Veo 3.
A 10-person startup team based in China, Catnip, has released MaineCoon, a streaming audio-video social model. The 22B-parameter model runs at 47.5 FPS on a single H100, produces its first frame within 1 second, and supports 30+ minutes of continuous generation. Costs are kept under $0.001 per second, dropping to just $0.00025 per second at full GPU utilization.
Core Capabilities
Streaming Audio-Video Generation
MaineCoon doesn't "generate first, play later" — it plays while generating, with audio and video output in sync. The first frame appears within 1 second of an instruction, streaming output follows, and you can feed in new instructions at any point mid-stream; the model switches over seamlessly.
The fundamental difference from existing models:
- Mainstream audio-video models: 6-7 FPS, watchable only after generation finishes, audio and video separate
- MaineCoon: 47.5 FPS, sub-second first frame, synchronized audio and video, switchable in real time
Unlimited-Duration Generation
An industry first: 30+ minutes of continuous audio-video generation while maintaining image quality, consistency, and audio-video sync without collapse. The official demo shows a 2-minute continuous generation video in which the characters show no obvious distortion by the end.
Real-Time Social Interaction
MaineCoon is the first to land this vertically in social interaction, with the core being a sense of "aliveness":
- Natural character details like eye movement, subtle lip motion, and speech rhythm
- Highly synchronized audio and video
- Instructions can be switched at any point during generation
- It picks up on what users say and responds with emotional feedback
Real-World Performance
Speed Advantage
| Model Type | Parameters | FPS |
|---|---|---|
| MaineCoon | 22B | 47.5 |
| Lightweight streaming video model | 1.3B | 19.1 |
| Comparable streaming audio-video model | - | 6-7 |
MaineCoon is a full 7 times faster than comparable streaming audio-video models, and even on an RTX Pro 6000 — half the cost of an H100 — it holds a steady 30 FPS or above.
Cost Comparison
- Inference cost per second: $0.001 (typical), $0.00025 (GPU at full load)
- vs. Veo 3: 1/2000 of the cost
- vs. Seedance: 1/560 of the cost
Benchmark Performance
Catnip built the first social short-video-specific benchmark, SocialVideo Bench, covering seven scenarios: dense speech, two-person interaction, musical performance, emotional acting, dance, creative challenges, and social memes.
MaineCoon scored 0.934 overall, surpassing 7 mainstream audio-video generation models and setting a new SOTA (the previous best baseline, SoulX-FlashTalk, scored 0.895).
Technical Architecture
Three-Stage Training
- Self-Resampling: exposing the model to degraded historical frames during training so it learns to stay stable under slight drift and noise, bridging the gap between training and inference
- Streaming representation alignment: introducing a frozen pretrained V-JEPA 2 vision encoder for distillation supervision, accelerating cross-modal semantic structure learning
- Domain-aware preference optimization + reinforced online policy distillation: training specialized preference expert models for different social scenarios (dance needs dynamics, conversation needs lip sync, wide shots need body structure), then unifying them into a deployable streaming policy
Three-Agent Inference Framework
On the inference side, three independent intelligent controllers work together:
- Director: handles narrative planning and quality correction, continuously monitoring generated content through an observer and repairing drift forward the moment it appears
- Cache Manager: manages retention and eviction of the KV cache, anchoring character appearance and key frames as long-term memory and periodically correcting global appearance drift
- Buffer Controller: balances real-time performance and interactive responsiveness, keeping ahead-of-time generation within a reasonable window
Engineering Optimization
The 22B model was trained on 64 H100s using only 10k GPU hours, with less than 1 million samples of data. By precomputing and storing video encodings, text embeddings, and teacher features, the GPU only does the most essential computation.
Use Cases
- Virtual social interaction: 1-on-1 video conversations with AI characters responding in real time
- Content creation: real-time streaming generation of short videos, adjusting as you create
- Virtual streamers/companions: long-duration stable operation with synchronized audio and video that never drops
- Multi-style content: both realistic style and animation styles (like Minecraft characters) are supported
How to Try It
The official team has opened a limited batch of 200 invite codes:
- Website: https://mainecoon.tech/
- Technical report: https://arxiv.org/abs/2606.17800
- Model blog: https://mainecoon.tech/blogs
About the Team
Catnip was founded a bit over half a year ago and has roughly 10 people. Founder Yang Shurui previously worked in product at TikTok and PixVerse, where she drove several hit template effects from zero to one. Chief scientist Xie Zeke is an assistant professor at HKUST (Guangzhou) and holds a doctorate from the University of Tokyo. The team has raised a seed round from top VC firms including Sequoia and Mingshi Capital.
The MaineCoon project kicked off in March 2026: 3 core researchers delivered the full stack — model training, training infrastructure, data foundations, and the inference system — in 2 months.