MaineCoon: The Fastest Streaming Audio-Video Social Model Yet
Catnip has unveiled MaineCoon, a 22B-parameter streaming audio-video model that hits 47.5 FPS on a single H100, runs at 1/2000th the cost of Veo 3, and supports 30+ minutes of synchronized audio-video output.


MaineCoon: The Fastest Streaming Audio-Video Social Model Yet
Catnip has unveiled MaineCoon, a 22B-parameter streaming audio-video model that hits 47.5 FPS on a single H100, runs at 1/2000th the cost of Veo 3, and supports 30+ minutes of synchronized audio-video output.
Existing audio-video generation models share two stubborn flaws: they are either too slow, forcing you to wait for the full render before seeing any result, or they handle video but drop the audio, so sound and picture always travel on separate tracks. MaineCoon, from the Catnip team, tries to solve both problems at once — a 22B-parameter model that runs at 47.5 FPS on a single H100, delivers its first frame within 1 second of a command, supports 30+ minutes of synchronized audio and video, and keeps per-second cost under $0.001.
What Is MaineCoon
MaineCoon is a streaming audio-video social model developed by the Catnip team. The name comes from the Maine Coon cat breed — nicknamed the "dog cat" because it follows you almost everywhere you go. The model behaves the same way: instead of generating everything in one go and stopping, it keeps tracking your state and continues in real time.
Give it a piece of text and it generates and plays back at the same time, with audio and video produced together, much like a 1V1 video call with a real person. Sessions can run past 30 minutes — a first for the industry.
Core Features
Streaming Audio-Video Generation
Streaming generation is not a new concept (ChatGPT emitting text one word at a time is streaming output), but a single video frame involves thousands of pixels that must stay precisely aligned with audio on the timeline — a challenge on an entirely different scale. The smaller the generated segment, the shorter the historical context each frame can rely on, and the more easily the model gives itself away.
MaineCoon compresses this unit down to the sub-second level: the first frame appears within 1 second of a command, pursuing low latency and high quality at the same time. Feed in new instructions mid-stream and the model adjusts on the fly.
The Fastest Inference Speed in the Industry
Comparable streaming audio-video models generally run at 6-7 FPS; MaineCoon is a full 7x faster.
- 22B parameters: deploys on a single H100 at up to 47.5 FPS
- RTX Pro 6000 (an inference card at half the cost of an H100): a stable 30+ FPS
- More than 2x faster than even a 1.3B lightweight streaming video model (19.1 FPS)
The speed comes with no sacrifice in quality — emotional expression is actually richer, and motion more coherent and stable.
Unlimited-Duration Generation
MaineCoon can continuously generate 10+ minutes of audio-video content while keeping image quality, consistency, and audio-video sync intact. The architecture is even fully capable of unbounded generation.
On the self-built SocialVideo Bench benchmark (covering seven scenarios: dense speeches, two-person interactions, musical performances, emotional acting, dance, creative challenges, and social memes), MaineCoon scored 0.934 overall, beating seven mainstream audio-video generation models and setting a new SOTA.
Cost Comparison
| Model | Inference Cost per Second | Compared with MaineCoon |
|---|---|---|
| MaineCoon (GPU fully utilized) | $0.00025 | Baseline |
| MaineCoon (standard) | <$0.001 | Baseline |
| Veo 3 | ~$0.5 | MaineCoon is 1/2000 of this |
| Seedance | ~$0.14 | MaineCoon is 1/560 of this |
Technical Architecture
Training: Three Progressive Stages
- Self-resampling: exposing the model to degraded historical frames during training so it learns to stay stable under imperfect conditions, closing the gap between training and inference
- Streaming representation alignment: introducing a frozen, pretrained V-JEPA 2 vision encoder for distillation supervision to speed up training convergence
- Domain-aware preference optimization (DPO) + reinforced online policy distillation (ROPD): training separate preference expert models for different social scenarios such as dance, conversation, and wide shots, then unifying them into a single deployable streaming policy
The 22B model finished training within 10k GPU hours, on fewer than 1 million samples.
Inference: Three Intelligent Controllers

- Director: the cognitive core, responsible for narrative and error correction. It generates structured prompts beat by beat, continuously monitors quality drift, and kicks off forward correction the moment something goes wrong
- Cache Manager: manages what stays in and what gets evicted from the KV cache, keeping character appearances, scene-establishing frames, and key dialogue frames as long-term memory anchors
- Buffer Controller: balances real-time performance with interactive responsiveness, controlling how far ahead the buffer runs so playback never stutters and user instructions take effect promptly
Use Cases
MaineCoon is the first model to land vertically in social interaction, targeting a sense of "aliveness" — the details that determine realism, like shifting gaze, lip twitches, and speaking rhythm.
- AI social conversation: simulating character dialogue where facial muscle movement and speech pauses follow instructions precisely
- Real-time interactive content: ask the character questions at any time, or feed in new instructions to adjust its behavior
- Long-video generation: sustaining quality and consistency across 10+ minutes of continuous content
💡 Tip: MaineCoon has opened a limited internal beta with 200 invitation codes. The team's next goal is a "social world model" — one that treats the human as the center of its coordinate system, actively reads user emotions, and simulates the course of social behavior from a human origin point. You can apply for the beta at mainecoon.tech.
Related articles

AnySearch: Turning Search into Infrastructure for Agents
AnySearch, which topped the Product Hunt weekly chart, isn't a search box — it's a unified API that feeds AI agents pre-filtered, deduplicated, structured information. This article breaks down its positioning, integration options, and why it saves tokens.

Claude Sonnet 5 Arrives: Near-Opus 4.8 Performance at 60% of the Price
Anthropic releases Claude Sonnet 5 with adaptive thinking on by default and a new tokenizer. Pricing stays at $3/$15, with a limited-time $2/$10 rate through August 31.

Mandol in Practice: Building Agent Long-Term Memory with Zero LLM Calls
The Institute of Software, Chinese Academy of Sciences, open-sources Mandol, which unifies KV/vector/graph storage via SemanticMap + SemanticGraph — zero LLM calls at retrieval, 5.4x faster retrieval, and best-in-class results on both LoCoMo and LongMemEval.

MCP Adds Enterprise-Managed Authorization: One Login, Every Connector Ready
Model Context Protocol ships the Enterprise-Managed Authorization extension: enterprises can centrally govern MCP Server access through IdPs like Okta, and users connect on first login with zero configuration.

MobileForge in Practice: Putting GUI Agents on an Unlabeled Data Flywheel
Kuaishou, together with Zhejiang University, open-sources MobileForge: MobileGym + HiFPO let mobile GUI agents self-explore, self-feed-back, and self-improve inside real apps, with full code and models released.

StepFun's Step Edge On-Device Family: 0.1-Second Local Agent Tool Calls
StepFun releases the Step Edge on-device model family — base/Audio/GUI/Gen — aimed at phones and cars, with local tool calls as fast as 0.1 seconds and sensitive data never leaving the device.