Xiaomi MiMo-V2.5-Pro-UltraSpeed: A 1T-Parameter Model Running at 1000 tokens/s
Xiaomi's flagship ultra-speed inference model breaks 1000 tokens/s of output on a standard 8-GPU commodity node, with output priced at 18 RMB per million tokens.


Xiaomi MiMo-V2.5-Pro-UltraSpeed: A 1T-Parameter Model Running at 1000 tokens/s
Xiaomi's flagship ultra-speed inference model breaks 1000 tokens/s of output on a standard 8-GPU commodity node, with output priced at 18 RMB per million tokens.
MiMo-V2.5-Pro-UltraSpeed, launched jointly by Xiaomi's MiMo team and the AI inference systems team TileRT, is a counterintuitive model. It puts a trillion-parameter (1T) MoE model on a single standard 8-GPU commodity node, breaks 1000 tokens/s of output for the first time, and peaks around 1200 tokens/s — without relying on dedicated hardware like Cerebras wafer-scale chips or Groq's custom SRAM. Within two weeks of launch it received more than 66,000 usage applications, forcing an extension of the trial window originally set to end on June 23. This article breaks down its technical approach, pricing, and how to apply, so you can judge where the speed actually comes from and whether it's worth using.
What Is MiMo-V2.5-Pro-UltraSpeed
One-line definition: Xiaomi's trillion-parameter flagship model running at 1000 tokens/s on commodity hardware.
- Architecture: MoE (Mixture of Experts)
- Total parameters: 1T (one trillion)
- Activated parameters: about 42 billion per forward pass
- Context length: supports 1M tokens
- Core breakthrough: stable 1000 tokens/s output on an 8-GPU commodity node, peaking around 1200 tokens/s
Xiaomi's MiMo announcement said applications had "far exceeded expectations," and the open window would be extended.
Speed: Where It Sits in the Industry
Placing 1000 tokens/s on the industry map makes the impact obvious. According to the AI benchmarking platform Artificial Analysis:
| Model | Output speed (approx.) |
|---|---|
| GPT-5.5 | 62-68 tokens/s |
| Claude Opus | 71 tokens/s |
| Gemini Flash | 192-200 tokens/s |
| MiMo-V2.5-Pro-UltraSpeed | 1000-1200 tokens/s |
That puts UltraSpeed's output speed an order of magnitude faster than frontier flagship models. For latency-sensitive scenarios — real-time conversation, streaming code generation, bulk long-form production — the difference in experience is a qualitative leap.
How It Does It: Joint Model-Side and System-Side Optimization
The key is that it relies on no dedicated hardware, instead squeezing performance out of commodity GPUs through joint optimization on the model side and the system side.
Model Side
- FP4 mixed quantization: FP4 is applied mainly to the MoE experts while other modules keep higher precision, cutting model size and memory-traffic pressure
- DFlash speculative decoding: block-wise masked parallel prediction replaces the token-by-token autoregression of a traditional draft model, letting the large model verify more candidate tokens at once
The open-sourced Hugging Face model card for MiMo-V2.5-Pro-FP4-DFlash is the base model behind UltraSpeed: an FP4-quantized backbone plus a BF16 DFlash drafter, under the MIT license.
System Side
- TileRT built a custom compilation engine and compute kernels for the FP4 quantization and DFlash pipeline
- A resident kernel engine and heterogeneous pipeline cooperation reduce kernel-launch and synchronization overhead
Pricing: 3x the Price for 10x the Speed
The UltraSpeed API uses a limited-time trial price set at 3x the standard MiMo-V2.5-Pro, delivering roughly a 10x boost in output speed.
| Billing item | Standard MiMo-V2.5-Pro | UltraSpeed |
|---|---|---|
| Cache-hit input | 0.025 RMB per million tokens | 3x the standard price |
| Cache-miss input | 3 RMB per million tokens | 3x the standard price |
| Output | 6 RMB per million tokens | 18 RMB per million tokens (about $2.65) |
For reference, Anthropic's flagship Claude Opus is publicly priced at $5 per million input tokens (about 34 RMB) and $25 per million output tokens (about 170 RMB). UltraSpeed's output price is only about 1.5% of Claude Opus's, while being more than ten times faster — of course, model capability can't be equated one-to-one, but that price-performance ratio is very friendly to throughput-heavy workloads.
Lei Jun posted on Weibo to announce the new MiMo-V2.5-Pro-UltraSpeed milestone.
How to Apply for Access
The model is still in a limited-time trial: users already approved can keep using it, and new users need to apply for the beta:
- API application:
https://platform.xiaomimimo.com/ultraspeed - Chat trial:
https://ultraspeed.xiaomimimo.com - Open-sourced base model:
https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash(MIT license)
As of June 23, more than 66,000 usage applications had come in, from Fortune Global 500 companies, industry leaders, and individual developers, across legal, finance, telecommunications, logistics, auto manufacturing, media, and higher education.
Use Cases
1000 tokens/s unlocks a batch of scenarios that previously "couldn't afford to wait":
- Real-time conversation and support: first-token and streaming latency near real time — a qualitative user-experience jump
- Streaming code generation: no more stutter in an coding agent's multi-round iterations
- Bulk long-form production: a major throughput lift for reports, marketing copy, and knowledge-base cleanup
- Multi-agent orchestration: agent-to-agent calls are latency-sensitive, and ultra-fast output cuts total wall time significantly
- Education and companionship: conversational, instant-feedback scenarios demand the fastest possible responses
Controversies and Limitations
- Comparability of "1T parameters": developers overseas have questioned how "trillion parameters" in a MoE architecture compares with dense models; the fact that only about 42 billion parameters activate needs to be spelled out
- Ecosystem maturity: compared with mature ecosystems like Claude and GPT, MiMo's tooling and agent-framework support are still catching up
- An objective gap in capability: speed advantages don't translate directly into reasoning advantages; for complex reasoning and long-chain agent tasks, run your own benchmarks through the Chat entry point first
💡 Tip: if your workload is throughput-heavy (support, bulk production, multi-agent orchestration), UltraSpeed's price-performance is highly attractive; if your core need is complex reasoning, run a few benchmarks through the Chat entry point before migrating.
Related articles

CAD Modeling Through Conversation: A Hands-On Guide to Zhejiang University's Open-Source CADDesigner
Type one sentence or sketch an outline and CADDesigner generates the 3D model. Zhejiang University's open-source LLM-driven CAD agent, with a workflow guide and pitfalls to avoid.

GoGo AI (Gege): Hands-On With China's First Pure-Chinese AI Music Model
Optimized for Chinese vocals, generates a full song on a single GPU in 10 seconds, and is already integrated with seven platforms including Douyin and CapCut. GoGo AI's model capabilities, access paths, and creator monetization loop.

MemSlides, Top of the HuggingFace Leaderboard: the PPT Agent That Remembers Your Preferences
Tsinghua, SJTU, and BUPT jointly open-sourced MemSlides, a memory-driven PPT generation agent supporting personalized style and multi-round local edits, topping the HuggingFace leaderboard.

ShotStream: the Open-Source Framework for Directing Multi-Shot Long Videos in Real Time (ECCV 2026)
CUHK MMLab and Kuaishou Kling jointly open-sourced ShotStream, the first real-time streaming multi-shot long-video generation framework — roughly 25x faster, with plot adjustments mid-generation.

Volcengine Seedance 2.0 API Integration in Practice: From Sign-Up to Your First Generated Video
A developer-oriented guide to the Seedance 2.0 video API: activating the Ark platform, API calls, SDK examples, TOS storage, and pricing gotchas.

Claude Artifacts Finally Gets Public Sharing + Real-Time Multiplayer Editing
Anthropic has added public link sharing and simultaneous multiplayer editing to Artifacts. This article explains what the capability is, how to use it, and how it differs from Claude Code Artifacts.