Xiaomi MiMo UltraSpeed: A Trillion-Parameter Model Running at 1000 tokens/s

·Toolin Editorial Team

Xiaomi's MiMo-V2.5-Pro UltraSpeed delivers 1000 tokens/s on a trillion-parameter model running on generic 8-GPU hardware — and it's the full-strength Pro version, not a dumbed-down Flash variant. The API is live and taking applications.

Xiaomi MiMo UltraSpeed: A Trillion-Parameter Model Running at 1000 tokens/s

Xiaomi's MiMo team and the inference systems team TileRT jointly announced that MiMo-V2.5-Pro's UltraSpeed mode has pushed a trillion-parameter (1T) flagship model past the 1000 tokens/s output mark for the first time. More importantly, this isn't a dumbed-down Flash variant — it's the full-strength Pro version.

That means roughly a 10x speedup while maintaining top-tier intelligence.

Real-World Speed

Take a complex visual dashboard generation task as an example:

  • UltraSpeed version: done in 13 seconds
  • Standard version: 6 minutes 15 seconds
  • Up to a 28x speedup for the same quality of output

In hands-on testing, peak speed even reached 1426 tokens/s, outputting 25624 tokens in 32 seconds and generating 1000 lines of code. A Snake game takes 10 seconds to generate, and a macOS system UI can be reproduced in 1 minute.

How It Works

Unlike dedicated-hardware approaches such as Cerebras' wafer-scale integration or Groq's custom pure on-chip SRAM chips, Xiaomi pulled this off on generic GPUs — a single standard 8-GPU node.

The core technology has three parts:

1. FP4 quantization: a dramatic slim-down without losing accuracy

  • Only the MoE experts get FP4 quantization; other modules keep their original precision
  • Through FP4 quantization-aware training (QAT), the model's overall capability stays essentially on par with the original
  • Model size and memory-access overhead are both dramatically reduced

2. DFlash speculative decoding: confirming multiple text segments in one go

  • Uses a block-level masked parallel prediction method, removing the serial constraint of autoregressive drafting
  • In coding scenarios the average acceptance length reaches 6.30, with 6-7 of every 8 drafted tokens accepted per verification round
  • The draft model uses sliding-window attention (SWA), making per-prediction compute constant

3. TileRT custom compiled kernels

  • Resident kernel engine: the compute pipeline lives inside the GPU and keeps flowing continuously
  • Heterogeneous pipeline collaboration: communication, data movement, and tensor compute are finely decomposed at the tile level
  • Microsecond-level hardware-software convergence: purpose-built for FP4 mixed quantization and DFlash

Pricing and API

The API is live with a limited-time trial price:

  • Priced at 3x MiMo-V2.5-Pro, in exchange for roughly a 10x output speedup
  • Based on MiMo-V2.5-Pro pricing, that works out to roughly ¥18 per million output tokens
  • API trial only; Token Plan not supported yet
  • Applications are open June 9 through June 23; approved users get two weeks of free limited-time Chat access

Why Speed Matters in Practice

The speedup isn't just "faster" — it unlocks entirely new usage patterns:

  • Agent scenarios: if you estimate a task takes one minute, you'll watch it through to the end. If it takes five minutes, you'll probably go do something else and waste time coming back. A 10x speedup directly changes your working rhythm
  • Sub-Agent concurrency: spin up one or two hundred sub-agents at once; with a 10x speedup and zero loss in model capability, the difference in experience is dramatic
  • Real-time decision loops: trillion-parameter models can plug into time-critical scenarios like high-frequency quantitative trading signal generation and instant fraud-detection interception

How to Apply

Tip: The current high acceptance rates are still concentrated mainly in structured tasks like coding; general conversation still has room for optimization. Inference capacity is tight, so broad commercial availability will take time.

Related articles

Claude Code Verification Loops: 4 Built-in Skills That Make AI Self-Check Code Before Delivery
AI Tutorials

Claude Code Verification Loops: 4 Built-in Skills That Make AI Self-Check Code Before Delivery

Build a verification layer with Claude Code's four built-in skills — /code-review, /simplify, /verify, and /design — so the AI self-checks its code before handing it over.

Toolin Editorial Team
Logi Options+: Turn an Old Mouse into an AI Workflow Controller (Zero-Cost Setup)
AI Tutorials

Logi Options+: Turn an Old Mouse into an AI Workflow Controller (Zero-Cost Setup)

Use Logitech's free official driver Logi Options+ — Smart Actions, Actions Ring, and per-app settings — to mod any Logitech mouse into a Vibe Mouse, a DIY alternative to the $230 Codex Micro keyboard.

Toolin Editorial Team
Loom: An Engineering State Layer for Coding Agents That Cures Long-Task Local Amnesia
AI Products

Loom: An Engineering State Layer for Coding Agents That Cures Long-Task Local Amnesia

The open-source tool Loom gives coding agents like Claude Code / Codex an independent, structured engineering state layer, enabling long-task auto-checkpointing and zero-cost multi-agent handoff.

Toolin Editorial Team
Jetson-PI: Peking University Open-Sources Real-Time On-Device VLA Control, 8.66× Higher Control Frequency on Jetson Orin
AI Products

Jetson-PI: Peking University Open-Sources Real-Time On-Device VLA Control, 8.66× Higher Control Frequency on Jetson Orin

Peking University, AIRS, and PrimeBot open-source Jetson-PI (Apache-2.0): FAAC asynchronous inference + confidence scheduling + a llama.cpp engine lift π0.5 on Jetson Orin from 0.70Hz to 6.06Hz, with no loss in accuracy.

Toolin Editorial Team
JiuwenSwarm: A Huawei-Backed Open-Source Multi-Agent Unified Workbench Where Humans Join the Agent Team
AI Products

JiuwenSwarm: A Huawei-Backed Open-Source Multi-Agent Unified Workbench Where Humans Join the Agent Team

openJiuwen open-sources JiuwenSwarm (Apache-2.0), merging office / Code / entertainment work into a unified workbench and launching HITS, a new human-agent collaboration paradigm — installable with a single pip line.

Toolin Editorial Team
OpenWorker: Andrew Ng's Open-Source Desktop AI Coworker That Ships Finished Work
AI Products

OpenWorker: Andrew Ng's Open-Source Desktop AI Coworker That Ships Finished Work

Andrew Ng open-sources OpenWorker (MIT), a local-first, model-agnostic desktop agent connecting to 25+ tools, turning "prepare a customer brief" into a finished document you can actually open.

Toolin Editorial Team