MiniMax M3 Released: Coding on Par with Opus, 1M Context + Native Multimodality

·Toolin Editorial Team

MiniMax M3 ships with the new sparse attention architecture MSA, supports a 1M context window, and beats GPT-5.5 on Coding — the first model in China to combine all three frontier capabilities

MiniMax M3 Released: Coding on Par with Opus, 1M Context + Native Multimodality

MiniMax M3 is officially released today. This model reaches frontier level on coding and Agent tasks, ships with the brand-new sparse attention architecture MSA (MiniMax Sparse Attention), supports up to 1M of ultra-long context, and is a natively multimodal model that supports image and video input and can operate the computer desktop.

M3 is the first model in China to combine all three of frontier Coding capability, ultra-long context, and native multimodality.

MiniMax M3 performance comparison

MiniMax M3's results across multiple benchmarks

Core Capabilities at a Glance

Coding and Agent Capability

On internationally authoritative evaluations of Coding capability, M3 reaches world-leading level:

  • SWE-Bench Pro: 59.0% (surpassing GPT-5.5 and Gemini 3.1 Pro, close to Opus 4.7)
  • Terminal Bench 2.1: 66.0%
  • SWE-fficiency: 34.8%
  • KernelBench Hard: 28.8%
  • MCP Atlas: 74.2%
  • SVG-Bench: surpasses Opus 4.7

M3 also scored highest on Claw-Eval, the Agent end-to-end evaluation framework.

The New Attention Architecture MSA

MSA is MiniMax's in-house sparse attention architecture, solving the problem of quadratic growth in computational complexity in full attention. Its core advantages:

  • Precise KV chunking: More precise than DSA and MoBA, achieving higher effective context coverage
  • Compute and memory-access optimization: Uses the KV outer gather Q approach, reading each block only once with contiguous memory access
  • Significant speedup: More than 4x faster than the open-source Flash-Sparse-Attention and FlashMoBA

MSA architecture comparison

Speed comparison of MSA against other sparse attention solutions

At 1 million context, M3's per-token compute is only 1/20 of the previous-generation model. Prefilling is sped up by more than 9x, and Decoding by more than 15x.

Native Multimodality

M3 performs multimodal mixed training from Step 0 — instead of attaching a vision encoder after the fact, the semantic spaces of different modalities of data fuse naturally. Interleaved data in the training corpus proved especially critical for the performance gains, and the training token scale has been raised to the order of 100 trillion.

Hands-On: Reproducing an Award-Winning Paper in 12 Hours Unattended

The MiniMax team handed M3 an ICLR 2025 Outstanding Paper Award winner — Learning Dynamics of LLM Finetuning — and asked it to reproduce the work independently.

M3 ran autonomously for nearly 12 hours, unattended throughout, and ultimately:

  • Produced 18 commits and 23 experiment charts
  • Successfully matched the predicted probability trends in the SFT stage
  • Clearly observed the squeezing effect in the DPO experiments
  • Validated the Extend mitigation method proposed in the original paper

This test simultaneously drew on M3's three strengths: 1M ultra-long context (reading the full paper), top-tier coding capability (writing experiment code), and native multimodality (generating and interpreting experiment charts).

How to Try It

You can experience MiniMax M3 first-hand on these platforms:

  • MiniMax Code: Online coding environment
  • Token Plan: API calling service
  • MiniMax API: Direct integration into your project

The Interactive User Simulator Framework

M3's Coding gains don't come from Benchmark training alone. MiniMax built an interactive user simulator framework that models how real developers behave during collaboration:

  • Requirement clarification and solution discussion
  • Feedback correction and continuous task switching
  • Multi-round iterative optimization on complex projects

This lets the Agent stop passively executing instructions and instead actively collaborate with the user to get tasks done. The next generation of Agent Coding will compete not just on code generation, but on long-horizon collaboration and human-Agent coordination efficiency.

Who It's For

  • Developers: Need an Agent with strong Coding capability to handle complex software engineering tasks
  • Researchers: Need to process ultra-long documents (papers, codebases) and perform complex reasoning
  • Agent builders: Need a foundation model that can simultaneously understand text, images, and video and operate a computer
  • Enterprise users: Need to handle long-context enterprise knowledge bases and document analysis scenarios

Related articles

Coze 3.0: Put a Team of AI Agents to Work for You
AI Products

Coze 3.0: Put a Team of AI Agents to Work for You

Coze 3.0 is officially out, with multi-agent collaboration, remote control of your computer from your phone, one-click import of local agents, and the ability to turn a story idea straight into video.

Toolin Editorial Team
Kimi Work Beta: From an Agent That Writes Code to an Agent That Does Work
AI Products

Kimi Work Beta: From an Agent That Writes Code to an Agent That Does Work

Moonshot AI launches Kimi Work Beta, a general-purpose local agent for knowledge workers supporting 300 sub-agents in parallel, 13-hour long-running tasks, browser control, and skill installation.

Toolin Editorial Team
OpenSquilla Meta Skill: Packing an Entire Workflow into a Single Skill
AI Products

OpenSquilla Meta Skill: Packing an Entire Workflow into a Single Skill

OpenSquilla's new Meta Skill feature nests multiple sub-skills inside one skill, running long-horizon workflows end to end while cutting token costs by 60-80%.

Toolin Editorial Team
Mashangfei: From One Sentence to a Complete, Running Business
AI Products

Mashangfei: From One Sentence to a Complete, Running Business

Mashangfei packages the three core links of a one-person company into a closed loop: generate a complete app with an admin backend from one sentence, an AI business assistant that auto-produces posters and copy, and 7x24 AI customer service plugged into WeChat in one click.

Toolin Editorial Team
Fully Automated AI Video Editing: A Three-Tool Stack for 100 Videos a Day
AI Tutorials

Fully Automated AI Video Editing: A Three-Tool Stack for 100 Videos a Day

A HyperFrames + Remotion + Git three-tool stack for fully automated AI video editing, from HTML to React components, with complete install commands and pitfall notes.

Toolin Editorial Team
A Hands-On Guide to Claude Code /workflows
AI Tutorials

A Hands-On Guide to Claude Code /workflows

A detailed look at when and how to use the Claude Code /workflows feature, using multi-agent parallelism for codebase sweeps and hard-problem research.

Toolin Editorial Team