Step 3.7 Flash Hands-On: 400 TPS Inference, Agent Tasks at 1/9 the Cost of Claude

·Toolin Editorial Team

StepFun has released Step 3.7 Flash: 400 tokens/second inference, 11B activated parameters delivering 97% of Claude Opus 4.6's performance, open source and locally deployable

Step 3.7 Flash Hands-On: 400 TPS Inference, Agent Tasks at 1/9 the Cost of Claude

StepFun has released Step 3.7 Flash, a new-generation Flash model built for production-grade Agents. Its core selling points are direct: 400 TPS inference speed, per-task cost at 1/9 of Claude Opus 4.6, while delivering 97% of the performance.

This is not another story of a small model chasing leaderboard scores. Step 3.7 Flash is an MoE model with 11B activated parameters, it is open source, and it supports local execution via mlx-vlm on Apple Silicon.

Step 3.7 Flash positioning

Step 3.7 Flash is positioned as a new-generation Agentic foundation model

What Is Step 3.7 Flash

Step 3.7 Flash is StepFun's latest-generation Flash model, following Step 3.5 Flash (which once topped OpenRouter Trending and ranked No.1 globally in OpenClaw call volume).

Core specs:

  • Architecture: MoE, 11B activated parameters
  • Vision understanding: 196B + 1.8B ViT
  • Inference speed: 400 TPS
  • Open source: Supports mlx-vlm; runs 32K context on 128GB Apple Silicon machines

Its design philosophy is interesting: for a Flash model with 11B activated parameters, cramming massive visual knowledge into the weights doesn't pay off. Step keeps only the core reasoning engine in the weights, externalizes perception boundaries and world knowledge to inference time, and relies on blistering speed — "take a few more looks, check a few more times" — to make up for the smaller parameter count.

Five Hands-On Scenarios

Scenario 1: Batch-Processing Invoice Reimbursements

Feed Step 3.7 Flash 12 casually snapped photos of invoices (skewed angles, blurry shots, dining, electronics, and travel all jumbled together). It not only recognizes the amount, tax, and merchant name on each invoice, but also judges which fields actually need to be filled in for reimbursement, automatically organizes everything into a unified table, and exports to Excel in one click.

What worked end to end was the full chain of "recognize -> understand -> organize -> export".

Invoice batch processing

Automatically recognizes invoice information and organizes it into a table

Scenario 2: Making Sense of Professional Software Interfaces

Given a screenshot of the Blender interface and asked "how do I delete this cube", the model automatically boxes the interface, reads the outliner, toolbar, and current edit mode, and gives an operation path specific down to every step. Being able to give executable operation advice inside a 3D application as information-dense as Blender means it is already capable of moving into professional tools.

Scenario 3: Airplane Cockpit Operation Guide

Give the model a densely packed screenshot of an airplane cockpit, with the single input "how to take off". It automatically boxes the cockpit area, recognizes what each key instrument means, works out the order of operations, and walks through step by step when to push the throttle and when to retract the landing gear.

Cockpit operation guide

Going from "reading the interface" to "teaching you how to operate it" is an order-of-magnitude jump in difficulty

Scenario 4: High-Speed Deep Research

Give it one sentence: "Around mass production of humanoid robots in 2026, give me a one-page decision summary I can act on." What it hands back is not a pile of links but a complete report that states its judgment up front, uses a table in the middle to compare the mass-production progress and risks of six companies (Tesla, Figure, Unitree, Agibot, 1X, Agility), and closes with three actionable focus points carrying time nodes. Every number carries a source attached.

Scenario 5: GUI Understanding and Computer Use

Give it a CapCut screenshot and one line: "export this clip as 1080P, 30 frames." It not only located the export button, but also proactively noticed that the current color format is 1080i (interlaced) rather than 1080P, and reminded you to change it manually. It even noticed there was more than one clip on the timeline, specifically pointing out that "what's being exported is the whole project, not this single clip."

Beyond reading the interface, it catches details users tend to miss

An Interesting Emergent Behavior

After writing a stretch of frontend code, Step 3.7 Flash will switch into the GUI by itself to test the page it just generated — check the rendering, click the interactive buttons — then go back and revise the code based on what it saw.

Write code -> look at the interface -> fix the code: nobody taught it this combination punch; it figured it out on its own.

Performance Numbers

MetricStep 3.7 FlashComparison
Inference speed400 TPSTop of the industry
Per-task cost1/9 of Claude Opus 4.689% cost reduction
Performance comparison97% of Claude Opus 4.6Gap of only 3%
Visual Benchmark (V)95.3Benchmarked against Kimi K2.6 (96.9)
Activated parameters11BMoE architecture

Feedback From Overseas Developers

  • One developer who switched back from Gemini 3.5 Flash to Step 3.7 Flash had it find more than 7 bugs in one go
  • Some describe the speed as "ridiculously fast"
  • mlx-vlm support means 4-bit quantization runs 32K context on Apple Silicon
  • One developer said they are seriously considering it as a replacement for other models for the first time

How to Use It

API calls: Connect through OpenRouter or StepFun's official API.

Local deployment: Open source, supports the mlx-vlm framework. On a 128GB Apple Silicon machine, the 4-bit quantized build runs 32K context.

Who It's For

  • Agent developers: Need a low-cost, high-throughput foundation model powering production-grade Agent workflows
  • Enterprise users: Need to batch-process invoices, documents, screenshots, and other vision-intensive tasks
  • Individual developers: Want a locally deployed, privacy-preserving AI model
  • Deep Research scenarios: Need decision summaries and information synthesis generated fast
  • GUI automation: Need Agents to understand and operate desktop application interfaces

Frequently Asked Questions

Q: Is the gap between Step 3.7 Flash and flagship models large? A: It reaches 97% of Claude Opus 4.6's performance on Agent tasks at 1/9 the cost. For the vast majority of production scenarios, that gap is acceptable.

Q: Can it run locally? A: Yes. It is open source and supports mlx-vlm. On a 128GB Apple Silicon machine, the 4-bit quantized build runs 32K context.

Q: Compared with Step 3.5 Flash, what improved? A: Mainly significant gains in vision understanding, GUI operation, Deep Research, and code generation, while keeping the Flash family's speed and cost advantages.

Q: Are there restrictions on commercial use? A: It is open source; check the official GitHub repository for the specific license terms.

Related articles

Hyra-1.0: Tencent Hunyuan's Scientific Discovery Agent
AI Products

Hyra-1.0: Tencent Hunyuan's Scientific Discovery Agent

A research agent capable of recursive self-improvement, covering AI training optimization, open math problems, quantum computing, and drug design, with multiple records broken.

Toolin Editorial Team
Meshy: Turning Text and Images into Usable 3D Models in 20 Seconds
AI Products

Meshy: Turning Text and Images into Usable 3D Models in 20 Seconds

AI text/image-to-3D tool with a two-step workflow and FBX/OBJ/GLB/STL export, plus an official API and platform plugins.

Toolin Editorial Team
Qwen-Image-3.0: Alibaba Qwen's Third-Generation Image Generation Model
AI Products

Qwen-Image-3.0: Alibaba Qwen's Third-Generation Image Generation Model

Up to 4.5k tokens of input, crisp rendering of 10px-scale small text, and 12 languages — renders complex layouts like posters, exam papers, and storyboard comics in a single pass.

Toolin Editorial Team
Build an Automated Reddit Overseas Customer Acquisition Workflow with Qwen 3.8
AI Tutorials

Build an Automated Reddit Overseas Customer Acquisition Workflow with Qwen 3.8

An eight-step fully automated pipeline: Apify scraping plus Qwen 3.8 scoring/drafting/self-review/retrospective — 260 posts filtered down to 22 opportunities, at a tenth of the cost.

Toolin Editorial Team
Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion
AI Products

Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion

Alibaba's new-generation TTS model controls laughter, gasps, and anger with tags, delivers 48kHz film-grade audio, and tops the global Speech Arena.

Toolin Editorial Team
Qwen3.8-Max Preview: A Hands-On Early Access Guide
AI Products

Qwen3.8-Max Preview: A Hands-On Early Access Guide

Qwen's 2.4T-parameter flagship preview is live on Token Plan, Qoder, and the Qwen website; officially rated second only to Fable 5 overall.

Toolin Editorial Team