Step 3.7 Flash: The Agent Efficiency Model at 409 tok/s

·Toolin Editorial Team

StepFun's new model hits 409 tokens/s output speed, costs 1/9 of Claude Opus 4.6 per task while matching 97% of its coding ability, and is designed for high-frequency Agent call scenarios.

Step 3.7 Flash: The Agent Efficiency Model at 409 tok/s

Now that Agents have become the mainstream deployment form, the key question in model competition is no longer "who is smarter" but "who can run more tasks, faster and more reliably, per unit of cost." StepFun's Step 3.7 Flash was born for exactly this battleground.

What Is Step 3.7 Flash

Step 3.7 Flash is StepFun's newly released Agent efficiency model, and it has taken multiple first places on Artificial Analysis (the AA leaderboard):

  • Output speed: 409 tokens/s, first among mainstream models (for comparison, GPT-5.3 runs at about 70 tok/s)
  • End-to-end response: 7.1 seconds to complete the prescribed benchmark task
  • Intelligent efficiency: delivers practically valuable task capability while maintaining extremely high output efficiency
  • Speed-price ratio: lands in the AA leaderboard's most attractive quadrant

What the Speed and Cost Actually Mean

In Agent scenarios, a single task requires dozens of model calls. Two extra seconds and a slightly higher price on each call accumulate over dozens of calls, and both latency and cost become a headache.

Step 3.7 Flash pricing: $0.2 per million input tokens and $1.15 per million output tokens. Its per-task cost is about 1/9 that of Claude Opus 4.6, yet its coding ability reaches 97% of Opus's.

One developer benchmarked Step 3.7 Flash alongside several mainstream models, and 3.7 Flash pulled far ahead of the pack at 2123 tok/s. Under NVFP4 settings, peak throughput even touched 6000 tok/s.

Multi-model speed comparison

Hands-On Scenarios

Multimodal Understanding

Upload a photo of a dexterous robot hand, and Step 3.7 Flash quickly identifies the product model from its visual details, automatically searches the web for full-dimension specifications, and organizes them into a structured table.

Tool Orchestration: Expense Report Filing

Throw a folder full of invoices at Step 3.7 Flash (via OpenClaw), and in under 60 seconds it produces an expense Excel and an explanatory memo for finance, with every item verified correct.

Multi-Agent Cluster: a 40-Person Product Review Panel

Have Step 3.7 Flash generate 40 differentiated virtual users to vote and rank 5 new features of a food-delivery app. All 40 Agents returned valid responses, with no role confusion or format drift. The final votes were clear, and the demographic segmentation sensible.

Multi-Agent cluster

Cache Hit Rate: Engineering Muscle on Display

One developer compiled 398 core data points from 60+ providers on OpenRouter into a "cache hit rate leaderboard." StepFun came in at 86.1%, S-tier and second in the world, behind only DeepSeek.

A high cache hit rate means the inference infrastructure is well engineered. In long-running scenarios like Agents and RAG, repeated context prefixes get reused efficiently, translating directly into lower costs and higher throughput.

How to Use It

Who It's For

  • Agent developers: the value-for-money pick for high-frequency call scenarios
  • Enterprise users: cost control in long-running, multi-turn interaction scenarios
  • Tool orchestration applications: scenarios that need a model with both speed and accuracy

Step 3.7 Flash's core value lies not in how "smart" it is at single-turn Q&A, but in its "completion efficiency" within Agent workflows. If your project needs a model that gets called repeatedly and runs for long stretches, this model deserves a serious evaluation.