Step 3.7 Flash in Claude Code: A Hands-On Guide
A hands-on test of StepFun's open-source Flash model integrated into Claude Code, using complex Agent workflows to see whether a Chinese model can stand in for closed-source foundations


Step 3.7 Flash in Claude Code: A Hands-On Guide
A hands-on test of StepFun's open-source Flash model integrated into Claude Code, using complex Agent workflows to see whether a Chinese model can stand in for closed-source foundations
StepFun has open-sourced Step 3.7 Flash under Apache 2.0, positioning it around "agent efficiency" — running the whole chain fast and steady inside real workflows. The official docs explicitly list direct integration with mainstream agent tools such as Claude Code, Cline, and Roo Code. This article records a complete hands-on test: driving Claude Code with Step 3.7 Flash as the base model through two high-complexity agent workflows, to see whether it can actually hold up.
What Is Step 3.7 Flash
Step 3.7 Flash is the Flash model StepFun released and open-sourced at the end of May 2025:
- Architecture: Sparse MoE (Mixture of Experts) — sizable overall, but only a small squad of the most relevant "experts" activates each time
- Speed: Peak generation speed of 400 tokens per second, with 256K context
- License: Apache 2.0, downloadable from GitHub, HuggingFace, and ModelScope
- Positioning: Not "the smartest", but "fast and steady on agent tasks"

On agent benchmarks like SWE-Bench and ClawEval, it posts genuinely competitive scores for its size. The real selling point isn't the highest score — it's delivering this level consistently with fewer activated parameters and faster speed.
How to Hook It Into Claude Code
StepFun's official docs list the tools it supports:

Configuration Steps
- Register on the StepFun console and obtain an API Key
- Configure the Claude Code route: Route
step-3.7-flashinto Claude Code via CCR (model routing) - Configure the launch command: Set up a
stepfuncommand that launches Claude Code driven by Step 3.7 Flash - Handle web search: After swapping the base model, Claude Code's native search stops working; switch to Tavily's MCP tool instead
# Reference configuration sketch (see StepFun's official docs for exact parameters)
# Add a step-3.7-flash route in Claude Code's model configuration
# See StepFun's official site's harness docs for the exact integration method
Tip: StepFun's website documents the integration method for every tool (Claude Code, Cline, and so on) in detail. If you'd rather not configure it yourself, try handing the integration docs to any Chinese desktop agent and letting it set things up for you.
Hands-On Results
Test 1: Nüwa (deep research + persona Skill generation)
Task goal: distill an investment-perspective Skill for a figure in the AI field.
Execution process:
- After confirming the subject, it spun up 6 sub-agents researching in parallel in one go (works, interviews, style, criticism, decision records, latest developments)
- The slowest of the 6 agents ran for 22 minutes, and Step 3.7 Flash kept the parallel state under control throughout, with no crossed results
- Once all agents reported back, it proactively stopped to present a research quality summary and waited for confirmation before continuing
- It distilled 6 core thinking models and 8 decision heuristics, and generated a runnable Skill in one shot
- It started its own independent review agent to poke holes, then patched in trigger words, fact-checking, and other details per the review comments

Test 2: Darwin 2.0 (multi-judge Skill optimization)
Task goal: optimize a stand-up comedy Skill with Darwin 2.0.
Execution process:
- Created a git branch, designed test cases, and ran a round of baseline scoring
- After locating the weakest dimension, iterated: each round relaunched two brand-new independent judge agents for blind scoring
- Committed after every change; rolled back whenever the score gain wasn't enough
- Once gains narrowed, the early-stopping mechanism triggered automatically

Cost Reference
Per StepFun console pricing:
- Input: 1.35 yuan per million tokens
- Output: 8.1 yuan per million tokens
That's Flash-tier pricing, well suited to high-frequency agent workflows.
Conclusion
In hands-on testing, Step 3.7 Flash showed two key traits:
- It skipped none of the steps it should have taken: It didn't cut corners just because the task was complex — 6 parallel agents plus multiple review rounds all executed in full
- It stopped honestly wherever it should stop and ask: What needed confirmation didn't get decided unilaterally; what needed reporting got reported proactively
It isn't flawless — a few edit operations errored out mid-run (mostly problems with the local tool environment), and the model went back and retried a different way to get past them. But for an open-source Flash model to deliver execution comparable to subscription Claude Code in complex agent workflows already exceeded expectations.
Where to Get It
- StepFun's site: https://platform.stepfun.com/
- GitHub: search for Step-3.7-Flash
- HuggingFace: search for step-3.7-flash
If you use tools like Claude Code or Codex but have cost concerns, a model that plugs into your existing workflow, is open source, and reliably runs the whole chain to completion is genuinely worth a try.