
StepFun
Officially listedA leading Chinese multimodal reasoning model family, with its Step 3 model open-sourced.
StepFun
StepFun is a leading Chinese AI large model company with its Step series of multimodal large models. The latest Step 3 uses a MoE architecture with 321B total parameters and 38B activated parameters, was officially open-sourced in July 2025, and reaches industry-leading levels on multiple benchmarks.
Core Capabilities
- Multimodal Reasoning: Strong visual perception and complex reasoning, with a 5B Vision Encoder
- Efficient Inference: MFA attention mechanism achieving 4039 token/gpu/s on Hopper GPUs
- Web Search: Online search for real-time information
- Image Generation: Text-to-image capability
- Knowledge Base Q&A: Supports custom knowledge bases
Use Cases Intelligent conversation, content creation, code generation, and document analysis. Developers can integrate via API, while ordinary users can use it directly through the web and app.
Unique Advantages Step 3 performs excellently on benchmarks including MMMU, MathVision, and AIME 2025, with inference efficiency significantly higher than DeepSeek V3. As an open-source model, it offers a high degree of technical transparency.
Editor's Recommendation A standout among Chinese large models — open-sourcing Step 3 demonstrates real technical strength. API pricing is competitive, well suited to cost-sensitive developers.
Pricing
### 定价模式:免费增值/API付费 **起步价**:免费体验 #### API定价(限时折扣) - **输入**:1.5元/百万token - **输出**:4元/百万token #### 其他信息 网页版和App可免费使用,API按量计费。 — Visit website
FAQ
Is StepFun free?
The web version and app are free to use. The API is billed by usage: 1.5 yuan per million input tokens and 4 yuan per million output tokens (limited-time discount).
What are StepFun's main features?
Core features include multimodal conversation, web search, image generation, knowledge base Q&A, and code generation.
What are the characteristics of StepFun's Step 3 model?
Step 3 uses a MoE architecture with 321B total parameters and 38B activated parameters. It has strong visual perception and complex reasoning capabilities and leads the industry in multiple benchmarks.
How does StepFun compare with other Chinese large models?
Step 3's inference efficiency is higher than DeepSeek V3's (4039 vs 2324 token/gpu/s), and it performs excellently on benchmarks such as MMMU and MathVision — the best among comparable open-source models.
Is StepFun suitable for developers?
Very much so. It provides an open platform API with text, vision, and other multimodal capabilities at competitive prices, and Step 3 is already open-sourced.
What access methods does StepFun support?
It supports three access methods: the web version (stepfun.com), the mobile app, and the open platform API (platform.stepfun.com).