What is Step 5 Preview?
Step 5 Preview is a reasoning-oriented model from StepFun, the China-based AI provider operated by Shanghai StepFun Intelligence Co., Ltd. It is currently positioned as the provider’s flagship preview model for demanding professional and agentic tasks. StepFun makes it available through its products and developer platform, while an open-weight release was announced for October 15, 2026. That planned release should not be treated as evidence that the model’s weights are already available.
The model is designed for work that may involve multiple steps, large documents, codebases, research material, financial information, or visual inputs. In practical terms, it is better suited to investigating a problem, planning a sequence of actions, reviewing substantial context, or helping with a complicated software task than to serving as a minimal-cost chatbot for short, routine prompts.
Where it fits in StepFun’s current catalog
Step 5 Preview sits at the high-capability end of StepFun’s current model catalog. StepFun’s platform also highlights models such as Step 3.7 Flash, Step 3.5 Flash, and StepAudio 3, but those names represent different positions and capabilities in the provider’s lineup. Step 5 Preview is the model specifically presented for agentic work, software engineering, professional knowledge work, and finance.
This positioning matters because the model is not a general-purpose media generator. Although the wider StepFun ecosystem includes image, video, speech, music, and other multimodal services, Step 5 Preview itself is documented as accepting text and vision input and producing text output. It should therefore be evaluated as a reasoning and work-assistance model rather than as an all-in-one creative model.
Core specifications at a glance
| Specification | Step 5 Preview |
|---|---|
| Provider | StepFun |
| Model family | Step 5 |
| Status | Current preview model |
| Architecture | Sparse Mixture of Experts |
| Total parameters | 600 billion |
| Active parameters per token | 27 billion |
| Context window | 1,000,000 tokens |
| Input | Text and vision |
| Output | Text |
| Tool use | Supported according to the supplied model data |
| Maximum output tokens | Not publicly specified in the supplied sources |
| Knowledge cutoff | Not publicly specified |
A sparse Mixture-of-Experts, or MoE, model contains many specialized parameter groups but activates only part of them for each token. StepFun reports 600 billion total parameters and 27 billion active parameters per token. These figures describe the architecture; they do not by themselves guarantee a particular response quality, operating cost, or speed in every workload.
One-million-token context and vision input
The defining technical feature of Step 5 Preview is its 1-million-token context window. A context window is the amount of information the model can consider in a request and its surrounding conversation, including text and other supported inputs. A window of this size can be useful for long code repositories, extensive research material, large collections of business documents, or extended agentic tasks.
The context limit should not be confused with a guaranteed output length. StepFun’s supplied materials specify the context window but do not specify a maximum output-token limit. Users should also expect practical constraints from an application’s interface, API configuration, latency, and cost even when the model’s nominal context capacity is very large.
Step 5 Preview supports vision input. This makes it suitable for tasks such as examining an image alongside written instructions, interpreting visual material during analysis, or using screenshots as part of a software or research workflow. The supplied specifications do not list audio or video input for this model, and they do not describe image, audio, video, or music generation. Its output is text only.
Reasoning, coding, and tool use
StepFun positions Step 5 Preview for reasoning-intensive work rather than only direct question answering. Its stated focus includes agentic workflows, software engineering, professional knowledge work, finance, and research. An agentic workflow is a task in which the model helps break down a goal, work through intermediate steps, use tools, and continue toward an outcome instead of producing a single isolated answer.
For coding, the model is intended for software engineering and coding assistance. Suitable examples include understanding a large codebase, proposing implementation steps, reviewing code, investigating a bug, or helping coordinate a multi-step development task. The supplied research does not provide a guaranteed programming-language list, a maximum repository size, or a specific coding benchmark score, so those details should be confirmed in the current StepFun documentation before selecting it for a production engineering process.
Tool use is listed as supported in the supplied model data. This can make the model more useful in applications that allow it to call external functions or services, but tool support does not mean that Step 5 Preview independently has unrestricted access to the web, databases, files, or software. Those capabilities depend on the surrounding StepFun product or developer implementation. The supplied data does not verify a separate built-in web-search capability for this model.
Pricing and cost trade-offs
StepFun’s supplied pricing documentation lists the following prices:
- Uncached input: ¥7 per 1 million tokens
- Cached input: ¥0.35 per 1 million tokens
- Output: ¥20 per 1 million tokens
Cached-input pricing is separate from the standard uncached input price. It may matter for applications that repeatedly send the same or substantially reusable context, but the exact eligibility and caching behavior should be checked in the current platform documentation. Output tokens cost more than uncached input tokens, so prompts that produce long responses or extended reasoning may have a material effect on total usage.
The model’s cost profile reflects a capability-versus-efficiency trade-off. A large-context reasoning model can reduce the need to divide a complex task into many smaller requests, but it may be unnecessary for short classification, simple extraction, or routine chat. A faster or less expensive model may be more appropriate when the task does not benefit from extended reasoning, visual understanding, or a very large context.
Main strengths and limitations
Strengths
- Very large context: The 1-million-token window is useful for long documents, codebases, and extended workflows.
- Agentic positioning: StepFun specifically targets tasks involving multiple steps, planning, and tool-assisted work.
- Professional focus: Software engineering, finance, research, and knowledge work are central intended uses.
- Vision input: The model can incorporate visual information alongside text.
- Tool support: Applications can use the model in function- or tool-enabled workflows, subject to the surrounding implementation.
- Parameter efficiency at inference: The reported MoE design activates 27 billion of 600 billion total parameters per token, although architecture figures should not be interpreted as a universal speed guarantee.
Limitations
- Preview status: Behavior, availability, pricing, and documentation may change while the model remains a preview.
- Text-only output: It is not the right choice when the model must natively generate images, audio, video, or music.
- Unspecified output limit: The supplied sources do not state a maximum output-token limit.
- Unspecified knowledge cutoff: StepFun has not publicly specified a knowledge cutoff in the supplied materials.
- Not immediately open weight: The announced open-weight date is in the future relative to the current availability described in the research.
- Regional availability: StepFun is oriented toward mainland China, and access, registration, billing, and developer services may have regional or eligibility restrictions.
- Vendor-reported evaluation claims: StepFun reports strong results across coding, agentic, finance, reasoning, and multimodal evaluations, but the supplied research identifies many of these results as vendor-reported. They should not be treated as independently verified guarantees.
When to choose Step 5 Preview
Choose Step 5 Preview when the task benefits from a large working context and sustained reasoning. It is a strong candidate for reviewing a substantial body of documentation, analyzing financial or professional material, assisting with a complex coding project, interpreting screenshots alongside instructions, or coordinating a tool-enabled research workflow.
It is also a reasonable choice when keeping related material in one context is more convenient than repeatedly summarizing and resending it. For example, a development assistant may need to inspect a large set of files and maintain a consistent understanding while proposing changes. A research workflow may need to compare many documents before producing a written synthesis.
Another model may be more appropriate when speed and minimum cost matter more than long-context reasoning. A smaller or faster model can be preferable for short replies, high-volume routine processing, basic extraction, or simple transformations. A dedicated image, video, speech, or music model is the better option when non-text generation is the central requirement. Users seeking an immediately open-weight model should also wait for verified availability rather than relying on the announced future release.
Availability and practical expectations
Step 5 Preview is currently available through StepFun products and the StepFun API platform according to the supplied research. The exact interface, account requirements, regional eligibility, and operational limits may differ between consumer products and developer access. The model’s API pricing is documented in Chinese yuan, and the provider’s broader ecosystem is primarily oriented toward China.
Overall, Step 5 Preview is best understood as a high-capability preview model for long-running, text-producing work with optional visual input and tool integration. Its strongest differentiators are the 1-million-token context window, the stated focus on agentic and professional tasks, and its positioning at the top of StepFun’s current model lineup. Its main trade-offs are preview uncertainty, regional availability, unspecified maximum output and knowledge cutoff details, and the fact that it does not replace dedicated generative media models.

