Step 5

Step 5 Preview

by StepFun · Current preview; available through StepFun products and API. Open-weight release planned for October 15, 2026.

Step 5 Preview is StepFun’s flagship preview reasoning model for long-running agentic workflows, coding, finance, research, and professional knowledge work. It offers a 1-million-token context window, text and vision input, text output, tool support, and sparse Mixture-of-Experts architecture. Pricing is ¥7 per million uncached input tokens, ¥0.35 per million cached input tokens, and ¥20 per million output tokens. Its main limitations are preview status, regional availability, unspecified maximum output and knowledge cutoff details, and the lack of native image, audio, or video output.

Text Reasoning Coding
Step 5 Preview is StepFun’s flagship preview model for tasks that require sustained reasoning across large amounts of information. It is aimed at agentic work, software engineering, professional analysis, finance, research, and other workflows where a short conversational response is not enough. The model accepts text and vision input, supports tool use, and offers a 1-million-token context window, but it does not generate images, audio, or video.
Outputs

What Step 5 Preview can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Prompt caching
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Step 5
Model type Reasoning
Context window 1M tokens
Release date 2026-09-20
Status Current preview; available through StepFun products and API. Open-weight release planned for October 15, 2026.
Knowledge cutoff notes

StepFun does not publicly specify a knowledge cutoff for Step 5 Preview in the available official model announcement or platform materials.

Model notes

Step 5 Preview is a sparse Mixture-of-Experts model with 600B total parameters and 27B active parameters per token. StepFun describes it as a flagship model for agentic work, software engineering, professional knowledge work, and finance. The model accepts text and vision input and provides a 1M-token context window. StepFun's launch material reports strong results on coding, agentic, finance, reasoning, and multimodal evaluations, but many benchmark results are vendor-reported. The model is currently available through StepFun products and API access. StepFun announced an open-weight release for October 15, 2026; that future release does not mean weights are currently available. Pricing is listed in Chinese yuan on StepFun's platform documentation; cached-input pricing is separate from standard uncached input pricing. Editorial scores are comparative estimates, not provider specifications.

Cost

Model pricing

Input ¥7 per 1 million uncached input tokens; ¥0.35 per 1 million cached input tokens
Output ¥20 per 1 million output tokens
Model guide

Step 5 Preview: StepFun’s Long-Context Model for Agentic Work

Step 5 Preview is StepFun’s current flagship preview model for complex agentic workflows, software engineering, professional knowledge work, finance, and vision-assisted tasks. It combines a sparse Mixture-of-Experts architecture with 600 billion total parameters, 27 billion active parameters per token, a 1-million-token context window, and text-and-vision input. The model is available through StepFun products and API access, with pricing of ¥7 per million uncached input tokens, ¥0.35 per million cached input tokens, and ¥20 per million output tokens.

What is Step 5 Preview?

Step 5 Preview is a reasoning-oriented model from StepFun, the China-based AI provider operated by Shanghai StepFun Intelligence Co., Ltd. It is currently positioned as the provider’s flagship preview model for demanding professional and agentic tasks. StepFun makes it available through its products and developer platform, while an open-weight release was announced for October 15, 2026. That planned release should not be treated as evidence that the model’s weights are already available.

The model is designed for work that may involve multiple steps, large documents, codebases, research material, financial information, or visual inputs. In practical terms, it is better suited to investigating a problem, planning a sequence of actions, reviewing substantial context, or helping with a complicated software task than to serving as a minimal-cost chatbot for short, routine prompts.

Where it fits in StepFun’s current catalog

Step 5 Preview sits at the high-capability end of StepFun’s current model catalog. StepFun’s platform also highlights models such as Step 3.7 Flash, Step 3.5 Flash, and StepAudio 3, but those names represent different positions and capabilities in the provider’s lineup. Step 5 Preview is the model specifically presented for agentic work, software engineering, professional knowledge work, and finance.

This positioning matters because the model is not a general-purpose media generator. Although the wider StepFun ecosystem includes image, video, speech, music, and other multimodal services, Step 5 Preview itself is documented as accepting text and vision input and producing text output. It should therefore be evaluated as a reasoning and work-assistance model rather than as an all-in-one creative model.

Core specifications at a glance

SpecificationStep 5 Preview
ProviderStepFun
Model familyStep 5
StatusCurrent preview model
ArchitectureSparse Mixture of Experts
Total parameters600 billion
Active parameters per token27 billion
Context window1,000,000 tokens
InputText and vision
OutputText
Tool useSupported according to the supplied model data
Maximum output tokensNot publicly specified in the supplied sources
Knowledge cutoffNot publicly specified

A sparse Mixture-of-Experts, or MoE, model contains many specialized parameter groups but activates only part of them for each token. StepFun reports 600 billion total parameters and 27 billion active parameters per token. These figures describe the architecture; they do not by themselves guarantee a particular response quality, operating cost, or speed in every workload.

One-million-token context and vision input

The defining technical feature of Step 5 Preview is its 1-million-token context window. A context window is the amount of information the model can consider in a request and its surrounding conversation, including text and other supported inputs. A window of this size can be useful for long code repositories, extensive research material, large collections of business documents, or extended agentic tasks.

The context limit should not be confused with a guaranteed output length. StepFun’s supplied materials specify the context window but do not specify a maximum output-token limit. Users should also expect practical constraints from an application’s interface, API configuration, latency, and cost even when the model’s nominal context capacity is very large.

Step 5 Preview supports vision input. This makes it suitable for tasks such as examining an image alongside written instructions, interpreting visual material during analysis, or using screenshots as part of a software or research workflow. The supplied specifications do not list audio or video input for this model, and they do not describe image, audio, video, or music generation. Its output is text only.

Reasoning, coding, and tool use

StepFun positions Step 5 Preview for reasoning-intensive work rather than only direct question answering. Its stated focus includes agentic workflows, software engineering, professional knowledge work, finance, and research. An agentic workflow is a task in which the model helps break down a goal, work through intermediate steps, use tools, and continue toward an outcome instead of producing a single isolated answer.

For coding, the model is intended for software engineering and coding assistance. Suitable examples include understanding a large codebase, proposing implementation steps, reviewing code, investigating a bug, or helping coordinate a multi-step development task. The supplied research does not provide a guaranteed programming-language list, a maximum repository size, or a specific coding benchmark score, so those details should be confirmed in the current StepFun documentation before selecting it for a production engineering process.

Tool use is listed as supported in the supplied model data. This can make the model more useful in applications that allow it to call external functions or services, but tool support does not mean that Step 5 Preview independently has unrestricted access to the web, databases, files, or software. Those capabilities depend on the surrounding StepFun product or developer implementation. The supplied data does not verify a separate built-in web-search capability for this model.

Pricing and cost trade-offs

StepFun’s supplied pricing documentation lists the following prices:

  • Uncached input: ¥7 per 1 million tokens
  • Cached input: ¥0.35 per 1 million tokens
  • Output: ¥20 per 1 million tokens

Cached-input pricing is separate from the standard uncached input price. It may matter for applications that repeatedly send the same or substantially reusable context, but the exact eligibility and caching behavior should be checked in the current platform documentation. Output tokens cost more than uncached input tokens, so prompts that produce long responses or extended reasoning may have a material effect on total usage.

The model’s cost profile reflects a capability-versus-efficiency trade-off. A large-context reasoning model can reduce the need to divide a complex task into many smaller requests, but it may be unnecessary for short classification, simple extraction, or routine chat. A faster or less expensive model may be more appropriate when the task does not benefit from extended reasoning, visual understanding, or a very large context.

Main strengths and limitations

Strengths

  • Very large context: The 1-million-token window is useful for long documents, codebases, and extended workflows.
  • Agentic positioning: StepFun specifically targets tasks involving multiple steps, planning, and tool-assisted work.
  • Professional focus: Software engineering, finance, research, and knowledge work are central intended uses.
  • Vision input: The model can incorporate visual information alongside text.
  • Tool support: Applications can use the model in function- or tool-enabled workflows, subject to the surrounding implementation.
  • Parameter efficiency at inference: The reported MoE design activates 27 billion of 600 billion total parameters per token, although architecture figures should not be interpreted as a universal speed guarantee.

Limitations

  • Preview status: Behavior, availability, pricing, and documentation may change while the model remains a preview.
  • Text-only output: It is not the right choice when the model must natively generate images, audio, video, or music.
  • Unspecified output limit: The supplied sources do not state a maximum output-token limit.
  • Unspecified knowledge cutoff: StepFun has not publicly specified a knowledge cutoff in the supplied materials.
  • Not immediately open weight: The announced open-weight date is in the future relative to the current availability described in the research.
  • Regional availability: StepFun is oriented toward mainland China, and access, registration, billing, and developer services may have regional or eligibility restrictions.
  • Vendor-reported evaluation claims: StepFun reports strong results across coding, agentic, finance, reasoning, and multimodal evaluations, but the supplied research identifies many of these results as vendor-reported. They should not be treated as independently verified guarantees.

When to choose Step 5 Preview

Choose Step 5 Preview when the task benefits from a large working context and sustained reasoning. It is a strong candidate for reviewing a substantial body of documentation, analyzing financial or professional material, assisting with a complex coding project, interpreting screenshots alongside instructions, or coordinating a tool-enabled research workflow.

It is also a reasonable choice when keeping related material in one context is more convenient than repeatedly summarizing and resending it. For example, a development assistant may need to inspect a large set of files and maintain a consistent understanding while proposing changes. A research workflow may need to compare many documents before producing a written synthesis.

Another model may be more appropriate when speed and minimum cost matter more than long-context reasoning. A smaller or faster model can be preferable for short replies, high-volume routine processing, basic extraction, or simple transformations. A dedicated image, video, speech, or music model is the better option when non-text generation is the central requirement. Users seeking an immediately open-weight model should also wait for verified availability rather than relying on the announced future release.

Availability and practical expectations

Step 5 Preview is currently available through StepFun products and the StepFun API platform according to the supplied research. The exact interface, account requirements, regional eligibility, and operational limits may differ between consumer products and developer access. The model’s API pricing is documented in Chinese yuan, and the provider’s broader ecosystem is primarily oriented toward China.

Overall, Step 5 Preview is best understood as a high-capability preview model for long-running, text-producing work with optional visual input and tool integration. Its strongest differentiators are the 1-million-token context window, the stated focus on agentic and professional tasks, and its positioning at the top of StepFun’s current model lineup. Its main trade-offs are preview uncertainty, regional availability, unspecified maximum output and knowledge cutoff details, and the fact that it does not replace dedicated generative media models.


Answers to Frequently Asked Questions

Is Step 5 Preview open weight and publicly available?
Step 5 Preview is currently available through StepFun products and its API platform, subject to account and regional requirements. StepFun announced an open-weight release for October 15, 2026, but that future date does not mean the model weights are already available.
Can Step 5 Preview generate images, audio, or video?
No. Step 5 Preview supports text and vision input but produces text output only. The supplied specifications do not describe native image, audio, video, or music generation, so dedicated media models are more suitable for those tasks.
How much does Step 5 Preview cost?
According to the supplied StepFun pricing documentation, uncached input costs ¥7 per 1 million tokens, cached input costs ¥0.35 per 1 million tokens, and output costs ¥20 per 1 million tokens. Actual eligibility and caching behavior should be checked in the current platform documentation.
How large is Step 5 Preview's context window?
Step 5 Preview has a 1,000,000-token context window. This can support large codebases, extensive research materials, business documents, and extended agentic workflows, although the maximum output-token limit has not been publicly specified.
What is Step 5 Preview?
Step 5 Preview is a reasoning-oriented model from StepFun designed for long-context, agentic, software engineering, finance, research, and other professional tasks. It accepts text and vision input, produces text output, and supports tool use through compatible applications.


Sources 3
Provider

About StepFun