Step 3

Step 3.5 Flash

by StepFun · Current; open-weight model available for local deployment and accessible through StepFun's API platform

Step 3.5 Flash is StepFun's open-weight sparse MoE model for coding, long-context reasoning, tool use, and agentic workflows. It has approximately 196.81B total parameters, 11B active parameters per token, a 256K context window, Multi-Token Prediction, and Apache 2.0 licensing. The model supports local inference through several frameworks and is also available through StepFun's platform, but verified API pricing and a separate maximum output limit were not supplied.

Text Reasoning Coding
Step 3.5 Flash is an open-weight foundation model from StepFun aimed at software engineering, long-context reasoning, tool-using agents, and work-focused automation. Its sparse architecture is designed to keep inference efficient: the model contains approximately 196.81 billion parameters, but activates about 11 billion for each token. It can be downloaded and run locally with several inference frameworks, while also being available through StepFun's platform.
Outputs

What Step 3.5 Flash can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Step 3
Model type Reasoning
Context window 262K tokens
Release date 2026-02-12
Status Current; open-weight model available for local deployment and accessible through StepFun's API platform
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date for the exact model was identified in the reviewed official model card, repository, launch material, or platform documentation.

Model notes

Step 3.5 Flash is a sparse Mixture-of-Experts transformer with approximately 196.81B total parameters and approximately 11B active parameters per token. Official model materials describe 45 transformer layers, a 128,896-token vocabulary, a 3:1 sliding-window-attention design, and a 256K context window. The model uses Multi-Token Prediction technology and is reported by StepFun to reach roughly 100–300 tokens per second in typical usage, with peaks up to 350 tokens per second for some single-stream coding workloads. Official benchmark claims include 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0. Tool calling is supported in the official deployment guidance through StepFun-specific reasoning and tool-call parsers. The model is released under the Apache 2.0 license and can be deployed with Transformers, vLLM, SGLang, and llama.cpp. Exact current per-token API input and output prices were not verified in an official public pricing source. Reasoning, coding, speed, and cost scores are editorial comparative estimates rather than vendor ratings. StepFun's official materials identify the model as a foundation language model; no native image, audio, or video input or output is documented.

Model guide

Step 3.5 Flash: Fast Open-Weight Reasoning for Coding and Agents

Step 3.5 Flash is StepFun's open-weight sparse Mixture-of-Experts language model for fast reasoning, coding, tool use, and agentic workflows. It combines approximately 196.81 billion total parameters with about 11 billion active parameters per token, a 256K context window, multi-token prediction, and Apache 2.0 licensing for local deployment.

What is Step 3.5 Flash?

Step 3.5 Flash is a text-based reasoning and coding model provided by StepFun. It is an open-weight model rather than a closed hosted-only service: StepFun publishes downloadable weights and deployment guidance, allowing organizations and technically capable users to run it on their own infrastructure. The model is released under the Apache 2.0 license.

The model is built for tasks that require more than short conversational answers. Its intended uses include software engineering, code generation, debugging, multi-step reasoning, tool calling, and long-running agent workflows. StepFun positions it within the current Step model lineup alongside models such as Step 5 Preview and Step 3.7 Flash, but Step 3.5 Flash's clearest distinction is its combination of open deployment, high-throughput generation, and emphasis on coding and agentic work.

StepFun announced the model on February 12, 2026. The official model materials identify it as a foundation language model; they do not document native image, audio, or video understanding or generation for this model.

Architecture and core specifications

Step 3.5 Flash uses a sparse Mixture-of-Experts, or MoE, architecture. In an MoE model, different subsets of the network can handle different tokens or tasks. This means the total network can be very large without activating every parameter for every generated token. For Step 3.5 Flash, the published figures are approximately 196.81 billion total parameters and approximately 11 billion active parameters per token.

That distinction matters in practice. The total parameter count indicates the scale of the full model, while the active count is more closely related to the computation required for an individual token. It does not make local deployment lightweight by default: storing the weights and serving the model still requires substantial hardware, especially at full precision or with multiple concurrent users. However, sparse activation can improve the relationship between model capacity and generation throughput.

Official materials also describe 45 transformer layers, a 128,896-token vocabulary, and a 3:1 sliding-window-attention design. The model supports a 256K-token context window, equivalent to 262,144 tokens in the supplied model metadata. A context window includes the prompt, conversation history, tool messages, and generated content that the serving system keeps available. The research does not specify a separate maximum output-token limit, so no independent output ceiling should be assumed beyond the limits imposed by the deployment or API configuration.

SpecificationReported detail
ProviderStepFun
Model familyStep 3
Model typeReasoning and coding language model
Total parametersApproximately 196.81 billion
Active parameters per tokenApproximately 11 billion
Context window256K tokens
LicenseApache 2.0
Input and outputText input and text output
Native image, audio, or video supportNot documented for this model

Reasoning, coding, and agent workflows

Step 3.5 Flash is primarily intended for problems where the model must maintain a plan, inspect intermediate information, or make several decisions before producing an answer. Examples include analyzing a software repository, proposing a sequence of code changes, diagnosing a failing test, or deciding which external tool to call next.

Its coding focus is supported by StepFun's reported benchmark results, including 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0. These are provider-reported claims rather than independent evaluations supplied in the research, so they should be treated as directional evidence rather than a guarantee of performance on a particular codebase. Real results will depend on the prompt, repository complexity, tool environment, test coverage, and serving configuration.

The model's tool-use support is especially relevant to agents. Official deployment guidance describes StepFun-specific reasoning and tool-call parsers, which are used to help a serving system distinguish ordinary reasoning or text from structured tool requests. This can support workflows such as reading files, running tests, querying a service, or taking other actions exposed by an orchestration layer. The model itself does not automatically provide those external tools; the application must define, authorize, execute, and return tool results.

Step 3.5 Flash also uses Multi-Token Prediction technology. Rather than predicting only one next token at a time during generation, this approach is intended to improve decoding efficiency. StepFun reports typical generation speeds of roughly 100 to 300 tokens per second, with peaks of up to 350 tokens per second for some single-stream coding workloads. These figures are provider claims and should not be treated as a universal speed guarantee. Hardware, quantization, batch size, context length, framework, and concurrency can all change observed throughput.

Context, modalities, and output limits

The 256K context window is one of the model's most useful specifications for practical development work. It can accommodate large code files, extensive logs, repository excerpts, long technical documents, and extended tool traces, subject to the serving framework's own memory and request limits. A larger context does not automatically mean that every detail will receive equal attention, so targeted retrieval and concise tool results may still improve reliability.

Step 3.5 Flash is text-only in the documented model configuration. It accepts text input and produces text output. There is no supplied evidence that this model directly accepts images, audio, or video, and it should not be confused with StepFun's broader ecosystem, which includes separate multimodal and media-generation services.

The research does not verify a maximum output-token value, structured-output guarantee, JSON mode, prompt caching, batch API, or fine-tuning availability for this exact model. Those fields should remain separate from general tool-calling support. In particular, the existence of a tool-call parser does not by itself prove that every deployment offers a general-purpose JSON mode.

Deployment and availability

Step 3.5 Flash is available through StepFun's official repositories and Hugging Face organization. The supplied deployment materials identify support for Transformers, vLLM, SGLang, and llama.cpp. These options cover different operating patterns, from direct model experimentation to production-oriented serving and local or quantized inference.

Users who need an easier hosted route can access the model through StepFun's platform, although the exact current API pricing was not verified in an official public pricing source supplied for this review. Therefore, there is no reliable per-token input or output price to report. The model's Apache 2.0 license can make self-hosting attractive for organizations that want more control over deployment, but infrastructure, hardware, engineering, monitoring, and maintenance costs still apply.

Local deployment also requires attention to memory requirements and compatibility. The approximately 196.81-billion-parameter total network is not comparable to an 11-billion-parameter dense model in terms of storage or operational simplicity. Quantization and optimized inference engines may reduce practical resource demands, but the research does not establish a single hardware requirement or performance level.

Main strengths and limitations

Where Step 3.5 Flash is strong

  • Open deployment: Downloadable weights and Apache 2.0 licensing provide more control than a hosted-only model.
  • Coding and software engineering: The model is designed for code generation, debugging, repository work, and terminal-oriented agent tasks.
  • Long context: The 256K-token window is suitable for large prompts, codebases, logs, and extended tool traces.
  • Agent integration: Official parser guidance supports reasoning and tool-call workflows.
  • Generation throughput: Multi-token prediction and sparse activation are intended to improve speed and efficiency.
  • Deployment flexibility: Transformers, vLLM, SGLang, and llama.cpp are identified as supported inference paths.

Where it is less suitable

  • Text-only operation: It is not the appropriate choice for direct image, audio, or video input and output.
  • Operational complexity: A large sparse model still requires substantial infrastructure and expertise for reliable local serving.
  • Unverified commercial pricing: Current official per-token API prices were not confirmed in the supplied research.
  • No verified output ceiling: The context window is documented, but a separate maximum output-token limit is not.
  • Benchmark uncertainty: The cited benchmark figures come from StepFun and may not predict performance in every environment.
  • Potential task variability: Highly specialized domains or very long multi-turn conversations may require a model with more consistently validated behavior for that workload.

When to choose Step 3.5 Flash

Choose Step 3.5 Flash when you want an open-weight model for coding agents, private inference, long-context software work, or tool-using automation. It is particularly compelling when deployment control matters and your team can operate the required serving infrastructure. A developer building an internal coding assistant, repository-analysis system, or terminal agent could use the model to generate patches, inspect test output, and coordinate a defined set of tools.

It is also a reasonable option when throughput matters more than using the largest possible dense model. The sparse architecture, active-parameter design, and reported decoding speed aim to provide a practical balance between reasoning capacity and generation cost. That balance is an intended design advantage, not a guaranteed cost result: the actual economics depend on hardware utilization, quantization, concurrency, and whether the model is self-hosted or accessed through a paid platform.

Another option may be more appropriate if the application needs native visual understanding, speech, media generation, a verified structured-output contract, or a clearly published API price. A hosted model may also be preferable for teams that do not want to manage large-model infrastructure. Conversely, a smaller dense model may be easier to deploy for simple classification, short answers, or low-resource applications where Step 3.5 Flash's reasoning and context capacity would be unnecessary.

Bottom line

Step 3.5 Flash is a specialized open-weight choice for users who value coding ability, long context, tool use, and deployment control. Its approximately 196.81-billion-parameter sparse MoE design activates about 11 billion parameters per token, while Multi-Token Prediction targets faster generation. The documented 256K context window and Apache 2.0 license strengthen its appeal for private engineering and agentic workloads.

Its trade-offs are equally important: it is text-only, large to operate despite sparse activation, and lacks a verified public price and documented maximum output limit in the supplied research. The best evaluation should therefore combine StepFun's reported benchmarks with tests on the user's own repositories, tools, hardware, and concurrency requirements.


Answers to Frequently Asked Questions

Does Step 3.5 Flash support images, audio, or video?
No native image, audio, or video support is documented for Step 3.5 Flash. The model is configured for text input and text output, so applications requiring direct multimodal understanding or media generation should consider another model.
How can Step 3.5 Flash be deployed?
Step 3.5 Flash can be accessed through StepFun's platform or deployed from its official repositories and Hugging Face organization. The documented inference options include Transformers, vLLM, SGLang, and llama.cpp. Local deployment requires substantial infrastructure because of the model's large total parameter count.
Is Step 3.5 Flash suitable for coding agents and tool use?
Yes. Step 3.5 Flash is intended for coding agents, repository analysis, debugging, terminal workflows, and multi-step reasoning. StepFun provides reasoning and tool-call parser guidance, but applications must define, authorize, execute, and return results from external tools.
What are the main specifications of Step 3.5 Flash?
Step 3.5 Flash uses a sparse Mixture-of-Experts architecture with approximately 196.81 billion total parameters and about 11 billion active parameters per token. It has a 256K-token context window, 45 transformer layers, a 128,896-token vocabulary, and supports text input and text output.
What is Step 3.5 Flash?
Step 3.5 Flash is an open-weight text-based reasoning and coding model from StepFun. It is designed for software engineering, code generation, debugging, tool calling, and long-running agent workflows, and is released under the Apache 2.0 license.


Sources 6
Provider

About StepFun