Model catalog

StepFun Models

Browse the AI models associated with StepFun. Compare current and historical models by family, capabilities, context window, availability and intended use.

12 models tracked
12 Total models
5 Model families
5 Model types
12 Current / accessible
All models

StepFun model catalog

Coding agents, software engineering, long-context reasoning, tool-using agents, private local inference, and work-centric automation

Type Reasoning
Context 262K
Reasoning 9/10
Speed 9/10
Tool use Streaming
Status

Current; open-weight model available for local deployment and accessible through StepFun's API platform

View model →

High-throughput coding agents, visual document and UI understanding, search-heavy research workflows, long-context analysis, and tool-using autonomous agents

Type Multimodal
Context 256K
Reasoning 8/10
Speed 9/10
Multimodal Image input Tool use
Status

Current and available

Input $0.20 per 1 million tokens for cache misses; $0.04 per 1 million tokens for cache hits
Output $1.15 per 1 million tokens
View model →

Long-running agentic workflows, software engineering, coding, professional knowledge work, financial analysis, research, and vision-assisted tasks

Type Reasoning
Context 1M
Reasoning 9/10
Speed 7/10
Multimodal Image input Tool use
Status

Current preview; available through StepFun products and API. Open-weight release planned for October 15, 2026.

Input ¥7 per 1 million uncached input tokens; ¥0.35 per 1 million cached input tokens
Output ¥20 per 1 million output tokens
View model →

Multilingual transcription, meetings, subtitles, live content, customer-service recordings, specialized terminology, noisy audio, dialects, code-switching, and singing or music transcription

Type Speech Recognition
Reasoning 5/10
Speed 9/10
Audio input Streaming
Status

Current and available through the StepFun API

Input 2.8 CNY per audio hour
View model →
StepFun logo
StepAudio 3

StepAudio 3 Gen

Zero-shot text-to-speech, natural-language voice design, singing and vocal generation, music, sound effects, ambience, and complete multi-element audio scenes

Type Other
Multimodal Audio input Media output
Status

Current research and platform-listed model; public API and access details are limited

View model →

Song generation, instrumental music, lyric-to-song workflows, accompaniment, cover-style synthesis, and rapid music prototyping

Type Other
Reasoning 2/10
Speed 3/10
Multimodal Audio input Media output
Status

Current; publicly showcased and available through StepFun's StepAudio 3 music experience

View model →

Natural realtime voice conversation, full-duplex interaction, interruption-aware assistants, emotional audio understanding and voice agents that use tools.

Type Realtime Audio
Reasoning 7/10
Speed 9/10
Multimodal Audio input Media output
Status

Current and publicly listed

View model →
StepFun logo
StepAudio 3

StepAudio 3 TTS

Controllable multilingual text-to-speech, narration, voice interfaces, localization, and expressive spoken-audio generation

Type Other
Reasoning 1/10
Speed 8/10
Media output Streaming
Status

Current and publicly accessible

View model →
StepFun logo
Step Edge

Step Edge

On-device visual assistants, GUI grounding, screen understanding, spatial reasoning, automotive interaction, and low-latency local agent workflows

Type Multimodal
Context 131K
Reasoning 7/10
Speed 8/10
Multimodal Image input Video input
Status

Current; officially announced as an edge-deployment base model

View model →

On-device speech recognition, audio understanding, voice assistants, in-vehicle interaction, privacy-sensitive audio processing, and low-latency edge AI.

Type Multimodal
Reasoning 6/10
Speed 8/10
Audio input
Status

Current; edge-deployment model with public capability information but no verified public hosted API pricing or general model-download documentation located.

View model →
StepFun logo
Step Edge

Step Edge Gen

On-device text-to-image generation, local image editing, privacy-sensitive creative features, and low-latency edge-device applications.

Type Other
Reasoning 1/10
Speed 9/10
Multimodal Image input Media output
Status

Current research model; public technical information is available, but official public API availability and downloadable weights were not verified.

View model →
StepFun logo
Step Edge

Step Edge GUI

Low-latency desktop and mobile GUI automation, visual grounding, local computer-use agents, and privacy-sensitive edge workflows

Type Other
Reasoning 6/10
Speed 9/10
Multimodal Image input Media output
Status

Current; edge-deployment GUI agent model

View model →