Coding agents, software engineering, long-context reasoning, tool-using agents, private local inference, and work-centric automation
Type
Reasoning
Context
262K
Reasoning
9/10
Speed
9/10
Tool use
Streaming
Status
Current; open-weight model available for local deployment and accessible through StepFun's API platform
View model
→
High-throughput coding agents, visual document and UI understanding, search-heavy research workflows, long-context analysis, and tool-using autonomous agents
Type
Multimodal
Context
256K
Reasoning
8/10
Speed
9/10
Multimodal
Image input
Tool use
Status
Current and available
Input
$0.20 per 1 million tokens for cache misses; $0.04 per 1 million tokens for cache hits
Output
$1.15 per 1 million tokens
View model
→
Long-running agentic workflows, software engineering, coding, professional knowledge work, financial analysis, research, and vision-assisted tasks
Type
Reasoning
Context
1M
Reasoning
9/10
Speed
7/10
Multimodal
Image input
Tool use
Status
Current preview; available through StepFun products and API. Open-weight release planned for October 15, 2026.
Input
¥7 per 1 million uncached input tokens; ¥0.35 per 1 million cached input tokens
Output
Â¥20 per 1 million output tokens
View model
→
Multilingual transcription, meetings, subtitles, live content, customer-service recordings, specialized terminology, noisy audio, dialects, code-switching, and singing or music transcription
Type
Speech Recognition
Reasoning
5/10
Speed
9/10
Audio input
Streaming
Status
Current and available through the StepFun API
Input
2.8 CNY per audio hour
View model
→
Zero-shot text-to-speech, natural-language voice design, singing and vocal generation, music, sound effects, ambience, and complete multi-element audio scenes
Multimodal
Audio input
Media output
Status
Current research and platform-listed model; public API and access details are limited
View model
→
Song generation, instrumental music, lyric-to-song workflows, accompaniment, cover-style synthesis, and rapid music prototyping
Type
Other
Reasoning
2/10
Speed
3/10
Multimodal
Audio input
Media output
Status
Current; publicly showcased and available through StepFun's StepAudio 3 music experience
View model
→
Natural realtime voice conversation, full-duplex interaction, interruption-aware assistants, emotional audio understanding and voice agents that use tools.
Type
Realtime Audio
Reasoning
7/10
Speed
9/10
Multimodal
Audio input
Media output
Status
Current and publicly listed
View model
→
Controllable multilingual text-to-speech, narration, voice interfaces, localization, and expressive spoken-audio generation
Type
Other
Reasoning
1/10
Speed
8/10
Media output
Streaming
Status
Current and publicly accessible
View model
→
On-device visual assistants, GUI grounding, screen understanding, spatial reasoning, automotive interaction, and low-latency local agent workflows
Type
Multimodal
Context
131K
Reasoning
7/10
Speed
8/10
Multimodal
Image input
Video input
Status
Current; officially announced as an edge-deployment base model
View model
→
On-device speech recognition, audio understanding, voice assistants, in-vehicle interaction, privacy-sensitive audio processing, and low-latency edge AI.
Type
Multimodal
Reasoning
6/10
Speed
8/10
Audio input
Status
Current; edge-deployment model with public capability information but no verified public hosted API pricing or general model-download documentation located.
View model
→
On-device text-to-image generation, local image editing, privacy-sensitive creative features, and low-latency edge-device applications.
Type
Other
Reasoning
1/10
Speed
9/10
Multimodal
Image input
Media output
Status
Current research model; public technical information is available, but official public API availability and downloadable weights were not verified.
View model
→
Low-latency desktop and mobile GUI automation, visual grounding, local computer-use agents, and privacy-sensitive edge workflows
Type
Other
Reasoning
6/10
Speed
9/10
Multimodal
Image input
Media output
Status
Current; edge-deployment GUI agent model
View model
→