What StepAudio 3 ASR Max is
StepAudio 3 ASR Max is a specialist automatic speech-recognition model from StepFun. Its job is to convert spoken or sung audio into text, rather than to generate conversational replies, synthesize speech, create images, or produce general-purpose written content. The documented output is text, delivered either as incremental recognition results or as a final transcription.
The model belongs to StepFun's current StepAudio 3 voice-model family. StepFun's catalog labels the broader product as StepAudio 3 ASR, while the model-specific documentation identifies StepAudio 3 ASR Max as the largest automatic speech-recognition model in that series. This distinction matters when comparing catalog names with the API model documentation: the page is about the Max variant, not the StepAudio family as a whole.
StepAudio 3 ASR Max is intended for situations where ordinary speech recognition may struggle. StepFun describes support for Chinese, English, dialects, Chinese-English code-switching, long audio, professional terminology, whispers, fast speech, unclear connected speech, environmental noise, singing, and background music. These are provider-described capabilities rather than a guarantee that every recording will be transcribed accurately.
How its recognition approach is positioned
Speech recognition normally has to interpret both the sound signal and the likely meaning of the words. StepAudio 3 ASR Max is described as combining acoustic recognition with language-model context and domain knowledge. In practical terms, this means the model is positioned to use surrounding words and subject context when the audio is ambiguous, rather than treating every sound as an isolated unit.
This approach can be useful for meetings, customer-service recordings, technical discussions, subtitles, and other audio containing specialist vocabulary. It may also help with speech that is fast, partially unclear, mixed between Chinese and English, or affected by background sound. However, the supplied documentation does not publish a universal accuracy rate, a fixed knowledge cutoff, a context-window size, or a guaranteed maximum audio duration for the exact model.
StepFun reports a 0.57% error rate on its contextual-reasoning evaluation. This is a vendor-reported result for a particular evaluation, not an independently verified accuracy figure for every language, speaker, recording environment, or use case. It should therefore be treated as evidence of the provider's positioning rather than as a general transcription guarantee.
Supported inputs and outputs
| Capability | Documented status |
|---|---|
| Input | Audio |
| Output | Incremental and final text |
| Streaming | Supported through the documented SSE endpoint |
| Image, video, or text input | Not documented for this model |
| Image, video, or audio output | Not supported as the model's documented output |
| Speaker diarization and timestamps | Not verified for the documented SSE endpoint |
The model's streaming behavior is particularly relevant for live or interactive workflows. Incremental results can be consumed as recognition progresses, while final text can be used after an utterance or recording is complete. The available research does not specify the exact event schema, latency target, punctuation behavior, timestamp format, or correction behavior for partial transcripts, so these details should be checked in the current StepFun documentation before implementation.
API access and pricing
StepAudio 3 ASR Max is available through StepFun's developer platform. The documented interface uses an HTTP POST request to /v1/audio/asr/sse and returns Server-Sent Events, commonly abbreviated as SSE. SSE is a web delivery method that lets a server send a sequence of updates over one open connection, which fits incremental transcription better than waiting for one response at the end.
The listed price is 2.8 CNY per audio hour. This is a duration-based price, so the relevant unit is the amount of audio submitted rather than the number of generated text tokens. The supplied research does not identify a separate output charge, free tier for this model, minimum billable duration, rounding rule, regional tax treatment, or enterprise discount. Those commercial details should be confirmed in the current StepFun platform pricing information.
Because the model is accessed through a provider API, an application will also need to handle authentication, audio upload or encoding, connection management, partial events, final results, and failures. The research confirms the protocol and endpoint but does not provide a complete current SDK example or the exact request parameters, so it would be unsafe to infer those implementation details here.
Strengths and practical use cases
StepAudio 3 ASR Max is most compelling when transcription quality across varied audio matters more than having a general-purpose language model generate an answer. Its documented areas of emphasis include:
- Multilingual and mixed-language audio: Chinese, English, dialects, and Chinese-English speech are specifically described.
- Difficult acoustic conditions: the provider cites environmental noise, whispers, fast speech, unclear connected speech, singing, and background music.
- Specialist vocabulary: StepFun describes support for more than 20 professional domains, which may benefit technical, business, or service recordings.
- Long-form transcription: long audio is listed among the target capabilities, although no exact maximum duration is published in the supplied documentation.
- Low-latency workflows: incremental results through SSE can support live captions, monitoring, or interfaces that display a transcript while speech continues.
These characteristics make the model a reasonable candidate for meeting notes, subtitle drafts, call-center quality review, lecture or event transcription, live content, multilingual media processing, and speech or music analysis where text output is the main requirement. Human review may still be appropriate for regulated, legal, medical, or publication-ready transcripts, especially when names, numbers, terminology, or speaker attribution must be exact.
Limits and unsupported capabilities
StepAudio 3 ASR Max should not be selected simply because it is part of a broader multimodal provider ecosystem. The model itself is documented as an audio-in, text-out speech-recognition system. It is not documented as a conversational audio model, text-to-speech system, image or video generator, coding model, or general-purpose assistant.
Several important specifications remain unverified for the exact model. StepFun does not publish a knowledge cutoff, context length, maximum output-token limit, fine-tuning availability, prompt caching, batch API, JSON mode, or function and tool-use support in the supplied research. Since the model's output is transcription text rather than a general response, a token-oriented output limit may not be the main operational constraint, but no maximum should be assumed. Likewise, speaker diarization and timestamps are not verified for the documented SSE endpoint.
The model's documented emphasis on contextual recognition should not be confused with general reasoning. It uses language-model context to improve recognition and reports a provider evaluation in that area, but it is not intended to solve complex problems, write software, browse the web, or take actions through tools. Its coding score in the supplied comparison data is an editorial estimate of 1 out of 10, not a StepFun-published benchmark or product rating.
Speed, cost, and editorial evaluation
The supplied comparative assessment gives StepAudio 3 ASR Max a speed score of 9 out of 10 and a cost score of 8 out of 10. These are editorial estimates for a specialist transcription model, not provider-published ratings. They reflect the model's apparent fit for low-latency audio processing and its listed duration-based price, but they do not establish a guaranteed response time or a universal cost advantage.
At 2.8 CNY per audio hour, the model's economics are easiest to evaluate against the volume and urgency of a transcription workload. Streaming can reduce perceived waiting time for live applications, while batch-style processing may be simpler when users only need a completed transcript. Actual total cost can also depend on audio preparation, retries, storage, downstream review, and any platform terms not included in the supplied price information.
When to choose StepAudio 3 ASR Max
Choose StepAudio 3 ASR Max when the primary task is converting challenging audio into text and the workflow benefits from incremental results. It is especially well suited to Chinese or mixed Chinese-English speech, dialects, specialist terminology, noisy recordings, long-form audio, and recordings that include singing or background music. The model is also a sensible option when a duration-based API price is easier to manage than token-based billing.
Consider another option when the application needs a verified speaker-diarization pipeline, reliable timestamps, structured JSON output, speech synthesis, real-time spoken dialogue, web search, tool calls, or general reasoning after transcription. A broader speech or multimodal system may be more appropriate when recognition and conversation must happen in one model. A separate post-processing or language model may be preferable when the transcript must be summarized, classified, translated, or converted into a strict business schema, because those capabilities are not documented for StepAudio 3 ASR Max itself.
Bottom line
StepAudio 3 ASR Max is a focused StepFun model for difficult and potentially low-latency speech transcription. Its strongest documented differentiators are contextual recognition, multilingual and dialect coverage, handling of challenging audio, specialist-domain support, and incremental text delivery through SSE. Its main uncertainties are equally important: exact input limits, output limits, diarization, timestamps, structured-output support, and several API features are not published in the supplied research. It is therefore best evaluated as a dedicated transcription component, not as a replacement for a general conversational or multimodal model.
Answers to Frequently Asked Questions
/v1/audio/asr/sse. This can support live captions and other low-latency workflows, although the exact latency, event schema, and partial-transcript correction behavior are not specified.
