StepAudio 3

StepAudio 3 ASR Max

by StepFun · Current and available through the StepFun API

StepAudio 3 ASR Max is StepFun's specialist speech-recognition model for converting challenging audio into incremental or final text. It supports multilingual and dialectal speech, mixed Chinese-English audio, specialist terminology, noisy recordings, long audio, singing, and background music. The model is available through an HTTP and SSE API at 2.8 CNY per audio hour, but exact context limits, maximum output limits, speaker diarization, timestamps, and several advanced API features are not documented.

Text Reasoning Coding
StepAudio 3 ASR Max is the largest automatic speech-recognition model identified in StepFun's StepAudio 3 family. It accepts audio and returns text through an HTTP plus Server-Sent Events interface, making it suitable for both completed transcription jobs and applications that need partial results while audio is being processed. Its main distinction is the use of language-model context alongside acoustic recognition, which StepFun positions for multilingual, domain-specific, noisy, dialectal, and mixed-language speech.
Outputs

What StepAudio 3 ASR Max can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

5/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family StepAudio 3
Model type Speech Recognition
Status Current and available through the StepFun API
Knowledge cutoff notes

StepFun's model documentation does not publish a knowledge cutoff for StepAudio 3 ASR Max. The model uses acoustic signals together with language-model context and domain knowledge during recognition, but this should not be interpreted as a documented fixed knowledge-cutoff date.

Model notes

The StepFun catalog labels the product StepAudio 3 ASR, while the model-specific documentation identifies StepAudio 3 ASR Max as the largest ASR model in the series and documents its API behavior. The documented input is audio and the output consists of incremental and final text. The API uses HTTP plus Server-Sent Events through POST /v1/audio/asr/sse. StepFun describes support for Chinese, English, dialects, mixed Chinese-English speech, long audio, more than 20 professional domains, whispers, fast speech, unclear connected speech, environmental noise, singing, and background music. StepFun reports a 0.57% error rate on its contextual-reasoning evaluation, but this is a vendor-reported result rather than an independently verified universal accuracy figure. No official knowledge cutoff, context window, maximum output-token limit, fine-tuning, caching, batch API, JSON-mode, or speaker-diarization specification was located for the exact model. Editorial scores are comparative estimates for a specialist transcription model, not provider-published ratings.

Cost

Model pricing

Input 2.8 CNY per audio hour
Model guide

StepAudio 3 ASR Max for Context-Aware, Low-Latency Transcription

StepAudio 3 ASR Max is StepFun's speech-recognition model for turning audio into incremental or final text. It is designed for multilingual and dialect-rich speech, noisy recordings, specialist terminology, long audio, singing, and other difficult transcription conditions. The model is available through StepFun's HTTP and Server-Sent Events API at a listed price of 2.8 CNY per audio hour.

What StepAudio 3 ASR Max is

StepAudio 3 ASR Max is a specialist automatic speech-recognition model from StepFun. Its job is to convert spoken or sung audio into text, rather than to generate conversational replies, synthesize speech, create images, or produce general-purpose written content. The documented output is text, delivered either as incremental recognition results or as a final transcription.

The model belongs to StepFun's current StepAudio 3 voice-model family. StepFun's catalog labels the broader product as StepAudio 3 ASR, while the model-specific documentation identifies StepAudio 3 ASR Max as the largest automatic speech-recognition model in that series. This distinction matters when comparing catalog names with the API model documentation: the page is about the Max variant, not the StepAudio family as a whole.

StepAudio 3 ASR Max is intended for situations where ordinary speech recognition may struggle. StepFun describes support for Chinese, English, dialects, Chinese-English code-switching, long audio, professional terminology, whispers, fast speech, unclear connected speech, environmental noise, singing, and background music. These are provider-described capabilities rather than a guarantee that every recording will be transcribed accurately.

How its recognition approach is positioned

Speech recognition normally has to interpret both the sound signal and the likely meaning of the words. StepAudio 3 ASR Max is described as combining acoustic recognition with language-model context and domain knowledge. In practical terms, this means the model is positioned to use surrounding words and subject context when the audio is ambiguous, rather than treating every sound as an isolated unit.

This approach can be useful for meetings, customer-service recordings, technical discussions, subtitles, and other audio containing specialist vocabulary. It may also help with speech that is fast, partially unclear, mixed between Chinese and English, or affected by background sound. However, the supplied documentation does not publish a universal accuracy rate, a fixed knowledge cutoff, a context-window size, or a guaranteed maximum audio duration for the exact model.

StepFun reports a 0.57% error rate on its contextual-reasoning evaluation. This is a vendor-reported result for a particular evaluation, not an independently verified accuracy figure for every language, speaker, recording environment, or use case. It should therefore be treated as evidence of the provider's positioning rather than as a general transcription guarantee.

Supported inputs and outputs

CapabilityDocumented status
InputAudio
OutputIncremental and final text
StreamingSupported through the documented SSE endpoint
Image, video, or text inputNot documented for this model
Image, video, or audio outputNot supported as the model's documented output
Speaker diarization and timestampsNot verified for the documented SSE endpoint

The model's streaming behavior is particularly relevant for live or interactive workflows. Incremental results can be consumed as recognition progresses, while final text can be used after an utterance or recording is complete. The available research does not specify the exact event schema, latency target, punctuation behavior, timestamp format, or correction behavior for partial transcripts, so these details should be checked in the current StepFun documentation before implementation.

API access and pricing

StepAudio 3 ASR Max is available through StepFun's developer platform. The documented interface uses an HTTP POST request to /v1/audio/asr/sse and returns Server-Sent Events, commonly abbreviated as SSE. SSE is a web delivery method that lets a server send a sequence of updates over one open connection, which fits incremental transcription better than waiting for one response at the end.

The listed price is 2.8 CNY per audio hour. This is a duration-based price, so the relevant unit is the amount of audio submitted rather than the number of generated text tokens. The supplied research does not identify a separate output charge, free tier for this model, minimum billable duration, rounding rule, regional tax treatment, or enterprise discount. Those commercial details should be confirmed in the current StepFun platform pricing information.

Because the model is accessed through a provider API, an application will also need to handle authentication, audio upload or encoding, connection management, partial events, final results, and failures. The research confirms the protocol and endpoint but does not provide a complete current SDK example or the exact request parameters, so it would be unsafe to infer those implementation details here.

Strengths and practical use cases

StepAudio 3 ASR Max is most compelling when transcription quality across varied audio matters more than having a general-purpose language model generate an answer. Its documented areas of emphasis include:

  • Multilingual and mixed-language audio: Chinese, English, dialects, and Chinese-English speech are specifically described.
  • Difficult acoustic conditions: the provider cites environmental noise, whispers, fast speech, unclear connected speech, singing, and background music.
  • Specialist vocabulary: StepFun describes support for more than 20 professional domains, which may benefit technical, business, or service recordings.
  • Long-form transcription: long audio is listed among the target capabilities, although no exact maximum duration is published in the supplied documentation.
  • Low-latency workflows: incremental results through SSE can support live captions, monitoring, or interfaces that display a transcript while speech continues.

These characteristics make the model a reasonable candidate for meeting notes, subtitle drafts, call-center quality review, lecture or event transcription, live content, multilingual media processing, and speech or music analysis where text output is the main requirement. Human review may still be appropriate for regulated, legal, medical, or publication-ready transcripts, especially when names, numbers, terminology, or speaker attribution must be exact.

Limits and unsupported capabilities

StepAudio 3 ASR Max should not be selected simply because it is part of a broader multimodal provider ecosystem. The model itself is documented as an audio-in, text-out speech-recognition system. It is not documented as a conversational audio model, text-to-speech system, image or video generator, coding model, or general-purpose assistant.

Several important specifications remain unverified for the exact model. StepFun does not publish a knowledge cutoff, context length, maximum output-token limit, fine-tuning availability, prompt caching, batch API, JSON mode, or function and tool-use support in the supplied research. Since the model's output is transcription text rather than a general response, a token-oriented output limit may not be the main operational constraint, but no maximum should be assumed. Likewise, speaker diarization and timestamps are not verified for the documented SSE endpoint.

The model's documented emphasis on contextual recognition should not be confused with general reasoning. It uses language-model context to improve recognition and reports a provider evaluation in that area, but it is not intended to solve complex problems, write software, browse the web, or take actions through tools. Its coding score in the supplied comparison data is an editorial estimate of 1 out of 10, not a StepFun-published benchmark or product rating.

Speed, cost, and editorial evaluation

The supplied comparative assessment gives StepAudio 3 ASR Max a speed score of 9 out of 10 and a cost score of 8 out of 10. These are editorial estimates for a specialist transcription model, not provider-published ratings. They reflect the model's apparent fit for low-latency audio processing and its listed duration-based price, but they do not establish a guaranteed response time or a universal cost advantage.

At 2.8 CNY per audio hour, the model's economics are easiest to evaluate against the volume and urgency of a transcription workload. Streaming can reduce perceived waiting time for live applications, while batch-style processing may be simpler when users only need a completed transcript. Actual total cost can also depend on audio preparation, retries, storage, downstream review, and any platform terms not included in the supplied price information.

When to choose StepAudio 3 ASR Max

Choose StepAudio 3 ASR Max when the primary task is converting challenging audio into text and the workflow benefits from incremental results. It is especially well suited to Chinese or mixed Chinese-English speech, dialects, specialist terminology, noisy recordings, long-form audio, and recordings that include singing or background music. The model is also a sensible option when a duration-based API price is easier to manage than token-based billing.

Consider another option when the application needs a verified speaker-diarization pipeline, reliable timestamps, structured JSON output, speech synthesis, real-time spoken dialogue, web search, tool calls, or general reasoning after transcription. A broader speech or multimodal system may be more appropriate when recognition and conversation must happen in one model. A separate post-processing or language model may be preferable when the transcript must be summarized, classified, translated, or converted into a strict business schema, because those capabilities are not documented for StepAudio 3 ASR Max itself.

Bottom line

StepAudio 3 ASR Max is a focused StepFun model for difficult and potentially low-latency speech transcription. Its strongest documented differentiators are contextual recognition, multilingual and dialect coverage, handling of challenging audio, specialist-domain support, and incremental text delivery through SSE. Its main uncertainties are equally important: exact input limits, output limits, diarization, timestamps, structured-output support, and several API features are not published in the supplied research. It is therefore best evaluated as a dedicated transcription component, not as a replacement for a general conversational or multimodal model.


Answers to Frequently Asked Questions

Does StepAudio 3 ASR Max provide speaker diarization and timestamps?
Speaker diarization and timestamps are not verified for the documented SSE endpoint. Applications that require reliable speaker labels or time-coded transcripts should confirm current StepFun documentation or use a separate processing solution.
What languages and audio conditions does StepAudio 3 ASR Max support?
StepFun describes support for Chinese, English, dialects, and Chinese-English code-switching. The model is also positioned for long audio, professional terminology, whispers, fast speech, unclear connected speech, environmental noise, singing, and background music. These are provider-described capabilities, not guaranteed accuracy levels for every recording.
How much does StepAudio 3 ASR Max cost?
The listed price is 2.8 CNY per audio hour. The available information does not specify whether there is a free tier, minimum billable duration, rounding rule, separate output charge, regional tax treatment, or enterprise discount.
Does StepAudio 3 ASR Max support real-time transcription?
Yes. StepAudio 3 ASR Max supports incremental transcription through StepFun's documented Server-Sent Events (SSE) endpoint, /v1/audio/asr/sse. This can support live captions and other low-latency workflows, although the exact latency, event schema, and partial-transcript correction behavior are not specified.
What is StepAudio 3 ASR Max used for?
StepAudio 3 ASR Max is an automatic speech-recognition model from StepFun that converts spoken or sung audio into incremental or final text. It is designed for challenging recordings, including Chinese and English speech, dialects, code-switching, specialist terminology, long audio, whispers, fast speech, noise, singing, and background music.


Sources 4
Provider

About StepFun