StepAudio 3

StepAudio 3 Realtime

by StepFun · Current and publicly listed

StepAudio 3 Realtime is StepFun’s specialized model for natural spoken interaction. It focuses on full-duplex conversation, interruption handling, emotional and paralinguistic audio understanding, realtime reasoning, streaming speech output and tool-enabled voice agents. Public documentation does not clearly confirm its exact pricing, API identifier, context length or maximum output limit.

Speech Reasoning
StepAudio 3 Realtime is a voice-first model from StepFun for applications that need to listen and respond during an ongoing conversation. Unlike a conventional speech pipeline that waits for a person to finish speaking before generating a reply, it is designed for overlapping turns, interruptions, natural pauses, expressive audio understanding, and low-latency responses. Its strongest fit is conversational voice agents rather than text-only generation, image creation, transcription-only processing, or music generation.
Outputs

What StepAudio 3 Realtime can produce

Speech
Inputs

What it can understand

Audio Multimodal input
Capabilities

Supported features

Tool use Streaming Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
9/10 Speed
Specifications

Technical details

Model family StepAudio 3
Model type Realtime Audio
Status Current and publicly listed
Knowledge cutoff notes

No authoritative knowledge-cutoff date for the exact StepAudio 3 Realtime model was found in the reviewed StepFun sources.

Model notes

StepFun presents StepAudio 3 Realtime as a specialized realtime voice-interaction model with natural turn-taking, full-duplex interaction, interruption handling, audio and emotional understanding, realtime reasoning and tool-enabled task completion. The official model page reports vendor benchmarks including 98.9 on AA Full-Duplex Bench, 56.0 on Ï„-Voice, 90.6 on MMSU, 86.5 on MMAR and 83.0 on GPQA Diamond. Exact public API model ID, pricing, context length, maximum output length, knowledge cutoff and detailed modality schema were not clearly published in the reviewed first-party sources. Reasoning, speed and other editorial scores are comparative estimates, not vendor specifications.

Model guide

StepAudio 3 Realtime: StepFun’s Full-Duplex Voice Model for Natural Conversation

StepAudio 3 Realtime is StepFun’s specialized realtime voice-interaction model for low-latency, full-duplex spoken conversations. It is designed to handle interruptions, turn-taking, emotional and paralinguistic cues, realtime reasoning, and tool-enabled voice-agent workflows. StepFun publicly positions it as a current model, but detailed API identifiers, pricing, context limits, and maximum output limits are not clearly documented in the reviewed sources.

What is StepAudio 3 Realtime?

StepAudio 3 Realtime is StepFun’s realtime voice-interaction model. It belongs to the StepAudio 3 family, alongside models and services associated with speech recognition, voice generation, text-to-speech, and music generation. The Realtime variant is specifically positioned for spoken conversations in which the system must listen, interpret audio, decide how to respond, and produce speech with minimal delay.

The defining distinction is its focus on full-duplex interaction. Full-duplex communication allows both sides of a conversation to speak in overlapping turns. In practical use, a person can interrupt the model, add a correction, or change the subject without waiting for the model to finish a complete response. This is closer to ordinary conversation than a strict listen-then-speak sequence.

Where it fits in StepFun’s lineup

StepAudio 3 Realtime is a specialist model within StepFun’s broader catalog, not a general replacement for every Step model. StepFun’s current ecosystem includes text and reasoning models, multimodal systems, image and video generation, speech services, and other StepAudio 3 variants. The Realtime model is aimed at the voice-interaction part of that lineup.

StepFun lists the model on its official model site and open platform. The reviewed sources describe it as currently available through StepFun’s platform context, although the exact public API model identifier and complete integration specification were not clearly exposed. Developers should verify the current identifier, access requirements, regional availability, and interface details directly in StepFun’s documentation before implementation.

How the realtime conversation experience works

Full-duplex interaction and interruptions

In a conventional voice assistant, the system may wait for an end-of-speech signal, transcribe the utterance, generate text, synthesize speech, and then play the answer. That approach can create noticeable pauses and makes interruptions awkward. StepFun presents StepAudio 3 Realtime as a model designed to reduce those conversational boundaries.

Its full-duplex positioning means an application can support more natural turn-taking. A user may begin speaking while the model is responding, stop the response, or add information during the exchange. This is particularly relevant for customer-service assistants, hands-free interfaces, and agents that need to clarify a task interactively rather than deliver a single uninterrupted answer.

Audio and paralinguistic understanding

The model is intended to interpret more than the literal words in an audio stream. StepFun highlights understanding of hesitation, laughter, pauses, vocal emotion, and other paralinguistic signals. These cues can help a voice agent distinguish between a confident request, uncertainty, frustration, amusement, or a user attempting to take back control of the conversation.

Such capabilities should be treated as provider-described behavior rather than a guarantee that every emotional or conversational cue will be interpreted correctly. Production systems should test the model with the accents, languages, background noise, speaking styles, interruptions, and accessibility scenarios expected in the target environment.

Reasoning and tool-enabled tasks

StepFun also positions StepAudio 3 Realtime as capable of reasoning while speaking. The intended behavior is not simply to convert speech to text and read a prerecorded answer. The model is described as responding quickly to straightforward requests, spending more time on difficult questions, and using tools when a conversation needs an external action or result.

Tool use is therefore one of the model’s stated capabilities, but the supplied public information does not define a complete function schema, supported tool protocol, or exact API calling format. A developer building a voice agent should confirm how tools are declared, how interruptions affect an in-progress tool call, and whether the model returns structured events or another platform-specific format.

What performance has StepFun reported?

StepFun’s official model page reports results across audio understanding, full-duplex interaction, voice-agent performance, general reasoning, and speech-related evaluations. The published figures include 98.9 on the AA Full-Duplex Bench, 56.0 on the τ-Voice benchmark, 90.6 on MMSU, 86.5 on MMAR, and 83.0 on GPQA Diamond.

These are vendor-reported results. They provide useful context for how StepFun presents the model, but they are not independent verification. Benchmark names, test conditions, prompting methods, and scoring procedures can materially affect results. They should not be treated as a substitute for testing the exact realtime workflow, audio conditions, response latency, interruption behavior, and tool calls required by a production application.

Supported modalities and technical scope

The model’s central modality is audio. The supplied model record identifies audio input and audio output, with direct speech output as its primary result. It also records streaming support and tool use. This makes it suitable for realtime speech-to-speech interaction rather than a batch-only audio-processing workflow.

CapabilityWhat is supported or known
Primary inputAudio input is identified for the model.
Primary outputAudio and speech output are identified; direct realtime voice responses are the main purpose.
StreamingStreaming is identified as supported.
Full-duplex interactionPromoted by StepFun as a central capability.
Tool useIdentified as supported, although the public tool schema was not clearly documented in the reviewed sources.
Image and video outputNot identified for this model.
Context lengthNot publicly documented in the reviewed sources.
Maximum output lengthNot publicly documented in the reviewed sources.
Knowledge cutoffNo authoritative date was found for this exact model.
PricingModel-specific token or realtime pricing was not clearly published in the reviewed sources.

The record does not clearly establish whether text input or text output is exposed as a standalone interface for this exact model, so those fields should not be assumed. The same caution applies to JSON or structured output, fine-tuning, caching, batch processing, and other conventional language-model features.

Main strengths and trade-offs

StepAudio 3 Realtime’s main strength is its specialization. Applications that depend on natural spoken interaction may benefit more from full-duplex behavior, interruption handling, and audio-aware responses than from a text model placed behind separate speech-recognition and speech-synthesis services.

  • Natural turn-taking: overlapping speech and interruption handling are central to the model’s positioning.
  • Low-latency interaction: StepFun presents the model for realtime responses rather than delayed batch processing.
  • Audio awareness: the model is intended to use emotional and paralinguistic signals in addition to words.
  • Agent workflows: tool-enabled task completion is part of the stated use case.
  • Streaming voice output: the model record identifies streaming and speech output, which are important for responsive interfaces.

These advantages come with trade-offs. A specialized realtime voice model is not automatically the best choice for long-form writing, image or video generation, transcription-only workloads, or applications that need fully documented token economics. The lack of clearly published pricing and context or output limits also makes capacity planning more difficult.

Best use cases

StepAudio 3 Realtime is a strong candidate for applications where the quality of the live conversation matters more than conventional text-model features:

  • Voice assistants that need interruption-aware dialogue
  • Customer-service and support agents that can ask and answer follow-up questions
  • Hands-free task assistants for devices, workplaces, or vehicles
  • Emotion-aware spoken interfaces that respond to vocal hesitation or frustration
  • Interactive agents that call tools during a conversation
  • Low-latency voice interfaces embedded in applications or hardware

For example, a support agent could listen to a customer’s explanation, detect that the customer is still speaking, avoid talking over them unnecessarily, and use a business tool after the request is sufficiently clear. The exact quality of this experience will depend on the application’s interruption policy, audio transport, tool orchestration, and safeguards as well as on the model itself.

When to choose StepAudio 3 Realtime

Choose this model when realtime spoken conversation is the primary product requirement and you need more than a simple speech-to-text and text-to-speech chain. It is especially relevant when users may interrupt, when vocal delivery carries useful information, or when an agent must complete tasks through tools while maintaining a live dialogue.

Consider another option when the workload is mainly text generation, coding, image or video creation, music generation, or transcription without interactive dialogue. A conventional text model may be easier to evaluate for document and coding tasks, while a dedicated transcription service may be more appropriate when accurate transcripts are the only required output. StepFun’s other StepAudio 3 variants may also be a better fit for speech recognition, text-to-speech, voice generation, or music-specific workloads.

Cost should be evaluated carefully. StepFun’s public materials reviewed here do not provide a confirmed model-specific price, so no reliable cost comparison can be made against other realtime voice systems. Before deployment, measure end-to-end latency, audio quality, interruption recovery, tool-call reliability, concurrent-session capacity, and the provider’s current billing rules rather than relying only on benchmark scores.

Limitations and verification checklist

The most important limitation is incomplete public technical documentation for the exact model. The reviewed sources do not clearly publish the model’s context window, maximum output length, knowledge cutoff, detailed modality schema, exact API identifier, or pricing. They also do not establish every conventional feature associated with text APIs, such as JSON mode, fine-tuning, caching, or batch processing.

Before using StepAudio 3 Realtime in production, confirm the following with StepFun:

  1. The current model identifier and realtime API protocol
  2. Supported audio formats, sampling rates, codecs, and transport methods
  3. Whether text messages, tool results, and audio events can be mixed in one session
  4. How interruption, cancellation, and turn detection are represented
  5. Latency expectations, concurrency limits, session duration, and rate limits
  6. Current input, output, or session-based pricing
  7. Data retention, regional availability, and handling of sensitive voice recordings

StepAudio 3 Realtime is best understood as a focused voice-agent model with promising realtime interaction characteristics and provider-reported evaluation results, not as a fully documented general-purpose model. Its suitability depends on whether natural, interruptible speech is more valuable for the project than the predictable limits, pricing, and broad text capabilities that may be available from other model types.


Answers to Frequently Asked Questions

Is pricing and API documentation available for StepAudio 3 Realtime?
The reviewed public sources do not clearly document model-specific pricing, the exact API identifier, context limits, maximum output length, or a complete tool schema. Developers should confirm the current realtime protocol, supported audio formats, interruption events, rate limits, pricing, data policies, and regional availability directly with StepFun.
What performance results has StepFun reported for StepAudio 3 Realtime?
StepFun reports scores of 98.9 on the AA Full-Duplex Bench, 56.0 on Ï„-Voice, 90.6 on MMSU, 86.5 on MMAR, and 83.0 on GPQA Diamond. These are vendor-reported results and should be independently validated under the audio, latency, interruption, and tool-use conditions of a production application.
What are the main use cases for StepAudio 3 Realtime?
The model is designed for interruption-aware voice assistants, customer-service agents, hands-free interfaces, emotion-aware spoken applications, interactive tool-using agents, and other low-latency voice interfaces embedded in software or hardware.
What is StepAudio 3 Realtime?
StepAudio 3 Realtime is StepFun’s realtime voice-interaction model for low-latency spoken conversations. It supports audio input and output, streaming, full-duplex interaction, interruption handling, and stated tool-use capabilities.
What does full-duplex interaction mean in StepAudio 3 Realtime?
Full-duplex interaction allows the user and model to speak in overlapping turns. Users can interrupt the model, correct information, or change the subject without waiting for the model to finish its entire response, creating a more natural conversational experience.


Sources 4
Provider

About StepFun