StepAudio 3

StepAudio 3 TTS

by StepFun · Current and publicly accessible

StepAudio 3 TTS is StepFun's dedicated text-to-speech model. It converts text into spoken audio and supports controllable voice delivery, language selection, speech rate, volume, pronunciation, and natural-language delivery instructions. The model is available through StepFun's voice experience and developer platform, while pricing, context limits, maximum output specifications, and the complete voice catalog remain undocumented in the reviewed sources.

Speech Reasoning Coding
StepAudio 3 TTS is a speech-generation model in StepFun's StepAudio 3 family. It is designed for applications that need spoken output rather than text responses, including narration, voice interfaces, localization, accessibility, and expressive audio generation. The model accepts text and returns audio, with controls intended to make the delivery more natural and adjustable.
Outputs

What StepAudio 3 TTS can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
Specifications

Technical details

Model family StepAudio 3
Model type Other
Release date 2026-09-09
Status Current and publicly accessible
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for this speech-generation model in the sources reviewed. Knowledge-cutoff terminology is generally less applicable to a dedicated TTS model than to a text-generation model.

Model notes

Canonical model identifier is stepaudio-3-tts. StepFun lists StepAudio 3 TTS as a current voice model and exposes it in the official speech studio. The model accepts text and generates spoken audio with voice and delivery controls. Public first-party materials reviewed do not specify a context window, maximum output-token limit, knowledge cutoff, fine-tuning process, batch API, caching, or model-specific token pricing. The release date is reported as September 9, 2026 by a current third-party model catalog; StepFun's public model pages confirm current availability but do not display an explicit release date. Editorial scores are comparative estimates for a specialist TTS model, not vendor benchmarks.

Model guide

StepAudio 3 TTS: Controllable Speech Generation for Voice Applications

StepAudio 3 TTS is StepFun's dedicated text-to-speech model for turning written text into spoken audio with controls for voice, language, pace, volume, pronunciation, and delivery style.

What is StepAudio 3 TTS?

StepAudio 3 TTS is StepFun's specialized text-to-speech model. Text-to-speech, often shortened to TTS, converts written words into spoken audio. In practical terms, an application can provide a script, message, or response and use the model to produce a voice recording instead of displaying or returning more text.

The model belongs to StepFun's StepAudio 3 audio family. StepFun's family materials distinguish text-to-speech from other audio functions such as speech recognition, real-time voice interaction, audio generation, and music generation. That positioning is important: StepAudio 3 TTS is intended to be selected when the required task is speech synthesis, not when an application needs transcription or a general conversational language model.

The canonical model identifier is stepaudio-3-tts. StepFun lists it in its online voice experience and developer platform, making it relevant both to people testing speech generation interactively and to developers integrating synthesized speech into software.

How the model works

The basic workflow is straightforward. An application supplies text, chooses an available voice or speech configuration, and receives spoken audio. The documented controls are aimed at the aspects of speech that matter in production: voice selection, language, speaking speed, volume, pronunciation, and delivery instructions.

Delivery instructions are useful when the same words need to sound different in different contexts. For example, a narration workflow may require a measured pace, while a voice interface may need a quicker and more conversational delivery. The available controls can also help with localization, where pronunciation and language selection are more important than simply producing a literal audio rendering of the input.

The supplied research identifies multilingual support and controllable delivery as central characteristics. However, it does not provide a verified list of every supported language, voice, output format, or control value. Those details should therefore be checked in the current StepFun interface or model-specific documentation before implementation.

Input and output modalities

CapabilityStepAudio 3 TTS
Text inputYes
Audio inputNot documented for this model
Image or video inputNo documented support
Text outputNo; the primary result is audio
Spoken-audio outputYes
Music outputNo; this is a speech model

StepAudio 3 TTS should be treated as a text-in, spoken-audio-out model. It is not documented as an image generator, video model, embedding model, speech-recognition system, or general-purpose text generator. If an application needs transcription, it should use a speech-recognition option rather than expecting this TTS model to interpret incoming audio.

Main strengths

  • Controllable delivery: Voice, language, rate, volume, pronunciation, and natural-language delivery instructions can be adjusted according to the available interface or API parameters.
  • Speech-focused design: The model is dedicated to producing spoken audio, which makes its purpose clearer than using a general language model with an attached voice layer.
  • Multilingual use cases: StepFun positions it for multilingual speech synthesis and localization workflows, although the full language list is not established by the supplied sources.
  • Developer access: The model is exposed through StepFun's platform as well as the company's voice experience, allowing it to be evaluated interactively and used in software workflows.
  • Suitability for expressive applications: Controls over pacing, pronunciation, and delivery can be useful for narration, assistants, accessibility features, and voice-over production.

These are model capabilities and positioning claims supported by StepFun's current product materials. They should not be interpreted as independently verified benchmark results. The supplied research does not include comparative listening tests, quality scores, latency measurements, or a published evaluation against competing TTS systems.

Best use cases

StepAudio 3 TTS is a reasonable fit when an application already has text and needs a spoken version of that text. Common examples include:

  • Product narration and voice-over: Generate spoken versions of instructional content, short videos, presentations, or product explanations.
  • Conversational interfaces: Give an assistant or customer-service application an audio response after another system has generated the response text.
  • Accessibility: Read interface messages, notifications, educational content, or other written material aloud.
  • Localization: Produce speech in different languages or adjust pronunciation for regional audiences, subject to the model's currently supported language and voice inventory.
  • Interactive applications: Add generated speech to games, agents, demonstrations, and other software that benefits from audible feedback.
  • Prosody experimentation: Test different rates, volumes, voices, and delivery instructions before settling on a production voice workflow.

For reliable production results, developers should test names, abbreviations, numbers, punctuation, specialist terms, and mixed-language text. These inputs often require pronunciation adjustments even when ordinary sentences sound natural.

Availability and implementation

StepAudio 3 TTS is currently listed by StepFun as an available voice model. It can be found in StepFun's Voice Experience Center, where the exact identifier is shown as stepaudio-3-tts. StepFun also provides a developer platform with speech-related services.

The available research does not establish one complete, current code example or a fixed set of request parameters, so an implementation should follow the current StepFun documentation rather than relying on an inferred API format. In particular, developers should verify the speech endpoint, authentication method, input schema, output format, streaming behavior, and regional availability before building an integration.

Streaming is marked as supported in the supplied model data, which may be useful for applications that want audio to begin arriving before an entire response has finished. The public materials reviewed do not specify latency targets, maximum audio duration, maximum input length, or maximum output size. Those limits may depend on the endpoint or account configuration and should be confirmed directly with StepFun.

Pricing and technical limits

No model-specific token or audio pricing was identified in the supplied authoritative sources. The available data therefore does not support quoting a price for StepAudio 3 TTS. Developers should check the current StepFun platform pricing and determine whether billing is based on characters, duration, requests, or another unit before estimating operating costs.

The following technical details are also not publicly established in the reviewed sources:

  • Context window or maximum input-character limit
  • Maximum output duration or audio size
  • Model-specific knowledge cutoff
  • Fine-tuning availability
  • Batch API availability
  • Caching support
  • A complete published list of voices, languages, and output formats

A knowledge cutoff is less meaningful for a dedicated TTS model than it is for a text-generation model. StepAudio 3 TTS does not need to know current events to pronounce supplied text, but the lack of a published cutoff still means that users should not infer any general language-model knowledge capability from the product name.

Reasoning, coding, and tool support

StepAudio 3 TTS is not a reasoning or coding model. It can synthesize speech from supplied text, but the supplied research does not establish that it independently plans tasks, writes reliable code, calls external tools, browses the web, or returns structured tool actions. The model data marks tool use and JSON-mode output as unsupported.

A typical application may combine it with other components: a language model can generate or summarize text, application code can validate that text, and StepAudio 3 TTS can then speak the approved result. In that architecture, the reasoning belongs to the upstream system and speech synthesis belongs to StepAudio 3 TTS. Keeping those roles separate helps avoid treating a voice model as a general-purpose assistant.

Speed, cost, and capability trade-offs

The supplied editorial data rates StepAudio 3 TTS highly for speed relative to the specialist model set, but that is an editorial estimate rather than a StepFun benchmark. No verified latency or cost comparison is available. Because the model is purpose-built for speech, it may be a more direct choice than routing a text response through a general model and a separate voice system, but the actual cost and responsiveness depend on StepFun's pricing, endpoint, audio format, and account conditions.

The central trade-off is specialization. StepAudio 3 TTS offers speech controls and audio output, but it does not replace a general reasoning model, a speech-recognition model, or a music-generation system. A general multimodal model may be more suitable when the application must interpret images, listen to audio, reason over documents, or use tools before producing a response. A speech-recognition model is more appropriate when the required direction is audio to text. A different audio model is needed for music or non-speech sound generation.

When to choose StepAudio 3 TTS

Choose StepAudio 3 TTS when the primary requirement is controllable spoken output from text, especially when voice, language, pacing, pronunciation, or delivery style matters. It is particularly relevant if the rest of the application already produces text and the missing component is natural-sounding audio.

Consider another option when the task requires any of the following:

  • Understanding or transcribing recorded speech
  • General reasoning, research, or coding
  • Image or video analysis
  • Music generation
  • Embeddings or semantic search
  • Documented fine-tuning, batch processing, or strict output limits that StepFun has not publicly specified for this model

For a final selection, test representative text rather than only short demonstrations. Include the languages, names, numbers, punctuation, emotional styles, and audio durations expected in production. Also confirm current pricing and regional access, since StepFun's services and model availability can change.

Bottom line

StepAudio 3 TTS is a focused StepFun model for turning text into configurable spoken audio. Its clearest value is the combination of speech synthesis, multilingual positioning, voice selection, and delivery controls. It is best evaluated as a TTS component, not as a general AI assistant. The model is currently available through StepFun's voice experience and platform, but important commercial and engineering details—including pricing, input limits, maximum output specifications, and the full voice and language catalog—remain unverified in the supplied public documentation.


Answers to Frequently Asked Questions

How can developers access StepAudio 3 TTS?
The model is available through StepFun's Voice Experience Center and developer platform under the identifier stepaudio-3-tts. Developers should consult the current StepFun documentation to verify the endpoint, authentication, request schema, output format, streaming behavior, pricing, limits, and regional availability.
Does StepAudio 3 TTS support speech recognition, reasoning, or music generation?
No. StepAudio 3 TTS is a text-in, spoken-audio-out model. It is not documented as a speech-recognition, general reasoning, coding, image or video analysis, music-generation, or tool-use model.
What is StepAudio 3 TTS used for?
StepAudio 3 TTS is a StepFun text-to-speech model that converts written text into spoken audio. It is designed for narration, conversational interfaces, accessibility features, localization, voice-over production, and other applications that need configurable speech output.
What controls does StepAudio 3 TTS provide?
StepAudio 3 TTS is positioned as supporting controls for voice selection, language, speaking rate, volume, pronunciation, and natural-language delivery instructions. The exact voices, languages, parameter values, and output formats should be confirmed in the current StepFun documentation.


Sources 3
Provider

About StepFun