StepAudio 3

StepAudio 3 Gen

by StepFun · Current research and platform-listed model; public API and access details are limited

StepAudio 3 Gen is StepFun’s broad audio-generation model for creating expressive speech, designed voices, singing, music, sound effects, ambience, multi-speaker dialogue, and composite acoustic scenes. Its RVQ-based architecture supports a unified approach to audio generation, but public pricing, API details, context limits, and maximum output length remain undocumented.

Speech Music
StepAudio 3 Gen is a StepFun audio model designed to generate more than spoken narration. It can create zero-shot text-to-speech, natural-language voice designs, vocals, music, sound effects, ambience, vibe speech, multi-speaker dialogue, and layered audio scenes. The model is listed in StepFun’s current platform catalog and is described in a provider-authored technical report, but important access and specification details remain unpublished, including public pricing, context limits, maximum output length, and a stable API model identifier.
Outputs

What StepAudio 3 Gen can produce

Speech Music
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Multimodal output
Specifications

Technical details

Model family StepAudio 3
Model type Other
Release date 2026-09-11
Status Current research and platform-listed model; public API and access details are limited
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for this audio-generation model. Its technical report describes model architecture and capabilities but does not specify a training-data cutoff.

Model notes

The model is independently verified through StepFun's current platform catalog, an official StepAudio 3 Gen showcase, and the provider-authored technical report submitted on September 11, 2026. It supports speech, vocals, music, sound effects, vibe speech, multi-speaker dialogue, and composite acoustic scenes. The technical report describes a discrete autoregressive generator over a shared 16 x 2048 RVQ code space operating at 12.5 Hz, with the first codebook generated across time and the remaining codebooks completed across the codebook dimension. The official showcase states that the experience center is forthcoming, while the platform lists StepAudio 3 Gen among its voice models. No authoritative public token pricing, context-window limit, maximum-output-token limit, exact API model identifier, downloadable weights, or lifecycle dates were found.

Model guide

StepAudio 3 Gen: One Model for Speech, Music, Vocals, and Sound Design

StepAudio 3 Gen is StepFun’s general-purpose audio-generation model for producing expressive speech, designed voices, singing, music, sound effects, ambience, multi-speaker dialogue, and composite acoustic scenes from natural-language instructions. Its main distinction is the attempt to handle many audio-generation tasks within one shared discrete autoregressive system rather than focusing only on conventional text-to-speech.

What is StepAudio 3 Gen?

StepAudio 3 Gen is StepFun’s general-purpose audio-generation model. Its purpose is to turn text instructions and content into audio, covering several tasks that are often handled by separate systems: speech synthesis, voice creation, singing, music generation, sound effects, ambience, and multi-speaker or multi-element scenes.

A conventional text-to-speech system usually takes written words and produces a spoken voice. StepAudio 3 Gen is positioned more broadly. A user can specify what should be said, but can also describe how the result should sound, such as a particular emotional delivery, vocal character, timing, musical style, atmosphere, or combination of sound elements. The official showcase presents it as a system for voice design, zero-shot text-to-speech, vocals, music, sound effects, and composite audio generation.

StepFun is the provider. The model appears in StepFun’s current platform catalog as a voice model, while the official StepAudio 3 Gen showcase and technical report provide the more detailed description of its capabilities and architecture. The model’s reported release or research date is September 11, 2026. The available evidence identifies it as a current research and platform-listed model, although the public API and access arrangements are not fully documented.

What can it generate?

StepAudio 3 Gen’s main strength is breadth within audio generation. Supported output categories documented in the supplied research include:

  • Speech: zero-shot text-to-speech and expressive spoken delivery.
  • Voice design: natural-language descriptions of a desired voice, rather than relying only on a fixed speaker list.
  • Vibe speech: speech shaped by a requested mood, style, or performance character.
  • Vocals and singing: vocal generation beyond ordinary spoken narration.
  • Music: generation of musical audio from instructions.
  • Sound effects: individual effects and designed sonic elements.
  • Ambience: environmental or atmospheric audio.
  • Multi-speaker dialogue: scenes involving more than one voice.
  • Composite acoustic scenes: combinations such as dialogue, background ambience, music, and effects in one generated result.

For example, a creator could request a short scene containing two speakers, a specified emotional exchange, distant environmental sound, and a musical bed. The research supports this type of composite generation as a model capability, but it does not establish a guaranteed production format, maximum scene duration, or a fixed set of controls for every public interface.

How the model works in technical terms

The technical report describes StepAudio 3 Gen as a discrete autoregressive audio-generation system. In plain language, it generates a sequence of discrete audio codes rather than directly predicting a waveform sample by sample. These codes are later decoded into audible sound.

The model uses a shared residual vector quantization, or RVQ, representation. RVQ compresses audio into several layers of codebooks. The report describes a 16 × 2048 RVQ code space operating at 12.5 Hz. The first codebook is generated across time, and the remaining codebooks are completed across the codebook dimension. This design gives the model a structured representation for coordinating speech, music, vocals, effects, and other acoustic components.

The shared representation is important to the model’s positioning. Rather than treating speech, music, and sound effects as completely unrelated tasks, StepAudio 3 Gen uses one audio-generation framework for them. That does not mean every category will have identical quality or control. It means the same model family is intended to handle a wider range of acoustic outputs than a speech-only system.

Inputs, outputs, and modalities

CapabilityResearch-supported status
Text inputSupported for natural-language generation instructions and text-to-speech.
Audio inputListed as supported in the model record, although the supplied sources do not define every audio-input workflow.
Audio outputSupported.
Speech outputSupported, including zero-shot text-to-speech and expressive speech.
Music outputSupported.
Image input or outputNot documented as supported.
Video input or outputNot documented as supported.
Text outputNot the model’s output purpose; the model record identifies text output as unsupported.

The available information confirms audio generation, but it does not provide a complete public contract for audio file formats, sampling rates, channel layouts, latency, or output-duration limits. Those details should be verified in the specific StepFun interface or API documentation available to an account.

Where StepAudio 3 Gen is strongest

The clearest advantage is the combination of speech and non-speech audio generation. A speech-focused model may be a better fit for predictable narration, but StepAudio 3 Gen is aimed at projects where the voice is only one part of the result. Its documented capabilities make it relevant to the following workflows:

  • Character and narrative voices: create voices for animation, games, audio fiction, or demonstrations without depending on a narrowly defined speaker inventory.
  • Expressive narration: produce speech with specified emotion, style, timing, or performance direction.
  • Song and vocal concepting: prototype vocals or singing as part of an early creative workflow.
  • Audio prototyping: explore music, sound effects, and ambience before commissioning separate production work.
  • Scene generation: combine dialogue and environmental elements for rough cuts, interactive media, or storyboards.
  • Sound-design experimentation: test natural-language descriptions of a sonic idea rather than manually assembling every layer.

Its breadth may also reduce the need to switch between a text-to-speech engine, a music generator, and a sound-effect tool during early ideation. That is a practical advantage, but it should not be interpreted as proof that one generated result will match the quality, consistency, or controllability of specialized tools in every category.

Limitations and unresolved specifications

The most important limitation is incomplete public documentation. The supplied research does not identify an authoritative token-style price, audio-generation price, context-window limit, maximum output length, exact API model identifier, downloadable weights, or lifecycle dates. It also does not document fine-tuning, caching, batch processing, streaming, tool use, or structured output support. These values should be treated as unknown rather than assumed to be unavailable in every StepFun environment.

The official showcase reportedly describes the experience center as forthcoming, while the platform catalog lists StepAudio 3 Gen among its voice models. This suggests that listing and actual hands-on availability may not be identical. Access can depend on the StepFun platform, account status, region, or the rollout state of a particular service.

StepAudio 3 Gen is not the appropriate choice for text-only language tasks, image or video generation, or speech recognition. The model record specifically positions it as an audio-generation system, not an automatic speech-recognition model. It is also a poor fit when an application requires a stable, fully documented public model contract and the necessary API details are not available to the intended user.

There are also creative-production limitations to consider. The research confirms broad categories of output, but it does not guarantee precise control over pronunciation, speaker identity across long projects, musical structure, timing accuracy, mixing quality, or the repeatability of a generated scene. For professional production, generated material may still require editing, selection, synchronization, mastering, and rights review.

Reasoning, coding, and tool support

StepAudio 3 Gen is not documented as a reasoning or coding model. It may interpret natural-language instructions as part of audio generation, but that should not be confused with a language model designed for multi-step reasoning, software development, or general-purpose conversation.

No authoritative tool or function-calling capability is supplied for this model. Web search, external actions, structured JSON output, and code execution should therefore not be assumed. The model’s output is audio rather than text, and the research does not describe a separate agentic workflow built into StepAudio 3 Gen itself.

Pricing and API access

No verified public price is available in the supplied research. There is also no documented context limit, maximum output limit, or stable public API identifier. StepFun does operate a broader platform with developer services, but the existence of that platform does not establish that every catalog-listed model has identical API availability or pricing.

For evaluation, users should check the current StepFun platform listing for account eligibility, regional availability, billing units, request limits, supported endpoints, and whether the model is accessible through an experience center, an API, or another product interface. Until those details are published or confirmed directly, cost and throughput comparisons with dedicated speech, music, or sound-effect services remain unverified.

When to choose StepAudio 3 Gen

Choose StepAudio 3 Gen when the project needs several kinds of generated audio from one model, especially when speech, vocals, music, ambience, and effects need to coexist in a single concept or scene. It is particularly interesting for rapid creative prototyping, character voices, audio fiction, game concepts, multimedia storyboards, and experiments where natural-language control is more useful than a fixed library of voices or sounds.

Choose a more specialized option when the priority is a documented production API, predictable pricing, long-term voice consistency, highly controlled narration, speech recognition, or a narrowly defined music or sound-effect workflow. A speech-only system may be easier to evaluate for dependable narration, while a dedicated music or sound-design tool may provide more focused controls for its category. StepAudio 3 Gen’s value is its cross-category scope, not a verified claim that it is the fastest, cheapest, or most accurate option for each individual task.

Bottom line

StepAudio 3 Gen is a distinctive StepFun model because it treats audio generation as a broad composition problem rather than only a text-to-speech problem. The documented scope includes expressive speech, voice design, singing, music, sound effects, ambience, multi-speaker dialogue, and composite scenes, supported by a shared RVQ-based discrete autoregressive architecture.

Its main practical uncertainty is not the stated capability range but the surrounding product contract. Public information does not yet establish pricing, context or duration limits, exact API access, or all operational controls. It is therefore best evaluated as a promising, broad audio-generation model for creative experimentation and multimodal sound design, with deployment decisions postponed until StepFun confirms the required access and production specifications.


Answers to Frequently Asked Questions

Is StepAudio 3 Gen available through a public API, and how much does it cost?
The supplied information does not verify a public price, stable API model identifier, context limit, or maximum output duration. Users should check the current StepFun platform documentation for availability, regional access, billing, request limits, and supported endpoints.
Who should use StepAudio 3 Gen?
It is suited to creators developing character voices, expressive narration, songs, music concepts, game audio, audio fiction, sound-design prototypes, and scenes that combine dialogue with ambience, music, or effects.
How does StepAudio 3 Gen work technically?
StepAudio 3 Gen is described as a discrete autoregressive audio-generation system. It uses a shared residual vector quantization representation with a 16 × 2048 RVQ code space operating at 12.5 Hz, generating audio codes that are decoded into sound.
What is StepAudio 3 Gen?
StepAudio 3 Gen is StepFun’s general-purpose audio-generation model for creating speech, designed voices, singing, music, sound effects, ambience, multi-speaker dialogue, and composite acoustic scenes from natural-language instructions.
What types of audio can StepAudio 3 Gen generate?
It can generate expressive speech, zero-shot text-to-speech, custom voice designs, vocals and singing, music, sound effects, ambience, multi-speaker conversations, and combinations of these elements in a single audio scene.


Sources 3
Provider

About StepFun