StepAudio 3

StepAudio 3 Music

by StepFun · Current; publicly showcased and available through StepFun's StepAudio 3 music experience

StepAudio 3 Music is StepFun’s dedicated music-generation model for creating songs, instrumentals, accompaniment, and cover-style arrangements. It accepts prompts, lyrics, and optional reference audio, uses an ABC-notation planning stage, and reportedly supports tracks up to approximately five minutes and thirty seconds. Public model-specific pricing and several API limits remain unverified.

Music Reasoning Coding
StepAudio 3 Music is a specialized audio-generation model in StepFun’s StepAudio 3 family. Unlike a general-purpose language model, it is designed to create music: users can provide a description, lyrics, an optional song title, and sometimes reference audio to generate a complete song, instrumental arrangement, accompaniment track, or cover-style result. The available research supports its use through StepFun’s music-generation experience, while model-specific API pricing and several production limits remain unverified.
Outputs

What StepAudio 3 Music can produce

Music
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
3/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family StepAudio 3
Model type Other
Release date 2026-09-11
Status Current; publicly showcased and available through StepFun's StepAudio 3 music experience
Knowledge cutoff notes

No model-specific knowledge cutoff was publicly documented. As a generative music model, a text knowledge cutoff is also not a central published specification.

Model notes

StepAudio 3 Music is a specialized music-generation model in the StepAudio 3 family rather than a text language model. StepFun's official demonstration supports prompt and lyric-based song creation, instrumental mode, reference-song input, accompaniment-oriented workflows, and an ABC-notation planning stage. The published technical report describes 48 kHz rendering and generation of songs or instrumentals up to approximately 5 minutes 30 seconds. A model-specific public API price, formal knowledge cutoff, context window, maximum output token limit, batch interface, caching support, and fine-tuning procedure were not verified in current first-party documentation. The September 11, 2026 date corresponds to the published StepAudio 3 Music technical report; a separate official launch date was not clearly stated.

Model guide

StepAudio 3 Music: StepFun’s Model for Songs, Covers, and Instrumentals

StepAudio 3 Music is StepFun’s dedicated music-generation model for creating songs, instrumentals, accompaniment, and cover-style arrangements from text prompts, lyrics, and optional reference audio. It combines musical planning with audio generation and can produce tracks of up to approximately five minutes and thirty seconds.

What is StepAudio 3 Music?

StepAudio 3 Music is StepFun’s specialized model for generating music and songs. Its primary output is rendered audio rather than text, images, video, or structured data. The model is intended for workflows where a user wants to turn a musical idea, lyrics, or an existing recording into a fuller arrangement.

Typical requests can include creating a song in a described style, generating an instrumental background, adding accompaniment to dry vocals, or producing a cover-style arrangement based on reference audio. This places StepAudio 3 Music closer to a dedicated text-to-music and audio-to-music system than to a conversational assistant.

StepFun presents the model as part of its StepAudio 3 family. The official demonstration and technical report identify music generation as a distinct capability, while StepFun’s broader catalog also includes models for speech, voice interaction, text, reasoning, image generation, and video generation. Those other product areas are relevant for positioning, but they are not the purpose of StepAudio 3 Music.

How the model creates music

The model accepts a combination of musical instructions and optional source material. The official demonstration supports descriptive prompts, lyrics, a song title, instrumental mode, and reference audio. Supported reference uploads are reported as WAV, FLAC, MP3, or Opus files.

A notable part of the workflow is an intermediate musical-planning stage based on ABC notation. ABC notation is a compact text representation of musical information such as melody and rhythm. In this system, planning before audio rendering is intended to improve control over musical structure, harmony, rhythm, melody, and arrangement. For users, the practical implication is that the model is not limited to producing an unstructured audio texture from a short prompt; it attempts to organize a more complete musical piece before rendering it.

The technical report describes a music tokenizer operating at 50 Hz with a 65,536-entry codebook, a flow-matching diffusion Transformer for latent prediction, and a 48 kHz audio decoder. In plain terms, the system converts audio into a machine-readable representation, predicts the representation used to form the music, and then decodes it into high-resolution audio. These implementation details are reported technical specifications, not guarantees that every generated track will have professional studio quality.

Inputs, outputs, and supported workflows

AreaWhat is supported or verified
Text inputMusical descriptions, prompts, lyrics, and a song title through the official demonstration
Audio inputOptional reference audio; the demonstration supports WAV, FLAC, MP3, and Opus uploads
Audio outputGenerated songs, instrumental music, accompaniment, and cover-style arrangements
Maximum reported durationApproximately five minutes and thirty seconds for songs or instrumentals
Audio decoding48 kHz output is described in the technical report
Image and video outputNot part of this model’s documented function
Text generationNot the model’s primary output and not documented as a general language capability

The model’s reference-audio workflows are especially important because they broaden its use beyond simple prompt-to-song generation. A user can supply source material for accompaniment or cover-style transformation rather than starting from a blank prompt. The available research does not establish every restriction on reference recordings, such as detailed duration, file-size, rights-management, or content-policy limits, so those details should be checked in the current StepFun interface.

What can you use it for?

  • Songwriting demos: Turn lyrics and a musical description into an early song arrangement.
  • Instrumental production: Generate background or standalone instrumental music from a prompt.
  • Vocal accompaniment: Build accompaniment around supplied dry vocals or other source material.
  • Cover and arrangement experiments: Explore a different arrangement or interpretation using reference audio.
  • Media previsualization: Create temporary music concepts for video, advertising, games, podcasts, or other creative projects.
  • Rapid iteration: Test ideas involving genre, mood, instrumentation, lyrics, and structure before committing to manual production.

These use cases make the model most valuable during ideation and prototyping. A generated track may provide a useful musical direction even when it still requires editing, mixing, rights review, or replacement of individual elements before commercial release.

Strengths and practical trade-offs

StepAudio 3 Music’s clearest strength is specialization. It is designed around songs and arrangements rather than being a general model that happens to produce audio. The combination of lyrics, prompt control, reference audio, instrumental mode, and musical planning gives users several ways to describe the desired result.

The reported generation length is also useful for complete song concepts. Tracks of up to approximately five minutes and thirty seconds are substantially more practical for demos and media drafts than systems limited to very short clips. The 48 kHz decoder is another technically relevant specification for workflows that need conventional high-resolution audio output, although sample rate alone does not determine artistic or production quality.

The main trade-off is that specialization comes with a narrow scope. StepAudio 3 Music is not documented as a general-purpose conversational model, coding assistant, reasoning system, web-search agent, or structured-output engine. It does not provide the broad tool-use surface that a text or multimodal assistant might offer. Its reported speed and cost scores in the supplied catalog are editorial database assessments rather than provider-published benchmarks, so they should be treated as comparative guidance rather than measured guarantees.

Reasoning, coding, and tool support

Reasoning and coding are not meaningful primary capabilities for this model. It may follow musical instructions and organize a composition through its planning stage, but that should not be confused with general-purpose reasoning. The supplied model record assigns a low reasoning score and a low coding score as editorial evaluations, not as official StepFun performance claims.

No tool use, web search, function calling, streaming interface, JSON mode, or structured-output capability was verified for StepAudio 3 Music. The model record lists tool use and streaming as unsupported or unavailable in the documented configuration. Users who need lyrics formatted as JSON, automated research, code generation, or an agent that calls external services should choose a general-purpose model instead and use a dedicated music system for the audio stage.

Pricing and availability

A model-specific public API price was not verified in the supplied first-party documentation. The available sources show StepAudio 3 Music in StepFun’s current platform catalog and showcase it through an official StepAudio 3 music experience, but they do not establish a confirmed per-generation, per-minute, token-based, or subscription price for this model.

It is therefore not accurate to present a numeric price or claim that a particular API plan includes the model. Access may depend on the current StepFun interface, platform eligibility, account requirements, region, or product changes. Users evaluating the model for production should confirm whether the current experience permits commercial use, whether an API is exposed, what quotas apply, and how generated audio is billed.

StepFun is a China-based provider, and its consumer ecosystem is oriented primarily toward mainland China. Registration and access conditions can differ across StepFun’s web, mobile, demonstration, and developer products. The model’s availability should be verified directly before building a workflow around it.

Important limitations and unknowns

The most concrete reported output limit is approximately five minutes and thirty seconds for songs or instrumentals. A conventional language-model context window, maximum output-token count, or formal text knowledge cutoff was not documented because these specifications are not central to this music model.

Several operational details remain unverified in the current sources, including public API pricing, exact quotas, generation latency, maximum reference-audio size, maximum prompt or lyric length, batch processing, caching, fine-tuning, and production-grade service-level guarantees. The technical report and demonstration describe the system’s intended functions, but they do not by themselves establish that every feature is available under every StepFun account or interface.

As with other generative music systems, users should also review the applicable terms and rights requirements before publishing generated tracks. The supplied research does not provide a definitive license policy for StepAudio 3 Music, so ownership, commercial usage, treatment of reference recordings, and distribution rights should not be assumed.

When to choose StepAudio 3 Music

Choose StepAudio 3 Music when the central task is generating a song, instrumental, accompaniment track, or cover-style arrangement and you want control through lyrics, musical descriptions, or reference audio. It is particularly suitable for rapid demos, creative exploration, and previsualization where a dedicated music workflow is more useful than a general chatbot.

Choose another type of model when the main requirement is conversation, coding, web research, image or video generation, structured JSON, external tool use, or low-latency interactive voice. A general audio model may also be more appropriate if the priority is speech recognition, text-to-speech, or real-time dialogue rather than composed music.

For production decisions, the practical comparison is not simply whether StepAudio 3 Music can generate audio. It is whether its specialized musical controls and reported long-form output justify the uncertainty around pricing, API access, quotas, and workflow limits. If those factors are acceptable, it offers a focused route from musical concept or reference recording to a complete draft.


Answers to Frequently Asked Questions

Is StepAudio 3 Music available through a public API, and how much does it cost?
A model-specific public API price was not verified in the available first-party documentation. Access, quotas, commercial-use terms, regional availability, and API eligibility should be confirmed directly through the current StepFun platform or interface.
How long can StepAudio 3 Music generate songs or instrumentals?
The reported maximum duration is approximately five minutes and thirty seconds for songs or instrumental tracks. The technical report also describes a 48 kHz audio decoder.
What can StepAudio 3 Music be used for?
It can be used for songwriting demos, instrumental production, vocal accompaniment, cover and arrangement experiments, media previsualization, and rapid testing of musical ideas involving genre, mood, instrumentation, lyrics, and structure.
What is StepAudio 3 Music?
StepAudio 3 Music is StepFun’s specialized model for generating songs, instrumental music, accompaniment, and cover-style arrangements. It produces rendered audio rather than text, images, video, or structured data.
What inputs does StepAudio 3 Music support?
The model supports musical descriptions, lyrics, song titles, instrumental-mode requests, and optional reference audio. The official demonstration supports WAV, FLAC, MP3, and Opus uploads.


Sources 5
Provider

About StepFun