MiniMax Music

MiniMax Music 3.0

by MiniMax · Open-weight model; MiniMax Token Plan access discontinued on 2026-08-20

MiniMax Music 3.0 is an open-weight text-to-music model that generates complete songs up to five minutes from lyrics and detailed musical descriptions. It supports section-aware arrangements, expressive vocals, and 32 kHz stereo WAV output. The official local workflow requires two CUDA GPUs, uses non-streaming inference, and limits prompts to 5,000 tokens and generation to 9,000 acoustic frames. No current model-specific hosted price was verified, and MiniMax Token Plan access ended on August 20, 2026.

Music Reasoning Coding
MiniMax Music 3.0 is a specialized AI music-generation model from MiniMax, released on August 13, 2026. Instead of generating a short loop or isolated sound effect, it is designed to compose and render a complete song from lyrics plus instructions about genre, tempo, key, mood, vocals, instrumentation, arrangement, and production style. The model is available as open weights for local deployment. MiniMax has also stated that Music model access was discontinued through its Token Plan from August 20, 2026, so hosted availability and commercial terms should be checked separately.
Outputs

What MiniMax Music 3.0 can produce

Music
Inputs

What it can understand

Text
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
4/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family MiniMax Music
Model type Other
Context window 5K tokens
Release date 2026-08-13
Status Open-weight model; MiniMax Token Plan access discontinued on 2026-08-20
Deprecation date 2026-08-20
Knowledge cutoff notes

No authoritative knowledge-cutoff date was published for this music-generation model. Its documented prompt limit is an inference constraint and should not be interpreted as a knowledge cutoff.

Model notes

MiniMax Music 3.0 is a specialized text-to-music model, not a conversational language model. It generates complete songs up to five minutes and produces 32 kHz, 16-bit stereo WAV audio. The official implementation uses an approximately 8B Global LLM, a 0.6B Local LLM, a 2.4B flow-matching component, and a 123M Flow-VAE decoder. The documented prompt limit is 5,000 tokens and the audio-generation limit is 9,000 acoustic frames. Local inference currently requires two CUDA GPUs and supports non-streaming generation only. MiniMax's Token Plan states that music models were discontinued from that plan starting August 20, 2026; the official open-weight checkpoint remains available. Pricing is not populated because no current official model-specific hosted price was verified from the authoritative sources used.

Model guide

MiniMax Music 3.0: Open-Weight AI for Complete Song Generation

MiniMax Music 3.0 is an open-weight text-to-music model for generating complete songs of up to five minutes from lyrics and detailed musical descriptions. It produces expressive vocals, evolving arrangements, and 32 kHz stereo WAV audio, but local use requires a substantial two-GPU CUDA setup and the reference implementation supports non-streaming generation only.

What is MiniMax Music 3.0?

MiniMax Music 3.0 is a text-to-music model that turns written lyrics and musical direction into a finished audio track. Its intended output is a complete song with vocals, instrumental accompaniment, section-level development, and a rendered stereo waveform, rather than a text response, symbolic MIDI file, or short sound effect.

The model is provided by MiniMax and sits within the company's broader audio and music portfolio. It is distinct from general-purpose language or coding models because it does not generate conversational text as its primary output. The official release presents it as an open-weight model, with a checkpoint and local inference instructions available through the MiniMax-Music3 GitHub and Hugging Face repositories.

A typical request might include lyrics marked with sections such as [Verse], [Pre-Chorus], and [Chorus], together with a description such as an electronic pop track at a specified tempo, featuring layered synthesizers, a warm lead vocal, backing harmonies, and a gradual build toward the final chorus.

How the generation workflow works

Music 3.0 accepts two main types of input. Lyrics provide the words and can include explicit structural tags. Musical instructions provide the creative and production guidance. These instructions can describe genre, subgenre, BPM, key, scale, emotional progression, intended listening context, vocal characteristics, instruments, groove, bass, percussion, effects, and the way the arrangement should change between sections.

The structured-caption workflow divides this information into three broad areas:

  • Global metadata: Genre, tempo, key, scale, mood progression, listening scenario, and production profile.
  • Vocal details: Vocal gender, timbre, delivery style, harmonies, backing vocals, and vocal effects.
  • Arrangement: Primary and secondary instruments, section-specific changes, rhythmic feel, bass, percussion, textures, and spatial effects.

MiniMax also describes a prompt-enhancement skill that can expand a short musical idea into a more detailed structured caption. This can help users who know the desired mood or genre but do not yet know how to specify the arrangement in production-oriented terms.

Song structure and output

The model is intended to maintain musical identity across relatively long sequences. It can represent common song sections such as intros, verses, pre-choruses, choruses, bridges, instrumental passages, solos, and outros. Explicit tags such as [Intro], [Verse], [Chorus], [Bridge], [Instrumental], [Solo], and [Outro] give the generation process clearer structural guidance.

According to the supplied specifications, Music 3.0 can generate songs up to five minutes long. The reference implementation produces 32 kHz, 16-bit stereo WAV audio. These are concrete output characteristics, but they should not be interpreted as guarantees that every requested lyric, instrument, tempo, key, or section will be reproduced exactly. The model treats such instructions as generative guidance rather than strict symbolic constraints.

Music 3.0 produces direct audio output, so it is multimodal in the output sense rather than a text-only model. Its documented workflow does not describe image, video, or audio input. Lyrics and musical descriptions are supplied as text; the result is music audio.

Technical architecture in plain language

The model uses separate components for long-range musical planning and detailed sound generation. Its approximately 8-billion-parameter Global LLM predicts a primary semantic music codebook. In practical terms, this component helps organize the higher-level musical content, such as the progression and identity of the song.

An approximately 0.6-billion-parameter Local LLM predicts additional acoustic codebooks within each frame. These codebooks represent more localized details of the sound. The system then combines hidden-state fusion, flow matching, and a Flow-VAE decoder to turn the internal representation into audio.

The implementation uses eight layers of residual vector quantization. One semantic codebook captures core musical structure, while seven acoustic codebooks represent progressively finer audio details. The official repository identifies an approximately 2.4-billion-parameter flow-matching component and a 123-million-parameter Flow-VAE decoder.

These architecture details explain why the model is more demanding than a lightweight audio-generation service. The system is not simply predicting a small sequence of labels and exporting it; it uses multiple stages to plan musical content, model acoustic detail, and reconstruct the final waveform.

Context limits and local deployment

The official implementation limits the tokenized text prompt to 5,000 tokens. This is the documented input constraint for the lyrics and musical description workflow. Audio generation is limited to 9,000 acoustic frames in the reference implementation. The maximum song duration is documented as up to five minutes, although actual results can depend on the supplied content and generation process.

Local serving is documented through SGLang-Omni and requires two CUDA GPUs. One GPU is used for language-model and music-token generation, while the other handles flow matching and waveform decoding. This makes the open-weight release more suitable for users with capable local hardware, a managed GPU environment, or an interest in self-hosting than for casual users seeking an instant browser-based generation experience.

The documented local inference endpoint is audio-speech compatible. Lyrics are passed in the input field, the musical description is passed in instructions, and the requested output can be set to WAV. The current reference implementation supports non-streaming generation only, so users should expect the complete result after generation rather than progressive audio delivery.

Pricing and availability

No current official model-specific hosted price was verified in the supplied sources, so a reliable per-song or per-minute price cannot be stated. The model is available as an open-weight checkpoint through the official MiniMax-Music3 repositories, but open weights do not mean that local use is cost-free: users still need suitable GPU hardware or rented compute, storage, and the software environment required by the implementation.

MiniMax's Token Plan availability page states that music models became unavailable through that plan on August 20, 2026. The open-weight checkpoint remains available, while access through MiniMax Audio or other hosted services, regional availability, and commercial usage rights should be verified directly with MiniMax. The discontinuation of Token Plan access should not be confused with the shutdown of the open-weight model itself.

Main strengths and limitations

The model's main strength is its focus on complete songs. It combines lyrics, vocal direction, arrangement guidance, and long-form musical development in one generation workflow. Explicit section tags and structured captions provide more control than a short, vague prompt, while the open-weight release gives technically capable users the option to run the checkpoint locally.

  • Complete-song orientation: It is designed for tracks up to five minutes, including vocals and evolving arrangements.
  • Detailed musical prompting: Users can specify genre, key, tempo, vocal qualities, instruments, section changes, and production character.
  • Open-weight deployment: The official checkpoint supports local or self-hosted experimentation.
  • Direct stereo audio output: The reference implementation produces 32 kHz, 16-bit stereo WAV files.
  • Long-range structure: The model is intended to preserve musical identity across verses, choruses, bridges, instrumental sections, and outros.

There are also important limitations. Music 3.0 is not a symbolic music sequencer, so it does not provide guaranteed note-by-note control or strict adherence to every requested parameter. A prompt asking for an exact BPM, key, instrument entrance, lyric pronunciation, or section length may produce an approximation rather than a formally constrained result.

Local deployment is resource-intensive because the documented setup requires two CUDA GPUs. The reference implementation is non-streaming, and there is no supplied evidence that every feature of a hosted MiniMax Audio product is available in the local open-weight workflow. The model also has no documented tool or function-calling capability, JSON output mode, web search, or general-purpose reasoning interface. Its reasoning and coding capabilities are therefore not meaningful use cases: it is an audio-generation model, not a general assistant.

When to choose MiniMax Music 3.0

Choose MiniMax Music 3.0 when the goal is to generate a complete song from lyrics and a detailed creative brief, especially when local deployment, open weights, or experimentation with structured musical prompts matters. It is a strong fit for songwriting demos, soundtrack concepts, arrangement exploration, vocal-production references, and self-hosted research into music-generation systems.

It may be a better choice than a short-loop generator when the desired result needs a recognizable verse-and-chorus structure or several minutes of development. It may also be preferable to a closed hosted service for teams that specifically need access to an official checkpoint and control over their inference environment.

Another option may be more appropriate when the priority is instant browser-based generation, low hardware requirements, streaming playback, exact symbolic composition, or reliable compliance with strict tempo, key, instrumentation, and lyric constraints. A hosted service may also be preferable when GPU costs and deployment work outweigh the value of local control. Because MiniMax's Token Plan no longer includes Music models according to the supplied availability notice, users seeking managed access should verify the current MiniMax Audio offering rather than assume that the Token Plan supports this model.

Bottom line

MiniMax Music 3.0 is best understood as an open-weight complete-song generator rather than a general AI assistant or conventional digital audio workstation. It combines lyric conditioning and detailed musical descriptions with a multi-stage architecture for planning, acoustic modeling, and waveform synthesis. Its five-minute output limit, vocals, section-aware prompting, and local checkpoint make it useful for serious creative prototyping. The trade-off is substantial: the documented deployment requires two CUDA GPUs, generation is non-streaming, pricing for hosted use is unverified, and musical instructions remain guidance rather than guarantees.


Answers to Frequently Asked Questions

Is MiniMax Music 3.0 available through a hosted service or priced per song?
No current official model-specific hosted price was verified in the supplied sources. The open-weight checkpoint is available through MiniMax repositories, but local use still requires suitable hardware or rented GPU compute. MiniMax's Token Plan availability notice says that music models became unavailable through that plan on August 20, 2026, so current hosted access and commercial rights should be confirmed directly with MiniMax.
What hardware is required to run MiniMax Music 3.0 locally?
The documented local deployment uses SGLang-Omni and requires two CUDA GPUs. One GPU handles language-model and music-token generation, while the other performs flow matching and waveform decoding. Users also need the model files, storage, and the required software environment.
What audio quality and song length does MiniMax Music 3.0 support?
The reference implementation can generate songs up to five minutes long and produces 32 kHz, 16-bit stereo WAV audio. These specifications describe the target output, but the model may not reproduce every requested lyric, tempo, key, instrument, or section exactly.
What is MiniMax Music 3.0?
MiniMax Music 3.0 is an open-weight text-to-music model that generates complete songs from lyrics and musical instructions. It can produce vocals, instrumental accompaniment, evolving song sections, and a rendered stereo WAV file.
How do you prompt MiniMax Music 3.0 to generate a song?
Provide lyrics with section tags such as [Verse], [Chorus], and [Bridge], along with musical instructions describing the genre, BPM, key, mood, vocal style, instruments, rhythm, effects, and arrangement changes. The model also includes a prompt-enhancement workflow that can expand a short idea into a more detailed structured caption.


Sources 4
Provider

About MiniMax