MiniMax Music

MiniMax Music 2.0

by MiniMax · Legacy or limited availability; MiniMax states that music models were no longer available through Token Plan from 2026-08-20, but a full model retirement date was not verified.

MiniMax Music 2.0 is a text-driven audio model for creating complete songs with vocals, melodies, instrumental arrangements, expressive styles, duets, a cappella passages, and structured sections such as verses and choruses. It launched in October 2025 and supports compositions of up to five minutes according to MiniMax. Exact pricing and current model access are unverified, while MiniMax states that music models left its Token Plan on August 20, 2026.

Music Reasoning Coding
MiniMax Music 2.0 turns lyrics and musical instructions into complete song-length audio. It is designed for users who want generated vocals, melodies, accompaniment, and recognizable song structures rather than isolated loops or short instrumental fragments. MiniMax describes support for multiple genres, expressive vocal delivery, male and female voices, duets, a cappella music, spoken or recited material, and compositions lasting up to five minutes. Current access is less certain: MiniMax has promoted newer music models, and its Token Plan documentation says music models were removed from that plan on August 20, 2026.
Outputs

What MiniMax Music 2.0 can produce

Music
Inputs

What it can understand

Text
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
6/10 Speed
Specifications

Technical details

Model family MiniMax Music
Model type Other
Release date 2025-10-31
Status Legacy or limited availability; MiniMax states that music models were no longer available through Token Plan from 2026-08-20, but a full model retirement date was not verified.
Knowledge cutoff notes

MiniMax has not published a verified knowledge-cutoff date for this generative music model.

Model notes

MiniMax announced Music 2.0 on October 31, 2025. The model generates complete structured songs with verses, choruses, bridges, vocals, and instrumental accompaniment, with compositions stated to reach up to five minutes. MiniMax describes support for multiple vocal styles, emotional expression, male and female duets, a cappella music, and prompt-based instrument control. The exact model's context length, output-token limit, knowledge cutoff, pricing, fine-tuning, caching, batch API, streaming, JSON mode, and structured-output support were not publicly verified. MiniMax's current Token Plan page says all music models were discontinued from Token Plan on August 20, 2026; this is not an independently verified full shutdown date for the model across all MiniMax services.

Model guide

MiniMax Music 2.0: Complete AI Songs with Vocals and Structured Arrangements

MiniMax Music 2.0 is a text-driven generative music model released on October 31, 2025. It creates complete songs from lyrics, musical ideas, and style descriptions, including vocals, instrumental arrangements, expressive singing styles, structured sections, duets, a cappella passages, and compositions of up to five minutes. It is now best treated as a legacy or limited-availability model because MiniMax states that music models became unavailable through its Token Plan on August 20, 2026.

What is MiniMax Music 2.0?

MiniMax Music 2.0 is a generative music model from MiniMax that creates musical audio from text-based creative direction. A prompt can combine lyrics, a genre, a desired mood, vocal characteristics, instruments, and an arrangement concept. The result is intended to be a complete song with singing, melody, accompaniment, and multiple sections rather than a text response or a short sound effect.

MiniMax officially announced Music 2.0 on October 31, 2025. The model belongs to the provider's music-generation and audio ecosystem, alongside newer offerings such as Music 2.5 and Music 3.0. Those newer products matter for current users because Music 2.0 may no longer be the most accessible or actively promoted option.

The model's primary output is audio. It does not function as a general-purpose language model, coding model, image generator, or video generator. Its value is concentrated in turning musical concepts into finished or near-finished song drafts.

How the model generates songs

Music 2.0 is designed to combine several parts of music production in one generation workflow. Users can provide lyrics or describe the kind of song they want, while the model handles elements such as vocal performance, melody, rhythm, instrumentation, and arrangement.

For example, a prompt might request a reflective pop song with a female vocal, piano and soft drums, an emotionally restrained verse, and a larger chorus. The supplied research supports this kind of control at the level of genre, instruments, emotional tone, vocal approach, and overall soundscape. It does not establish that every prompt will produce precise, deterministic control over individual notes or a professional multitrack project.

MiniMax says compositions can include recognizable sections such as verses, choruses, and bridges. This section-level structure is important for users creating songs, advertisements, background music, demos, or narrative musical pieces because it provides more than a repeating musical texture.

Vocals, genres, and performance control

One of Music 2.0's main distinctions is its emphasis on sung vocals. MiniMax describes the model as producing human-like vocal timbres and following directions about singing technique, emotion, and performance style. These are provider claims rather than independent benchmark results, but they identify the model's intended use more clearly than a generic text-to-audio label.

The documented use cases include pop, jazz, blues, rock, folk, and electronic music. The model is also described as supporting male and female vocals, duets, and a cappella arrangements. A duet prompt can therefore specify an interaction between two vocal parts, while an a cappella request can focus the generation on voices without conventional instrumental accompaniment.

Prompt-based instructions can describe delivery and mood, such as intimate singing, energetic rock performance, melancholy phrasing, or dramatic vocal expression. MiniMax also presents the model as able to preserve a core vocal identity while changing singing styles within a composition. The supplied sources do not provide a formal control specification, voice-identity guarantee, or consistency benchmark, so these features should be evaluated through actual samples rather than treated as guaranteed production behavior.

Arrangement and specialized audio capabilities

Music 2.0 can follow instructions about instrumental accompaniment. MiniMax specifically references instruments such as jazz horns, drums, and piano, as well as layered arrangements more broadly. A user can combine instrumentation with a genre and atmosphere to request, for example, a piano-led ballad, a horn-driven jazz section, or an electronic track with a particular emotional character.

The model is not limited to conventional verse-and-chorus songs. MiniMax also reports support for emotionally directed monologue soundtracks and poetry recitation with music. These applications combine spoken or recited material with musical development and atmospheric accompaniment. They may be useful for short films, audio storytelling, spoken-word projects, creative presentations, or dramatic readings.

The officially stated maximum composition length is up to five minutes. This is a provider-published claim about the model's intended generation capacity, not a guarantee that every prompt will produce a coherent five-minute arrangement. Longer projects may still require multiple generations, editing, or external audio production.

Supported inputs and outputs

The documented input is primarily text: lyrics, musical ideas, style descriptions, arrangement instructions, and performance direction. The model's documented output is music audio containing vocals, instrumental accompaniment, or both.

CapabilityVerified or documented status
Text inputSupported for lyrics, musical concepts, styles, instruments, and mood instructions.
Audio outputSupported; the main output is generated music and song audio.
Vocal outputSupported according to MiniMax's model description.
Instrumental outputSupported, including prompted accompaniment and arrangements.
Maximum composition lengthUp to five minutes, according to MiniMax.
Image, video, or text outputNot documented for this model.
Context length or output-token limitNot publicly verified.

Music 2.0 should therefore be evaluated as an audio-generation system, not through language-model expectations such as chat quality, token throughput, or structured JSON responses. The available research does not verify a public context-window specification, token limit, streaming interface, function calling, tool use, or structured-output mode.

Reasoning, coding, and API limitations

Reasoning and coding scores are not meaningful primary measures for this model. Music 2.0 is built for musical generation, not multi-step text reasoning, software development, data analysis, or general question answering. It should not be selected for coding agents or workflows that need a language model to return executable code.

The supplied sources also do not verify tool or function support, web search, fine-tuning, caching, batch processing, streaming, JSON mode, or structured outputs. These missing specifications do not prove that no interface ever supported such features; they mean that the capabilities were not publicly verified in the supplied research.

Pricing for the exact Music 2.0 model was not verified. MiniMax's current Token Plan page states that all music models became unavailable through Token Plan beginning August 20, 2026. That statement establishes a change in Token Plan availability, but it does not independently establish a complete retirement of Music 2.0 from every MiniMax product, audio platform, or other access route.

Current status and availability

Music 2.0 should currently be treated as a legacy or limited-availability model. MiniMax has since promoted newer Music 2.5 and Music 3.0 offerings, and the Token Plan documentation says that music models were removed from that plan on August 20, 2026. The research does not include a model-specific shutdown notice confirming that Music 2.0 was permanently disabled across all services.

Users who need this exact model should check the current MiniMax Audio or other official MiniMax product surfaces rather than assuming that a historical launch page means the model remains available. Access, quotas, product placement, and pricing may change independently across MiniMax services.

Strengths and limitations

Main strengths

  • Generates complete song concepts rather than only short musical fragments.
  • Combines vocals, melody, instrumentation, and arrangement in one workflow.
  • Supports prompts describing vocal emotion, delivery, genre, and performance style.
  • Can produce structured sections such as verses, choruses, and bridges.
  • Includes documented support for duets, a cappella concepts, and multiple musical genres.
  • Can generate music for poetry recitation and emotionally directed monologues.
  • Has a provider-stated composition length of up to five minutes.

Main limitations

  • Current access and pricing for the exact model are not verified.
  • Token Plan access was discontinued for music models on August 20, 2026.
  • It may be less suitable than newer MiniMax music offerings for users seeking an actively supported model.
  • Public documentation does not specify context length, token limits, streaming, tool use, fine-tuning, or structured outputs.
  • The sources describe broad creative control but do not establish precise note-level editing, multitrack export, or guaranteed vocal consistency.
  • It is not intended for general reasoning, coding, image generation, video generation, or speech-only synthesis.

When to choose MiniMax Music 2.0

Choose Music 2.0 when the goal is to turn lyrics or a detailed musical brief into a complete song with vocals and accompaniment, and when the model is still available through the MiniMax surface you use. It is particularly relevant for fast songwriting drafts, concept albums, demo tracks, background songs, poetic narration, and experiments with duets or a cappella arrangements.

It may be preferable to a basic loop or instrumental generator when a recognizable song form and sung performance are important. It can also be a useful choice when one prompt needs to specify both the vocal character and the instrumental setting.

Another option may be more appropriate when current availability is the deciding factor, when you need confirmed pricing or API guarantees, or when you require detailed control over individual tracks and notes. Newer MiniMax music models may be worth checking for an actively maintained alternative, while a dedicated digital audio workstation or specialized music-production tool may be better for precise editing. A general language model is more appropriate for songwriting assistance, coding, or reasoning before the musical generation step.

Bottom line

MiniMax Music 2.0 is a song-oriented generative audio model whose defining capability is the creation of structured, vocal-led music from text instructions. Its documented strengths include expressive singing, genre and instrument control, duets, a cappella arrangements, verses and choruses, and compositions of up to five minutes. Its main practical concern is not the intended feature set but its uncertain current availability: MiniMax has moved on to newer music offerings and ended music-model access through Token Plan. Treat it as a legacy model unless an official current MiniMax product page confirms access.


Answers to Frequently Asked Questions

What are the main limitations of MiniMax Music 2.0?
Public documentation does not verify precise note-level editing, multitrack export, guaranteed vocal consistency, context length, token limits, streaming, fine-tuning, structured outputs, or API features. The model is designed for music generation rather than coding, reasoning, image generation, video generation, or general question answering.
Is MiniMax Music 2.0 still available?
Music 2.0 should be treated as a legacy or limited-availability model. MiniMax has promoted newer Music 2.5 and Music 3.0 offerings, and music models were removed from the Token Plan on August 20, 2026. Users should check current official MiniMax product pages for access and pricing.
How long can songs generated by MiniMax Music 2.0 be?
MiniMax states that Music 2.0 can generate compositions of up to five minutes. This is a provider-published maximum and does not guarantee that every five-minute generation will have a coherent or polished arrangement.
What is MiniMax Music 2.0?
MiniMax Music 2.0 is a generative music model that turns text prompts into complete or near-complete songs with vocals, melody, instrumental accompaniment, and structured sections such as verses, choruses, and bridges.
What types of music can MiniMax Music 2.0 generate?
It is documented as supporting genres including pop, jazz, blues, rock, folk, and electronic music. Prompts can also specify instruments, mood, vocal style, performance technique, duets, a cappella arrangements, poetry recitation, or emotionally directed monologue soundtracks.


Sources 5
Provider

About MiniMax