MiniMax Music

MiniMax Music Cover

by MiniMax · Current; specialized music-cover model introduced with MiniMax Music 2.6

MiniMax Music Cover is a specialist audio-to-audio model for transforming authorized recordings into new genres and arrangements. It uses reference audio and a style prompt, aims to preserve the source melody, and generates a new performance with changed vocals, instrumentation, and production direction.

Music Reasoning Coding
MiniMax Music Cover is designed for transforming existing songs rather than creating music from a blank prompt. You provide a reference recording and a description of the desired style, and the model generates a new musical performance with changed vocals, instruments, arrangement, and genre while retaining the original melodic foundation.
Outputs

What MiniMax Music Cover can produce

Music
Inputs

What it can understand

Text Audio Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
6/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family MiniMax Music
Model type Other
Release date 2026-04-10
Status Current; specialized music-cover model introduced with MiniMax Music 2.6
Knowledge cutoff notes

Knowledge-cutoff information is not applicable or publicly documented for this specialized music-generation model.

Model notes

Canonical model identifier is music-cover. The workflow requires reference audio and a target-style prompt. MiniMax introduced Cover as a new capability in Music 2.6 on April 10, 2026. Official MiniMax agent documentation lists MP3, WAV, and FLAC reference audio, typically six seconds to six minutes and up to 50 MB. Lyrics can be supplied or extracted from the source audio through the surrounding workflow. This is a specialist audio-generation model, so token context, maximum output tokens, knowledge cutoff, and conventional LLM pricing do not apply.

Cost

Model pricing

Input Not applicable to token pricing; audio-generation pricing varies by MiniMax platform or deployment
Output Not applicable to token pricing; charged per generated audio result where applicable
Model guide

MiniMax Music Cover: Transform Existing Songs into New Genres

MiniMax Music Cover is a specialized audio-to-audio model that reinterprets an existing song in a different genre, arrangement, vocal style, or production direction while aiming to preserve the source melody.

What Is MiniMax Music Cover?

MiniMax Music Cover is a specialized audio-to-audio music-generation model provided by MiniMax. Its defining feature is that it starts with an existing song. Instead of asking the system to invent a melody from text, you upload a reference track and describe how you want that track reinterpreted.

The model is intended to preserve the recognizable melodic structure of the source while rebuilding the surrounding performance. Depending on the instruction, the result can change the genre, instrumentation, vocal character, arrangement, mood, tempo direction, and overall production style.

MiniMax introduced Cover as a capability in its Music 2.6 release, announced on April 10, 2026. The canonical model identifier used in the current agent tooling is music-cover. Although it is related to MiniMax's broader music-generation offering, it should not be confused with a standard text-to-music request: the cover workflow requires reference audio.

How the Cover Workflow Works

A typical request combines two inputs:

  • Reference audio: An existing song that provides the melodic and structural anchor.
  • Style instruction: A text prompt describing the desired musical treatment.

The reference recording supplies the melody and other musical cues. The text prompt tells MiniMax how to rebuild the song. For example, a user might request a jazz arrangement with brushed drums and a smoky vocal style, an orchestral cinematic version, or an electronic dance interpretation with a different production character.

MiniMax describes the process as extracting the melodic skeleton of the source and reconstructing the surrounding musical elements. This makes Music Cover different from a conventional remix, which normally retains more of the original recording, and different from voice cloning, whose primary goal is to reproduce a particular speaker or singer.

Reference Audio Requirements

Reference audio is required for the cover mode. MiniMax's current agent-skill documentation lists MP3, WAV, and FLAC as supported source formats. The documented guidance generally expects clips to be between six seconds and six minutes long, with a maximum file size of 50 MB.

Clear vocals and a recognizable melody are recommended. A source with heavy noise, unclear singing, substantial distortion, or a weak melodic structure may give the system less useful information to preserve. These duration and size values are workflow guidance from the available documentation rather than a general language-model context window.

Lyrics and Style Instructions

Lyrics can be supplied separately, but they are optional in the cover workflow. When lyrics are not provided, the surrounding MiniMax workflow can extract them from the source audio using automatic speech recognition. This makes it possible to work from a recording without manually transcribing every line.

The style prompt is the main creative control. It can describe genre, mood, instrumentation, vocal treatment, arrangement, and production qualities. More specific instructions are generally more useful than a single broad label. For example, a prompt that identifies the genre, desired instruments, vocal character, and atmosphere gives the system more direction than simply asking for a “different version.”

Main Capabilities

  • Reinterprets an existing song in another genre or production style.
  • Uses the source melody as the main musical constraint.
  • Changes vocals, instrumentation, arrangement, and mix direction.
  • Accepts optional replacement lyrics for a creative reinterpretation.
  • Generates audio output through the documented MiniMax music workflow.
  • Supports use cases such as alternate demos, creative experimentation, game or video concepts, advertising ideas, and licensed content production.

The model's core output is newly generated music audio. The supplied research identifies audio output formats such as MP3, WAV, and PCM in the documented workflow, but it does not provide a separate maximum output duration or a guaranteed output file-size limit.

Music Cover Versus Standard Music Generation

MiniMax's standard music-generation capability is intended to create original tracks from a text prompt and, where applicable, lyrics. Music Cover has a different purpose: it uses an existing recording to constrain the melodic identity of the result.

CharacteristicMiniMax Music CoverStandard music generation
Starting pointExisting reference audioText prompt and optional lyrics
Primary purposeReinterpretation and style transferOriginal song creation
MelodyIntended to remain recognizably related to the sourceGenerated as part of the new track
LyricsOptional, or potentially extracted from the sourceUsually supplied or generated as part of the music request
Best starting materialA clear song with an identifiable melodyA musical idea expressed through text

This distinction matters when choosing a workflow. If the goal is to hear an existing composition as a reggae, metal, orchestral, or electronic track, Music Cover is the more directly suitable option. If there is no reference recording and the goal is to invent a new song, standard music generation is a better fit.

Technical Profile and Limitations

Music Cover is a specialist audio-generation model, not a general-purpose language model. Its primary input is reference audio plus a text style instruction, and its primary output is generated music audio.

SpecificationAvailable information
ProviderMiniMax
Model identifiermusic-cover
Model familyMiniMax Music
Audio inputYes; MP3, WAV, and FLAC are documented
Text inputYes; used for style direction
Audio outputYes
Image or video input/outputNo documented support for this model
Context lengthNot applicable or not publicly specified
Maximum output tokensNot applicable
Native tool or function useNot documented for the model itself
JSON or structured outputNot applicable to the core audio result

There is no token-based context window or maximum-token setting for the core music result. The documented reference-audio duration and file-size guidance should not be interpreted as a text context limit. The available material also does not specify a maximum generated-audio duration, guaranteed latency, streaming behavior, fine-tuning support, or a model-specific caching or batch interface.

Music Cover is not designed for general reasoning, coding, chat, document analysis, or voice cloning. Its text-processing role is limited to interpreting the creative direction for the audio transformation. It should also not be treated as a precise multitrack editor: generated vocals, timing, instrumentation, phrasing, and mixing can vary from the source and from one generation to another.

Pricing and Availability

A conventional token price does not apply to this specialist audio-generation model. The supplied documentation indicates that audio-generation pricing can vary by MiniMax platform or deployment and may be charged per generated audio result where applicable. No verified universal per-generation price is provided in the available research.

MiniMax's broader product ecosystem includes consumer and developer-facing services, and availability, quotas, payment options, and model access can differ by country and product. The current Token Plan documentation states that music models became unavailable through Token Plan on August 20, 2026, while users may use the MiniMax Audio platform or open-weight releases instead. Because product access can change, users should check the current MiniMax service documentation before planning a production workflow.

Quality and Practical Trade-offs

The main benefit of Music Cover is control over musical identity: the user can begin with a melody that already works and explore alternative arrangements without asking the model to invent the entire song. This can make it useful for songwriting, arrangement exploration, creative pitching, and early-stage production.

The trade-off is that the output is a reinterpretation, not a guaranteed preservation of every detail in the original recording. Vocal phrasing may change, instruments may enter or leave at different times, and the mix may not match a professional multitrack production. The strongest results are likely to come from clean source material with a clear melody and a style direction that is specific without trying to control every individual note.

Compared with a conventional remix or digital audio workstation, Music Cover offers faster high-level transformation but less precise control over individual stems, effects, timing, and edits. Compared with text-to-music generation, it offers stronger melodic continuity but less freedom to start from nothing. Compared with voice-cloning tools, it focuses on rebuilding a complete musical performance rather than reproducing a particular singer's identity.

When to Choose MiniMax Music Cover

Choose MiniMax Music Cover when you have an authorized recording and want to explore how its melody might work in another musical setting. It is particularly appropriate for:

  • Testing alternate genres for an original song.
  • Creating arrangement demos before investing in a full production.
  • Developing music directions for videos, games, podcasts, or advertising.
  • Exploring new vocal and instrumental treatments while retaining a familiar melodic idea.
  • Producing personal or commercial reinterpretations for which you have the necessary rights.

A standard music-generation model may be more appropriate when no reference audio exists or when the desired result should have a completely original melody. A traditional digital audio workstation or professional remix workflow may be preferable when precise control over stems, timing, vocal takes, effects, and mastering is more important than rapid style exploration. A dedicated voice tool is a better choice when the primary requirement is reproducing a specific voice rather than transforming a whole song.

Rights and Responsible Use

Users should upload only audio they own, have licensed, or are otherwise authorized to process. A generated cover or reinterpretation can still involve rights in the original sound recording, composition, lyrics, and vocal performance. Changing the genre or arrangement does not remove copyright or permission obligations.

For that reason, Music Cover is best treated as a creative transformation tool whose suitability depends on both the quality of the reference recording and the user's rights to use it. Its clearest value is high-level musical experimentation: preserving a recognizable melodic foundation while testing a substantially different performance direction.


Answers to Frequently Asked Questions

Can I upload any song to MiniMax Music Cover?
No. Users should upload only audio they own, have licensed, or are otherwise authorized to process. A generated reinterpretation may still involve rights in the original recording, composition, lyrics, and vocal performance.
Can MiniMax Music Cover change a song’s lyrics, vocals, and instrumentation?
Yes. Users can provide optional replacement lyrics and describe changes to the genre, mood, instrumentation, arrangement, vocal character, tempo direction, and production style. Lyrics may also be extracted from the source audio when they are not supplied.
What audio formats and file limits does MiniMax Music Cover support?
The documented workflow supports MP3, WAV, and FLAC reference files. Source clips are generally expected to be between six seconds and six minutes long, with a maximum file size of 50 MB.
How is MiniMax Music Cover different from text-to-music generation?
Music Cover requires an existing reference recording and is designed to reinterpret its melody. Standard text-to-music generation creates a song from a text prompt and optional lyrics without needing a source track.
What is MiniMax Music Cover?
MiniMax Music Cover is an audio-to-audio music-generation model that transforms an existing song into a new genre, arrangement, vocal style, or production direction while aiming to preserve the source melody.


Sources 3
Provider

About MiniMax