What Is MiniMax Music Cover?
MiniMax Music Cover is a specialized audio-to-audio music-generation model provided by MiniMax. Its defining feature is that it starts with an existing song. Instead of asking the system to invent a melody from text, you upload a reference track and describe how you want that track reinterpreted.
The model is intended to preserve the recognizable melodic structure of the source while rebuilding the surrounding performance. Depending on the instruction, the result can change the genre, instrumentation, vocal character, arrangement, mood, tempo direction, and overall production style.
MiniMax introduced Cover as a capability in its Music 2.6 release, announced on April 10, 2026. The canonical model identifier used in the current agent tooling is music-cover. Although it is related to MiniMax's broader music-generation offering, it should not be confused with a standard text-to-music request: the cover workflow requires reference audio.
How the Cover Workflow Works
A typical request combines two inputs:
- Reference audio: An existing song that provides the melodic and structural anchor.
- Style instruction: A text prompt describing the desired musical treatment.
The reference recording supplies the melody and other musical cues. The text prompt tells MiniMax how to rebuild the song. For example, a user might request a jazz arrangement with brushed drums and a smoky vocal style, an orchestral cinematic version, or an electronic dance interpretation with a different production character.
MiniMax describes the process as extracting the melodic skeleton of the source and reconstructing the surrounding musical elements. This makes Music Cover different from a conventional remix, which normally retains more of the original recording, and different from voice cloning, whose primary goal is to reproduce a particular speaker or singer.
Reference Audio Requirements
Reference audio is required for the cover mode. MiniMax's current agent-skill documentation lists MP3, WAV, and FLAC as supported source formats. The documented guidance generally expects clips to be between six seconds and six minutes long, with a maximum file size of 50 MB.
Clear vocals and a recognizable melody are recommended. A source with heavy noise, unclear singing, substantial distortion, or a weak melodic structure may give the system less useful information to preserve. These duration and size values are workflow guidance from the available documentation rather than a general language-model context window.
Lyrics and Style Instructions
Lyrics can be supplied separately, but they are optional in the cover workflow. When lyrics are not provided, the surrounding MiniMax workflow can extract them from the source audio using automatic speech recognition. This makes it possible to work from a recording without manually transcribing every line.
The style prompt is the main creative control. It can describe genre, mood, instrumentation, vocal treatment, arrangement, and production qualities. More specific instructions are generally more useful than a single broad label. For example, a prompt that identifies the genre, desired instruments, vocal character, and atmosphere gives the system more direction than simply asking for a “different version.”
Main Capabilities
- Reinterprets an existing song in another genre or production style.
- Uses the source melody as the main musical constraint.
- Changes vocals, instrumentation, arrangement, and mix direction.
- Accepts optional replacement lyrics for a creative reinterpretation.
- Generates audio output through the documented MiniMax music workflow.
- Supports use cases such as alternate demos, creative experimentation, game or video concepts, advertising ideas, and licensed content production.
The model's core output is newly generated music audio. The supplied research identifies audio output formats such as MP3, WAV, and PCM in the documented workflow, but it does not provide a separate maximum output duration or a guaranteed output file-size limit.
Music Cover Versus Standard Music Generation
MiniMax's standard music-generation capability is intended to create original tracks from a text prompt and, where applicable, lyrics. Music Cover has a different purpose: it uses an existing recording to constrain the melodic identity of the result.
| Characteristic | MiniMax Music Cover | Standard music generation |
|---|---|---|
| Starting point | Existing reference audio | Text prompt and optional lyrics |
| Primary purpose | Reinterpretation and style transfer | Original song creation |
| Melody | Intended to remain recognizably related to the source | Generated as part of the new track |
| Lyrics | Optional, or potentially extracted from the source | Usually supplied or generated as part of the music request |
| Best starting material | A clear song with an identifiable melody | A musical idea expressed through text |
This distinction matters when choosing a workflow. If the goal is to hear an existing composition as a reggae, metal, orchestral, or electronic track, Music Cover is the more directly suitable option. If there is no reference recording and the goal is to invent a new song, standard music generation is a better fit.
Technical Profile and Limitations
Music Cover is a specialist audio-generation model, not a general-purpose language model. Its primary input is reference audio plus a text style instruction, and its primary output is generated music audio.
| Specification | Available information |
|---|---|
| Provider | MiniMax |
| Model identifier | music-cover |
| Model family | MiniMax Music |
| Audio input | Yes; MP3, WAV, and FLAC are documented |
| Text input | Yes; used for style direction |
| Audio output | Yes |
| Image or video input/output | No documented support for this model |
| Context length | Not applicable or not publicly specified |
| Maximum output tokens | Not applicable |
| Native tool or function use | Not documented for the model itself |
| JSON or structured output | Not applicable to the core audio result |
There is no token-based context window or maximum-token setting for the core music result. The documented reference-audio duration and file-size guidance should not be interpreted as a text context limit. The available material also does not specify a maximum generated-audio duration, guaranteed latency, streaming behavior, fine-tuning support, or a model-specific caching or batch interface.
Music Cover is not designed for general reasoning, coding, chat, document analysis, or voice cloning. Its text-processing role is limited to interpreting the creative direction for the audio transformation. It should also not be treated as a precise multitrack editor: generated vocals, timing, instrumentation, phrasing, and mixing can vary from the source and from one generation to another.
Pricing and Availability
A conventional token price does not apply to this specialist audio-generation model. The supplied documentation indicates that audio-generation pricing can vary by MiniMax platform or deployment and may be charged per generated audio result where applicable. No verified universal per-generation price is provided in the available research.
MiniMax's broader product ecosystem includes consumer and developer-facing services, and availability, quotas, payment options, and model access can differ by country and product. The current Token Plan documentation states that music models became unavailable through Token Plan on August 20, 2026, while users may use the MiniMax Audio platform or open-weight releases instead. Because product access can change, users should check the current MiniMax service documentation before planning a production workflow.
Quality and Practical Trade-offs
The main benefit of Music Cover is control over musical identity: the user can begin with a melody that already works and explore alternative arrangements without asking the model to invent the entire song. This can make it useful for songwriting, arrangement exploration, creative pitching, and early-stage production.
The trade-off is that the output is a reinterpretation, not a guaranteed preservation of every detail in the original recording. Vocal phrasing may change, instruments may enter or leave at different times, and the mix may not match a professional multitrack production. The strongest results are likely to come from clean source material with a clear melody and a style direction that is specific without trying to control every individual note.
Compared with a conventional remix or digital audio workstation, Music Cover offers faster high-level transformation but less precise control over individual stems, effects, timing, and edits. Compared with text-to-music generation, it offers stronger melodic continuity but less freedom to start from nothing. Compared with voice-cloning tools, it focuses on rebuilding a complete musical performance rather than reproducing a particular singer's identity.
When to Choose MiniMax Music Cover
Choose MiniMax Music Cover when you have an authorized recording and want to explore how its melody might work in another musical setting. It is particularly appropriate for:
- Testing alternate genres for an original song.
- Creating arrangement demos before investing in a full production.
- Developing music directions for videos, games, podcasts, or advertising.
- Exploring new vocal and instrumental treatments while retaining a familiar melodic idea.
- Producing personal or commercial reinterpretations for which you have the necessary rights.
A standard music-generation model may be more appropriate when no reference audio exists or when the desired result should have a completely original melody. A traditional digital audio workstation or professional remix workflow may be preferable when precise control over stems, timing, vocal takes, effects, and mastering is more important than rapid style exploration. A dedicated voice tool is a better choice when the primary requirement is reproducing a specific voice rather than transforming a whole song.
Rights and Responsible Use
Users should upload only audio they own, have licensed, or are otherwise authorized to process. A generated cover or reinterpretation can still involve rights in the original sound recording, composition, lyrics, and vocal performance. Changing the genre or arrangement does not remove copyright or permission obligations.
For that reason, Music Cover is best treated as a creative transformation tool whose suitability depends on both the quality of the reference recording and the user's rights to use it. Its clearest value is high-level musical experimentation: preserving a recognizable melodic foundation while testing a substantially different performance direction.

