MiniMax ASR

ASR 1.0

by MiniMax · Current public model

MiniMax ASR 1.0 is a specialized speech-to-text model for uploaded audio. It supports 19 documented language selections, automatic mixed-language recognition, speaker diarization, timestamps, subtitle export, and streaming JSON responses. Requests are limited to 500 seconds and 50 MB, with pricing based on processed audio duration at $0.38 per hour.

Text Reasoning Coding
MiniMax ASR 1.0 is MiniMax's public speech-to-text model for turning supported audio recordings into searchable text, speaker-labeled transcripts, or subtitles. It is aimed at applications such as meeting transcription, podcast processing, multilingual media workflows, live captions, and call analysis. The model accepts common audio formats, supports automatic recognition of mixed-language speech, and can return either a straightforward transcript or more detailed timestamped and diarized output.
Outputs

What ASR 1.0 can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family MiniMax ASR
Model type Other
Status Current public model
Knowledge cutoff notes

Knowledge cutoff is not applicable to this speech-recognition model; MiniMax documents transcription capabilities and input limitations rather than a text knowledge cutoff.

Model notes

The canonical API model identifier is asr-1.0. The model accepts WAV, AIFF, FLAC, ALAC in M4A, MP3, AAC, Opus, and Ogg containers up to 500 seconds and 50 MB per request. It supports JSON, verbose JSON, SRT, and VTT responses. Verbose JSON provides speaker diarization and timestamps. Streaming uses server-sent events and supports JSON output only; it is file-upload streaming rather than native microphone audio streaming. The API documentation lists 19 supported language hints and automatic mixed-language recognition when the language hint is omitted. Editorial capability scores are not vendor benchmarks.

Cost

Model pricing

Input $0.38 per hour of processed audio
Output No separate output charge; billing is based on input audio duration
Model guide

MiniMax ASR 1.0: Multilingual Speech-to-Text with Diarization and Subtitles

MiniMax ASR 1.0 is a developer-facing automatic speech recognition model for converting audio files into text. It supports multilingual and mixed-language transcription, streaming JSON responses, speaker diarization, timestamps, and direct SRT or WebVTT subtitle output. Requests can contain up to 500 seconds or 50 MB of audio, and pricing is based on processed audio duration at $0.38 per hour.

What is MiniMax ASR 1.0?

MiniMax ASR 1.0 is an automatic speech recognition model from MiniMax. Automatic speech recognition, or ASR, converts spoken audio into written text. Unlike a general-purpose language model, ASR 1.0 is specialized for listening to an uploaded audio file and producing a transcription rather than writing essays, answering questions, generating code, or creating media.

The model is available through MiniMax's developer-facing speech-to-text service. Its canonical model identifier is asr-1.0, and requests use the /v1/speech_to_text endpoint with multipart form data. A MiniMax API key is supplied through Bearer authentication.

Its main practical distinction is the combination of multilingual transcription, optional speaker and time alignment, subtitle export, and file-upload streaming. That makes it useful when an application needs more than a single block of recognized text.

Where ASR 1.0 fits in the MiniMax lineup

ASR 1.0 belongs to MiniMax's developer-oriented audio and speech capabilities rather than its general text, image, video, music, or agent products. It is a recognition model: it consumes audio and produces text or subtitle-oriented transcription data.

This distinction matters when evaluating the model. MiniMax also offers products and models for speech synthesis, voice and music workflows, video generation, coding, and agentic tasks, but those capabilities should not be attributed to ASR 1.0. The supplied documentation does not describe this model as a conversational assistant, a speech generator, or a general multimodal reasoning system.

Supported audio formats, limits, and languages

Each request can contain up to 500 seconds of audio and 50 MB of data. The service accepts WAV, AIFF, FLAC, ALAC in an M4A container, MP3, AAC, Opus, and Ogg. Raw PCM without a supported container is not accepted.

For large or high-fidelity recordings, MiniMax recommends mono 16 kHz audio or a compressed format to help remain below the 50 MB limit. The documentation states that higher sample rates and stereo audio do not improve recognition quality for this service. This is a practical file-preparation recommendation rather than a claim that every source recording must be converted before use.

The API provides language hints for Chinese, Cantonese, English, Japanese, Korean, Thai, Vietnamese, Indonesian, Malay, Filipino, Arabic, Turkish, French, German, Spanish, Italian, Portuguese, Polish, Russian, and Ukrainian. When the language parameter is omitted or empty, ASR 1.0 can automatically recognize multilingual or code-switched audio. Explicitly setting the expected language may still be useful for short clips or recordings containing specialized terminology.

Transcription and output options

The basic JSON response includes the complete transcription, detected audio duration, and a trace identifier. This is the simplest option for applications that need plain text, such as searchable archives, rough notes, or downstream text processing.

Applications can request verbose_json when they need more structure. This format provides detected speaker information and timestamped segments. Timestamp granularity can be configured for sentence- or segment-level output, as well as word-level timestamps for English and character-level timestamps for Chinese.

ASR 1.0 can also return subtitle files directly in SRT or WebVTT format. These formats are useful for videos, recorded presentations, online courses, and media review workflows because they package recognized text with timing information in a format commonly accepted by subtitle tools and players.

Speaker diarization and timestamps

Speaker diarization means estimating which portions of a recording belong to which speaker. With verbose_json, the response includes the number of detected speakers and timestamped segments containing speaker labels. This can make a meeting or interview transcript easier to review than an undifferentiated block of text.

Diarization should not be interpreted as perfect identity recognition. The supplied documentation describes speaker labels and detected speaker counts, but does not establish a benchmark for speaker-attribution accuracy. Recordings with overlapping speech, background noise, or very similar voices may require human review.

Streaming behavior

Streaming mode sends incremental transcription events using server-sent events, allowing an application to display partial results before the entire request has finished processing. This can help with low-latency captions or interfaces that need progressive feedback.

Streaming currently supports JSON output only. SRT, WebVTT, and diarization-oriented formats cannot be combined with streaming because they require alignment and segment processing before the final result is assembled. The streaming feature is also file-upload streaming, not documented native microphone capture or a direct real-time audio socket.

Pricing and cost profile

MiniMax lists ASR 1.0 at $0.38 per hour of processed audio. Billing is based on the duration reported for the input audio, not on the number of generated text tokens. There is no separate output charge identified in the supplied pricing information.

This pricing model makes cost relatively predictable for transcription workloads: a longer recording costs more because it contains more audio, while the amount of text produced does not create a separate token bill. Actual project costs can still depend on how many files are submitted, whether recordings are retried, and the account or regional terms applied by MiniMax.

Capabilities and limitations

AreaASR 1.0 status
Primary taskAutomatic speech recognition and audio transcription
Audio inputSupported uploaded audio files
Text outputPlain transcription, verbose JSON, SRT, or WebVTT
StreamingSupported through incremental server-sent events; JSON only
Speaker diarizationSupported in verbose JSON
Maximum request500 seconds and 50 MB per request
Language handling19 documented language selections plus automatic mixed-language recognition
Context windowNot published for this speech-recognition model
Maximum output tokensNot published and not expressed as a token-generation limit
Tool or function callingNot supported or documented for this model
Fine-tuning and batch APINot published for this exact model

ASR 1.0 does not generate images, audio, video, or speech. It is not documented as having reasoning, coding, web-search, tool-use, prompt-caching, or fine-tuning capabilities. Editorial capability scores may rate it highly for transcription speed or cost, but those are evaluations rather than MiniMax-published benchmarks.

Main strengths and trade-offs

The model's strongest feature is task specialization. A service that needs transcripts, speaker segments, or subtitles can use output formats designed for those jobs instead of asking a general language model to interpret an audio file indirectly. The supported language list and automatic mixed-language recognition are also useful for international meetings, interviews, and media containing code-switching.

Its output flexibility is another advantage. Plain JSON is suitable for application logic, verbose JSON exposes timing and speaker information, and SRT or WebVTT can move directly into subtitle workflows. Streaming JSON adds a lower-latency option when an application can work with partial transcription events.

The main trade-off is that every request is bounded by 500 seconds and 50 MB. Longer recordings must be divided into multiple requests, which can require application-side file management and careful handling of segment boundaries. Streaming does not remove this limitation, and it does not provide the documented functionality of a native microphone-streaming service.

ASR 1.0 also does not replace a general language model. It recognizes speech but does not, according to the supplied specifications, summarize a meeting, answer questions about a transcript, execute tools, or write application code. Those tasks would require a separate processing step.

Best use cases

  • Meetings and interviews: Use verbose JSON when speaker labels and timestamps make review easier.
  • Podcasts, webinars, and recorded presentations: Convert spoken content into searchable text or prepare it for editing.
  • Subtitle production: Request SRT or WebVTT when the output will be attached to a video.
  • Live captions and progressive interfaces: Use streaming JSON when partial transcription is more useful than waiting for a completed response.
  • Multilingual media: Use the documented language hints or automatic recognition for recordings containing more than one language.
  • Call-quality analysis and archives: Generate text for later search, indexing, or analysis by another system.

When to choose MiniMax ASR 1.0

Choose ASR 1.0 when the central requirement is affordable file-based transcription with multilingual support, optional speaker segmentation, timestamps, subtitle export, or incremental JSON results. Its audio-duration pricing is straightforward for teams estimating transcription costs, and its format choices reduce the amount of post-processing needed for common media workflows.

A different type of option may be more appropriate when the application requires a documented native microphone or telephone stream, very long single-file processing without client-side splitting, domain-specific customization, or integrated summarization and question answering. The supplied research does not establish that ASR 1.0 provides those capabilities, so they should not be assumed from its streaming feature.

For workflows that need both transcription and interpretation, ASR 1.0 can serve as the speech-recognition stage, while a separate text model or application component handles summarization, extraction, translation, or actions. Keeping those roles separate also makes it easier to inspect the original transcript and diagnose recognition errors.

Practical implementation guidance

Upload each recording as multipart form data, identify the model as asr-1.0, and authenticate with a MiniMax API key. Select basic JSON for a simple transcript, verbose_json for speaker and timing data, or SRT and WebVTT when the output is intended for subtitles. Use streaming only when incremental JSON events are useful to the application.

Before submission, check both the duration and file size. Recordings longer than 500 seconds or larger than 50 MB need to be split or compressed. Preserve enough surrounding context when splitting audio so that words are not cut at boundaries, and validate timestamps if separate transcript segments will later be combined.


Answers to Frequently Asked Questions

How much does MiniMax ASR 1.0 cost?
MiniMax lists ASR 1.0 at $0.38 per hour of processed audio. Billing is based on the duration of the input audio rather than the number of generated text tokens, although retries, file volume, account terms, and regional conditions may affect total project costs.
Can MiniMax ASR 1.0 generate SRT and WebVTT subtitles?
Yes. MiniMax ASR 1.0 can return recognized speech as SRT or WebVTT subtitle files, making it suitable for videos, online courses, presentations, podcasts, and other media workflows. Subtitle formats are not available with streaming mode, which supports JSON only.
Does MiniMax ASR 1.0 support speaker diarization and timestamps?
Yes. When the response format is set to verbose JSON, MiniMax ASR 1.0 provides detected speaker counts, speaker labels, and timestamped segments. It supports sentence- or segment-level timestamps, as well as word-level timestamps for English and character-level timestamps for Chinese.
Which audio formats, languages, and file limits does MiniMax ASR 1.0 support?
The service accepts WAV, AIFF, FLAC, ALAC in an M4A container, MP3, AAC, Opus, and Ogg files. Each request is limited to 500 seconds and 50 MB. It provides language hints for 19 languages and can automatically recognize multilingual or code-switched audio when no language is specified.
What is MiniMax ASR 1.0 used for?
MiniMax ASR 1.0 is an automatic speech recognition model that converts uploaded audio into written transcripts. It supports plain text transcription, speaker diarization, timestamps, subtitle export, and incremental JSON results through streaming.


Sources 4
Provider

About MiniMax