What is MiniMax ASR 1.0?
MiniMax ASR 1.0 is an automatic speech recognition model from MiniMax. Automatic speech recognition, or ASR, converts spoken audio into written text. Unlike a general-purpose language model, ASR 1.0 is specialized for listening to an uploaded audio file and producing a transcription rather than writing essays, answering questions, generating code, or creating media.
The model is available through MiniMax's developer-facing speech-to-text service. Its canonical model identifier is asr-1.0, and requests use the /v1/speech_to_text endpoint with multipart form data. A MiniMax API key is supplied through Bearer authentication.
Its main practical distinction is the combination of multilingual transcription, optional speaker and time alignment, subtitle export, and file-upload streaming. That makes it useful when an application needs more than a single block of recognized text.
Where ASR 1.0 fits in the MiniMax lineup
ASR 1.0 belongs to MiniMax's developer-oriented audio and speech capabilities rather than its general text, image, video, music, or agent products. It is a recognition model: it consumes audio and produces text or subtitle-oriented transcription data.
This distinction matters when evaluating the model. MiniMax also offers products and models for speech synthesis, voice and music workflows, video generation, coding, and agentic tasks, but those capabilities should not be attributed to ASR 1.0. The supplied documentation does not describe this model as a conversational assistant, a speech generator, or a general multimodal reasoning system.
Supported audio formats, limits, and languages
Each request can contain up to 500 seconds of audio and 50 MB of data. The service accepts WAV, AIFF, FLAC, ALAC in an M4A container, MP3, AAC, Opus, and Ogg. Raw PCM without a supported container is not accepted.
For large or high-fidelity recordings, MiniMax recommends mono 16 kHz audio or a compressed format to help remain below the 50 MB limit. The documentation states that higher sample rates and stereo audio do not improve recognition quality for this service. This is a practical file-preparation recommendation rather than a claim that every source recording must be converted before use.
The API provides language hints for Chinese, Cantonese, English, Japanese, Korean, Thai, Vietnamese, Indonesian, Malay, Filipino, Arabic, Turkish, French, German, Spanish, Italian, Portuguese, Polish, Russian, and Ukrainian. When the language parameter is omitted or empty, ASR 1.0 can automatically recognize multilingual or code-switched audio. Explicitly setting the expected language may still be useful for short clips or recordings containing specialized terminology.
Transcription and output options
The basic JSON response includes the complete transcription, detected audio duration, and a trace identifier. This is the simplest option for applications that need plain text, such as searchable archives, rough notes, or downstream text processing.
Applications can request verbose_json when they need more structure. This format provides detected speaker information and timestamped segments. Timestamp granularity can be configured for sentence- or segment-level output, as well as word-level timestamps for English and character-level timestamps for Chinese.
ASR 1.0 can also return subtitle files directly in SRT or WebVTT format. These formats are useful for videos, recorded presentations, online courses, and media review workflows because they package recognized text with timing information in a format commonly accepted by subtitle tools and players.
Speaker diarization and timestamps
Speaker diarization means estimating which portions of a recording belong to which speaker. With verbose_json, the response includes the number of detected speakers and timestamped segments containing speaker labels. This can make a meeting or interview transcript easier to review than an undifferentiated block of text.
Diarization should not be interpreted as perfect identity recognition. The supplied documentation describes speaker labels and detected speaker counts, but does not establish a benchmark for speaker-attribution accuracy. Recordings with overlapping speech, background noise, or very similar voices may require human review.
Streaming behavior
Streaming mode sends incremental transcription events using server-sent events, allowing an application to display partial results before the entire request has finished processing. This can help with low-latency captions or interfaces that need progressive feedback.
Streaming currently supports JSON output only. SRT, WebVTT, and diarization-oriented formats cannot be combined with streaming because they require alignment and segment processing before the final result is assembled. The streaming feature is also file-upload streaming, not documented native microphone capture or a direct real-time audio socket.
Pricing and cost profile
MiniMax lists ASR 1.0 at $0.38 per hour of processed audio. Billing is based on the duration reported for the input audio, not on the number of generated text tokens. There is no separate output charge identified in the supplied pricing information.
This pricing model makes cost relatively predictable for transcription workloads: a longer recording costs more because it contains more audio, while the amount of text produced does not create a separate token bill. Actual project costs can still depend on how many files are submitted, whether recordings are retried, and the account or regional terms applied by MiniMax.
Capabilities and limitations
| Area | ASR 1.0 status |
|---|---|
| Primary task | Automatic speech recognition and audio transcription |
| Audio input | Supported uploaded audio files |
| Text output | Plain transcription, verbose JSON, SRT, or WebVTT |
| Streaming | Supported through incremental server-sent events; JSON only |
| Speaker diarization | Supported in verbose JSON |
| Maximum request | 500 seconds and 50 MB per request |
| Language handling | 19 documented language selections plus automatic mixed-language recognition |
| Context window | Not published for this speech-recognition model |
| Maximum output tokens | Not published and not expressed as a token-generation limit |
| Tool or function calling | Not supported or documented for this model |
| Fine-tuning and batch API | Not published for this exact model |
ASR 1.0 does not generate images, audio, video, or speech. It is not documented as having reasoning, coding, web-search, tool-use, prompt-caching, or fine-tuning capabilities. Editorial capability scores may rate it highly for transcription speed or cost, but those are evaluations rather than MiniMax-published benchmarks.
Main strengths and trade-offs
The model's strongest feature is task specialization. A service that needs transcripts, speaker segments, or subtitles can use output formats designed for those jobs instead of asking a general language model to interpret an audio file indirectly. The supported language list and automatic mixed-language recognition are also useful for international meetings, interviews, and media containing code-switching.
Its output flexibility is another advantage. Plain JSON is suitable for application logic, verbose JSON exposes timing and speaker information, and SRT or WebVTT can move directly into subtitle workflows. Streaming JSON adds a lower-latency option when an application can work with partial transcription events.
The main trade-off is that every request is bounded by 500 seconds and 50 MB. Longer recordings must be divided into multiple requests, which can require application-side file management and careful handling of segment boundaries. Streaming does not remove this limitation, and it does not provide the documented functionality of a native microphone-streaming service.
ASR 1.0 also does not replace a general language model. It recognizes speech but does not, according to the supplied specifications, summarize a meeting, answer questions about a transcript, execute tools, or write application code. Those tasks would require a separate processing step.
Best use cases
- Meetings and interviews: Use verbose JSON when speaker labels and timestamps make review easier.
- Podcasts, webinars, and recorded presentations: Convert spoken content into searchable text or prepare it for editing.
- Subtitle production: Request SRT or WebVTT when the output will be attached to a video.
- Live captions and progressive interfaces: Use streaming JSON when partial transcription is more useful than waiting for a completed response.
- Multilingual media: Use the documented language hints or automatic recognition for recordings containing more than one language.
- Call-quality analysis and archives: Generate text for later search, indexing, or analysis by another system.
When to choose MiniMax ASR 1.0
Choose ASR 1.0 when the central requirement is affordable file-based transcription with multilingual support, optional speaker segmentation, timestamps, subtitle export, or incremental JSON results. Its audio-duration pricing is straightforward for teams estimating transcription costs, and its format choices reduce the amount of post-processing needed for common media workflows.
A different type of option may be more appropriate when the application requires a documented native microphone or telephone stream, very long single-file processing without client-side splitting, domain-specific customization, or integrated summarization and question answering. The supplied research does not establish that ASR 1.0 provides those capabilities, so they should not be assumed from its streaming feature.
For workflows that need both transcription and interpretation, ASR 1.0 can serve as the speech-recognition stage, while a separate text model or application component handles summarization, extraction, translation, or actions. Keeping those roles separate also makes it easier to inspect the original transcript and diagnose recognition errors.
Practical implementation guidance
Upload each recording as multipart form data, identify the model as asr-1.0, and authenticate with a MiniMax API key. Select basic JSON for a simple transcript, verbose_json for speaker and timing data, or SRT and WebVTT when the output is intended for subtitles. Use streaming only when incremental JSON events are useful to the application.
Before submission, check both the duration and file size. Recordings longer than 500 seconds or larger than 50 MB need to be split or compressed. Preserve enough surrounding context when splitting audio so that words are not cut at boundaries, and validate timestamps if separate transcript segments will later be combined.

