What is GLM-ASR-2512?
GLM-ASR-2512 is Z.AI’s cloud-based automatic speech recognition model. Automatic speech recognition, or ASR, converts spoken audio into written text. Unlike a general-purpose language model, GLM-ASR-2512 is focused on recognizing speech and returning a transcription rather than answering questions, writing code, generating images, or operating tools.
The model is available through Z.AI’s audio transcription API. It is intended for applications that need hosted speech recognition without deploying an audio model locally. Typical examples include live captions, short voice commands, meeting excerpts, customer-service quality checks, office dictation, and transcription of specialized vocabulary.
Z.AI identifies the model as glm-asr-2512. It belongs to the GLM-ASR family and should be distinguished from GLM-ASR-Nano-2512, an open-weight model described separately in Zhipu AI’s research material. The subject of this page is the hosted GLM-ASR-2512 API model.
How the model processes audio
GLM-ASR-2512 accepts audio input and produces text output. The official API documentation lists WAV and MP3 support, with a maximum file size of 25 MB and a maximum audio duration of 30 seconds for each request. Longer recordings therefore need to be divided into shorter segments by the calling application before transcription.
The API supports both synchronous and streaming transcription. In synchronous mode, an application waits for the transcription result. In streaming mode, the service sends transcription content in chunks through an event stream as processing progresses. Streaming is useful for captions, voice interfaces, and other experiences where waiting for the entire audio segment would add noticeable delay.
The response includes identifying and status information such as a task identifier, creation timestamp, request identifier, model name, and the completed transcription text. The supplied research does not document a general context window or maximum output-token value for this model. Those omissions are expected for an audio transcription endpoint whose primary limit is expressed in audio duration rather than text tokens.
Languages, dialects, and recognition quality
Z.AI describes GLM-ASR-2512 as a multilingual and dialect-aware recognition model. The documented Chinese varieties include Mandarin, Sichuanese, Cantonese, Min Nan, and Wu. The documentation also lists English accents and languages including French, German, Japanese, Korean, and Spanish.
This language and dialect coverage is one of the model’s main practical distinctions. It may be a better fit than a narrowly focused speech recognizer when an application must handle multilingual recordings, regional Chinese speech, or conversations that mix Chinese and English. The provider also highlights performance in mixed Chinese-English speech, command-style text, industry terminology, long sentences, colloquial speech, and varied accents.
Z.AI reports a character error rate of 0.0717 in its latest competitive evaluation. Character error rate is a measure of transcription mistakes relative to the reference text; lower values generally indicate fewer character-level errors. This figure is a provider-reported evaluation result, not a guarantee for every language, microphone, speaker, accent, noise condition, or application. Real-world accuracy should be tested with representative recordings before production use.
Prompts and hotwords for specialized vocabulary
GLM-ASR-2512 supports a prompt parameter that can provide previous transcription context. This can help application developers maintain continuity when processing longer material in multiple short segments, although the documented audio-duration limit still applies to each individual request. Z.AI recommends keeping the prompt below 8,000 characters.
The API also supports hotwords. Hotwords are terms that the recognizer should pay particular attention to, such as names, product names, locations, project codes, medical terminology, or internal abbreviations. The documentation recommends a maximum of 100 hotwords.
For example, a customer-support application could supply the names of products and service plans that commonly occur in calls. A meeting transcription system could provide participant names, department names, or project identifiers. Hotwords are a recognition aid, not a guarantee that every term will be transcribed correctly, especially when the recording is noisy or the speaker is unclear.
API access and integration
GLM-ASR-2512 is accessed through Z.AI’s audio transcription endpoint:
/api/paas/v4/audio/transcriptions
Requests use multipart form data and require a Z.AI API bearer token, an audio file or base64-encoded audio, and the model identifier glm-asr-2512. The stream parameter determines whether the application receives a completed response or event-stream updates.
This makes the model suitable for backend services, captioning pipelines, voice-entry features, and batch-style processing of short audio clips. Developers should design around the 30-second and 25 MB limits rather than assuming that a single request can handle a complete meeting, call, interview, or lecture. Segmenting longer audio also creates application-level concerns such as preserving context, ordering partial results, and avoiding duplicated text at segment boundaries.
Main strengths and limitations
Strengths
- Focused transcription: The model is purpose-built for converting speech into text rather than sharing resources with unrelated generation tasks.
- Multilingual and dialect support: Z.AI documents several Chinese varieties, English accents, and multiple additional languages.
- Real-time delivery: Streaming responses can support captions and interactive voice workflows.
- Terminology controls: Prompts and up to 100 recommended hotwords can help with names, codes, and industry-specific vocabulary.
- Hosted deployment: Applications can use the model through an API without managing speech-recognition weights or local inference infrastructure.
- Short-request workflow: The 30-second limit can work well for commands, snippets, captions, and segmented audio pipelines.
Limitations
- Short audio limit: Each request is limited to 30 seconds and 25 MB according to the supplied documentation.
- Not a general-purpose assistant: The model does not provide documented reasoning, coding, image generation, audio generation, embeddings, or tool/function calling capabilities.
- No documented general context window: Z.AI does not publish a conventional context-length or maximum-output-token specification for this endpoint.
- Cloud dependency: The hosted API requires network access and an active Z.AI account or API arrangement; it is not the same as running an open-weight model locally.
- Pricing uncertainty: The exact model documentation does not publish a detailed official price. A secondary model listing reports approximately CNY 0.06 per audio minute, but this should be confirmed in the current Z.AI billing interface.
- Accuracy varies by conditions: Provider-reported evaluation results may not predict performance for every accent, language, recording device, background-noise level, or speaking style.
Reasoning, coding, and tool support
GLM-ASR-2512 should not be selected for reasoning or coding. Its output is transcribed text, not a reasoned answer or generated program. The supplied model information does not document tool use, function calling, web search, structured output, batch processing, or fine-tuning for this model.
A practical architecture can still place the transcription model inside a larger workflow. For example, an application may use GLM-ASR-2512 to convert a voice recording to text and then pass that text to a separate language model for summarization, extraction, classification, or question answering. Those later capabilities would come from the additional model and application logic, not from GLM-ASR-2512 itself.
Pricing and cost trade-offs
The available research reports a transcription price of approximately CNY 0.06 per audio minute. This is a secondary listing rather than a clearly published price in the supplied official model documentation, so it should be treated as an indication rather than a guaranteed current rate. Regional billing, account terms, quotas, and future pricing changes may affect the actual cost.
Because billing is reported by audio duration rather than input and output tokens, cost estimation is relatively straightforward for a transcription pipeline: estimate the total audio minutes, account for any repeated processing or retries, and verify the active rate in the Z.AI account. The model may be economically attractive for short hosted transcription tasks, but applications processing large archives should compare per-minute pricing, concurrency, storage, and segmentation overhead with other speech-to-text services or local inference options.
Streaming can improve responsiveness but does not by itself imply lower cost. It should be chosen when partial results are valuable, such as live captions or voice interfaces. Synchronous processing may be simpler for short clips where the user can wait for the completed transcript.
Best use cases for GLM-ASR-2512
- Live captions for short conversations, presentations, or voice interactions
- Voice commands and short spoken input fields
- Meeting snippets and segmented meeting-note workflows
- Customer-service call sampling and quality assurance
- Dictation into office, CRM, or documentation systems
- Multilingual transcription involving supported languages or Chinese dialects
- Specialized transcription where names, product terms, codes, or other hotwords matter
- Applications that prefer an API over deploying speech-recognition infrastructure locally
When should you choose GLM-ASR-2512?
Choose GLM-ASR-2512 when the primary requirement is hosted multilingual speech-to-text, especially if streaming responses, Chinese dialect coverage, short audio segments, or custom hotwords are important. It is also a reasonable candidate for teams already using Z.AI and wanting transcription through the same provider’s API ecosystem.
Another speech-recognition option may be more appropriate when recordings routinely exceed 30 seconds and the application cannot conveniently segment them, when offline or local processing is mandatory, or when an independently verified price, compliance policy, accuracy benchmark, or language list is required. A general-purpose language model is more suitable when the main task begins after transcription and requires summarization, reasoning, coding, or tool execution.
For production deployment, test representative audio rather than relying only on the reported 0.0717 character error rate. Include the languages, dialects, speakers, background noise, microphone types, terminology, and segment lengths expected in the real application. Also confirm the current API limits, availability, data-handling terms, and per-minute price before committing to a high-volume workflow.

