What is Grok Voice Transcribe 2.0?
Grok Voice Transcribe 2.0 is xAI’s current model for automatic speech recognition. Its primary job is to listen to audio and produce a written transcript, rather than to hold a conversation, generate general text or perform broad reasoning. The model accepts audio input and returns text together with transcript metadata such as timestamps, confidence values, channel information and speaker labels.
xAI provides the model through its Speech to Text API. It can process uploaded or referenced audio in batch mode, or transcribe an active audio stream over a WebSocket connection. The exact API model identifier is grok-voice-transcribe-2.0.
Within xAI’s current speech-to-text lineup, Grok Voice Transcribe 2.0 is the default model. It replaces Grok Voice Transcribe 1.0, which xAI announced for deprecation. This positioning makes version 2.0 the relevant choice for new integrations that need xAI’s transcription service.
Core capabilities for practical transcription
The model is intended for audio that is more complicated than a clean, single-speaker recording. xAI describes support for telephone calls, conversations with overlapping or competing voices, accents, noisy surroundings and spoken information such as phone numbers and email addresses.
- Automatic language detection: xAI documents support for more than 38 languages without requiring the application to identify the language in advance.
- Word-level timestamps: Individual words can be associated with their position in the audio, which is useful for captions, search, editing and playback synchronization.
- Confidence scores: Transcript results can include an indication of how certain the system is about recognized content.
- Speaker diarization: The service can label speech by speaker, helping separate participants in interviews, meetings and calls.
- Multichannel transcription: Up to eight audio channels are supported, which can be useful when different microphones or call participants are recorded separately.
- Key-term biasing: Applications can provide important vocabulary to improve recognition of domain-specific names, products or terminology.
- Text formatting: The API can format numbers, dates, currencies, phone numbers and email addresses in more readable forms.
- Filler-word removal: Optional removal of words such as conversational fillers can produce cleaner transcripts.
- Smart turn detection: The model supports turn detection for voice-agent workflows, where the application needs to determine when a speaker has finished speaking.
These features distinguish Grok Voice Transcribe 2.0 from a minimal speech-recognition endpoint that returns only an unstructured block of text. Developers can use the additional metadata to build searchable recordings, synchronized captions, speaker-attributed notes and responsive voice interfaces.
Batch and streaming modes
Grok Voice Transcribe 2.0 supports two operating patterns. Batch transcription is designed for audio that already exists, such as a meeting recording, podcast, uploaded video soundtrack or archived support call. The application submits the audio through the REST API and receives a completed transcription.
Streaming transcription is designed for live or near-live use. Audio is sent through a WebSocket connection while it is being recorded, and the service returns interim transcript events followed by completed transcript results. Interim text allows an application to display captions or begin processing speech before the speaker has finished. However, earlier interim text may be corrected as additional audio arrives, so applications should treat provisional results differently from finalized segments.
xAI documents support for common audio formats including WAV, MP3, WebM, OGG and M4A. The service is available in the us-east-1 region. Documented service limits include 10 requests per second for REST use and 10 requests per second for streaming use, with up to 100 concurrent streaming sessions per team.
Pricing and cost trade-offs
xAI prices Grok Voice Transcribe 2.0 by the amount of audio processed rather than by the number of generated text tokens. Batch REST transcription costs $0.10 per hour of audio. Real-time streaming transcription costs $0.20 per hour of audio.
| Mode | Price | Typical purpose |
|---|---|---|
| Batch REST transcription | $0.10 per hour of audio | Recorded meetings, calls, interviews and media archives |
| Real-time WebSocket transcription | $0.20 per hour of audio | Live captions, voice agents and interactive applications |
The streaming rate is twice the batch rate, but it supports a different operational requirement: receiving transcript events while audio is being produced. Batch processing is therefore the more economical option when low latency is not important. xAI’s launch information states that diarization, timestamps and key terms are included without an additional charge.
The editorial speed and cost assessments for this model are favorable because it is specialized for transcription, offers a low per-hour price and supports real-time use. These are evaluations of its practical positioning, not scores published by xAI and not a substitute for application-specific testing.
Inputs, outputs and documented limits
The model’s input is audio, and its primary output is text. It can also return transcript metadata, including timestamps, confidence scores, channel details and speaker labels. It does not directly generate images, video, music or speech audio.
Grok Voice Transcribe 2.0 is not documented as a general multimodal conversational model. In particular, the supplied specifications do not provide a text context-window size, maximum output-token limit or separate reasoning-token limit. Those language-model limits are not applicable in the usual sense because this product is specialized for speech recognition and returns transcripts rather than open-ended generated responses.
The model also has no documented general-purpose tool or function-calling capability. Its API features should be understood as transcription controls and metadata options, not as tools that let the model browse the web, call external services or execute application functions.
Strengths and limitations
The main strength of Grok Voice Transcribe 2.0 is the combination of low usage pricing and features that matter in real production transcription. Automatic language detection reduces setup work for multilingual inputs. Speaker diarization and multichannel support are useful for calls and meetings, while timestamps and formatting controls make transcripts easier to index, search and display.
Its streaming mode is another important advantage for applications that need immediate text. Voice agents, live captions and monitoring systems can begin reacting to interim results rather than waiting for an entire recording to finish.
The model remains specialized. It transcribes speech but does not replace a reasoning model, text-generation model, coding assistant or speech-synthesis system. An application that needs to summarize a meeting, answer questions about a transcript or take an action will need additional application logic or another model after transcription. Similarly, the accuracy of the final result depends on recording quality, background noise, overlapping speakers, channel configuration, language and accents.
Streaming applications must also account for revisions to interim text. A user interface should mark provisional words as changeable and update them when completed transcript events arrive. Systems that require a stable record should store finalized results rather than treating every interim event as permanent.
When to choose Grok Voice Transcribe 2.0
Choose Grok Voice Transcribe 2.0 when the central requirement is converting audio into structured, usable text at a per-hour price. It is a good fit for:
- Customer-support and contact-center transcription.
- Live captions and accessibility features.
- Voice-agent input and turn-aware conversational interfaces.
- Meetings, interviews and multi-speaker conversations.
- Transcription of telephone or other multichannel recordings.
- Video and audio indexing with searchable timestamps.
- Multilingual voice commands and applications that need automatic language detection.
- Specialized vocabulary that can benefit from key-term biasing.
Batch mode is the better choice when recordings can be processed after they are created and minimizing cost is the priority. Streaming mode is more appropriate when users need visible or actionable transcription during the conversation and the additional $0.10 per hour is justified by lower latency.
When another type of option may be more appropriate
A general-purpose language model is more suitable when the main task is reasoning over text, writing code, summarizing a transcript or producing a conversational answer. Grok Voice Transcribe 2.0 can provide the transcript that feeds those tasks, but it is not documented as performing them itself.
A speech-generation or voice-conversation system is more appropriate when the application must produce spoken audio in response to the user. Grok Voice Transcribe 2.0 handles the listening and transcription side only; it does not synthesize speech.
Another transcription service may be preferable if an application requires a region, integration, language, retention policy or accuracy profile that xAI’s documented service does not provide. The supplied documentation identifies the us-east-1 region and more than 38 supported languages, but it does not specify a broader regional availability list, a universal accuracy guarantee or a custom context-length limit. Those factors should be checked against the requirements of the intended deployment.
Bottom line
Grok Voice Transcribe 2.0 is a focused xAI speech-to-text model rather than a general AI assistant. Its strongest practical case is affordable transcription for recorded or live audio, especially when an application benefits from multilingual detection, timestamps, speaker labels, multichannel input, key-term biasing or smart turn detection. Batch processing offers the lower price, while streaming provides the responsiveness needed for live interfaces. For reasoning, coding, summarization, external actions or speech output, it should be paired with other components rather than used as a standalone replacement.

