Grok Voice Transcribe

Grok Voice Transcribe 2.0

by xAI · Current

xAI’s Grok Voice Transcribe 2.0 converts recorded and streaming audio into text. It supports more than 38 languages, timestamps, confidence scores, speaker diarization, up to eight audio channels, key-term biasing, formatting controls and smart turn detection. Batch transcription costs $0.10 per hour, while streaming costs $0.20 per hour.

Text Reasoning Coding
Grok Voice Transcribe 2.0 is xAI’s specialized speech-to-text model for turning recorded or live audio into written transcripts. It is designed for multilingual conversations, telephone audio, noisy environments and voice-agent applications, and is available through both batch REST requests and real-time WebSocket streaming. Pricing is based on audio duration: $0.10 per hour for batch transcription and $0.20 per hour for streaming.
Outputs

What Grok Voice Transcribe 2.0 can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming Batch API
Model profile

Performance characteristics

0/10 Reasoning
0/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Grok Voice Transcribe
Model type Other
Context window tokens
Maximum output tokens
Release date September 17, 2026
Status Current
Knowledge cutoff notes

xAI does not publish a separate knowledge cutoff for this speech-recognition model. Its behavior is primarily determined by audio-transcription training and API configuration rather than a stated world-knowledge cutoff.

Model notes

Canonical API identifier: grok-voice-transcribe-2.0. The model supports batch REST transcription and real-time WebSocket streaming. xAI documents more than 38 languages with automatic language detection, word-level timestamps, confidence scores, speaker diarization, up to eight-channel multichannel transcription, key-term biasing, text formatting, filler-word removal, and smart turn detection. The model was announced as the successor to Grok Voice Transcribe 1.0. xAI's release notes state that Grok Voice Transcribe 1.0 reached end of life on October 2, 2026 and that requests to the old identifier are routed to Grok Voice Transcribe 2.0 at the same price. Editorial speed and cost scores reflect its specialized real-time positioning and low per-hour pricing; reasoning and coding scores are not applicable.

Cost

Model pricing

Input $0.10 per hour of audio for batch REST transcription; $0.20 per hour of audio for streaming transcription
Model guide

Grok Voice Transcribe 2.0: Low-Cost Batch and Real-Time Speech Transcription

Grok Voice Transcribe 2.0 is xAI’s current speech-recognition model for converting recorded or streaming audio into text. It supports batch REST transcription and real-time WebSocket streaming, with automatic language detection across more than 38 languages, word-level timestamps, confidence scores, speaker diarization, multichannel audio, key-term biasing, formatting controls, filler-word removal and smart turn detection. Batch transcription costs $0.10 per hour of audio, while streaming costs $0.20 per hour.

What is Grok Voice Transcribe 2.0?

Grok Voice Transcribe 2.0 is xAI’s current model for automatic speech recognition. Its primary job is to listen to audio and produce a written transcript, rather than to hold a conversation, generate general text or perform broad reasoning. The model accepts audio input and returns text together with transcript metadata such as timestamps, confidence values, channel information and speaker labels.

xAI provides the model through its Speech to Text API. It can process uploaded or referenced audio in batch mode, or transcribe an active audio stream over a WebSocket connection. The exact API model identifier is grok-voice-transcribe-2.0.

Within xAI’s current speech-to-text lineup, Grok Voice Transcribe 2.0 is the default model. It replaces Grok Voice Transcribe 1.0, which xAI announced for deprecation. This positioning makes version 2.0 the relevant choice for new integrations that need xAI’s transcription service.

Core capabilities for practical transcription

The model is intended for audio that is more complicated than a clean, single-speaker recording. xAI describes support for telephone calls, conversations with overlapping or competing voices, accents, noisy surroundings and spoken information such as phone numbers and email addresses.

  • Automatic language detection: xAI documents support for more than 38 languages without requiring the application to identify the language in advance.
  • Word-level timestamps: Individual words can be associated with their position in the audio, which is useful for captions, search, editing and playback synchronization.
  • Confidence scores: Transcript results can include an indication of how certain the system is about recognized content.
  • Speaker diarization: The service can label speech by speaker, helping separate participants in interviews, meetings and calls.
  • Multichannel transcription: Up to eight audio channels are supported, which can be useful when different microphones or call participants are recorded separately.
  • Key-term biasing: Applications can provide important vocabulary to improve recognition of domain-specific names, products or terminology.
  • Text formatting: The API can format numbers, dates, currencies, phone numbers and email addresses in more readable forms.
  • Filler-word removal: Optional removal of words such as conversational fillers can produce cleaner transcripts.
  • Smart turn detection: The model supports turn detection for voice-agent workflows, where the application needs to determine when a speaker has finished speaking.

These features distinguish Grok Voice Transcribe 2.0 from a minimal speech-recognition endpoint that returns only an unstructured block of text. Developers can use the additional metadata to build searchable recordings, synchronized captions, speaker-attributed notes and responsive voice interfaces.

Batch and streaming modes

Grok Voice Transcribe 2.0 supports two operating patterns. Batch transcription is designed for audio that already exists, such as a meeting recording, podcast, uploaded video soundtrack or archived support call. The application submits the audio through the REST API and receives a completed transcription.

Streaming transcription is designed for live or near-live use. Audio is sent through a WebSocket connection while it is being recorded, and the service returns interim transcript events followed by completed transcript results. Interim text allows an application to display captions or begin processing speech before the speaker has finished. However, earlier interim text may be corrected as additional audio arrives, so applications should treat provisional results differently from finalized segments.

xAI documents support for common audio formats including WAV, MP3, WebM, OGG and M4A. The service is available in the us-east-1 region. Documented service limits include 10 requests per second for REST use and 10 requests per second for streaming use, with up to 100 concurrent streaming sessions per team.

Pricing and cost trade-offs

xAI prices Grok Voice Transcribe 2.0 by the amount of audio processed rather than by the number of generated text tokens. Batch REST transcription costs $0.10 per hour of audio. Real-time streaming transcription costs $0.20 per hour of audio.

ModePriceTypical purpose
Batch REST transcription$0.10 per hour of audioRecorded meetings, calls, interviews and media archives
Real-time WebSocket transcription$0.20 per hour of audioLive captions, voice agents and interactive applications

The streaming rate is twice the batch rate, but it supports a different operational requirement: receiving transcript events while audio is being produced. Batch processing is therefore the more economical option when low latency is not important. xAI’s launch information states that diarization, timestamps and key terms are included without an additional charge.

The editorial speed and cost assessments for this model are favorable because it is specialized for transcription, offers a low per-hour price and supports real-time use. These are evaluations of its practical positioning, not scores published by xAI and not a substitute for application-specific testing.

Inputs, outputs and documented limits

The model’s input is audio, and its primary output is text. It can also return transcript metadata, including timestamps, confidence scores, channel details and speaker labels. It does not directly generate images, video, music or speech audio.

Grok Voice Transcribe 2.0 is not documented as a general multimodal conversational model. In particular, the supplied specifications do not provide a text context-window size, maximum output-token limit or separate reasoning-token limit. Those language-model limits are not applicable in the usual sense because this product is specialized for speech recognition and returns transcripts rather than open-ended generated responses.

The model also has no documented general-purpose tool or function-calling capability. Its API features should be understood as transcription controls and metadata options, not as tools that let the model browse the web, call external services or execute application functions.

Strengths and limitations

The main strength of Grok Voice Transcribe 2.0 is the combination of low usage pricing and features that matter in real production transcription. Automatic language detection reduces setup work for multilingual inputs. Speaker diarization and multichannel support are useful for calls and meetings, while timestamps and formatting controls make transcripts easier to index, search and display.

Its streaming mode is another important advantage for applications that need immediate text. Voice agents, live captions and monitoring systems can begin reacting to interim results rather than waiting for an entire recording to finish.

The model remains specialized. It transcribes speech but does not replace a reasoning model, text-generation model, coding assistant or speech-synthesis system. An application that needs to summarize a meeting, answer questions about a transcript or take an action will need additional application logic or another model after transcription. Similarly, the accuracy of the final result depends on recording quality, background noise, overlapping speakers, channel configuration, language and accents.

Streaming applications must also account for revisions to interim text. A user interface should mark provisional words as changeable and update them when completed transcript events arrive. Systems that require a stable record should store finalized results rather than treating every interim event as permanent.

When to choose Grok Voice Transcribe 2.0

Choose Grok Voice Transcribe 2.0 when the central requirement is converting audio into structured, usable text at a per-hour price. It is a good fit for:

  • Customer-support and contact-center transcription.
  • Live captions and accessibility features.
  • Voice-agent input and turn-aware conversational interfaces.
  • Meetings, interviews and multi-speaker conversations.
  • Transcription of telephone or other multichannel recordings.
  • Video and audio indexing with searchable timestamps.
  • Multilingual voice commands and applications that need automatic language detection.
  • Specialized vocabulary that can benefit from key-term biasing.

Batch mode is the better choice when recordings can be processed after they are created and minimizing cost is the priority. Streaming mode is more appropriate when users need visible or actionable transcription during the conversation and the additional $0.10 per hour is justified by lower latency.

When another type of option may be more appropriate

A general-purpose language model is more suitable when the main task is reasoning over text, writing code, summarizing a transcript or producing a conversational answer. Grok Voice Transcribe 2.0 can provide the transcript that feeds those tasks, but it is not documented as performing them itself.

A speech-generation or voice-conversation system is more appropriate when the application must produce spoken audio in response to the user. Grok Voice Transcribe 2.0 handles the listening and transcription side only; it does not synthesize speech.

Another transcription service may be preferable if an application requires a region, integration, language, retention policy or accuracy profile that xAI’s documented service does not provide. The supplied documentation identifies the us-east-1 region and more than 38 supported languages, but it does not specify a broader regional availability list, a universal accuracy guarantee or a custom context-length limit. Those factors should be checked against the requirements of the intended deployment.

Bottom line

Grok Voice Transcribe 2.0 is a focused xAI speech-to-text model rather than a general AI assistant. Its strongest practical case is affordable transcription for recorded or live audio, especially when an application benefits from multilingual detection, timestamps, speaker labels, multichannel input, key-term biasing or smart turn detection. Batch processing offers the lower price, while streaming provides the responsiveness needed for live interfaces. For reasoning, coding, summarization, external actions or speech output, it should be paired with other components rather than used as a standalone replacement.


Answers to Frequently Asked Questions

Is Grok Voice Transcribe 2.0 a general-purpose conversational AI model?
No. Grok Voice Transcribe 2.0 is specialized for speech recognition and does not provide general reasoning, summarization, coding, tool calling or speech synthesis. Applications that need those capabilities must connect the transcript to additional logic or other models.
What features does Grok Voice Transcribe 2.0 support?
The model supports automatic language detection for more than 38 languages, word-level timestamps, confidence scores, speaker diarization, up to eight audio channels, key-term biasing, text formatting, filler-word removal and smart turn detection.
What is the difference between batch and streaming transcription?
Batch transcription processes existing recordings and returns a completed transcript, making it the lower-cost option. Streaming transcription sends audio over WebSocket and returns interim and finalized results during recording, making it suitable for live captions, voice agents and interactive applications.
What is Grok Voice Transcribe 2.0 used for?
Grok Voice Transcribe 2.0 is xAI’s speech-to-text model for converting recorded or live audio into written transcripts. It supports metadata such as timestamps, confidence scores, speaker labels and channel information.
How much does Grok Voice Transcribe 2.0 cost?
Batch REST transcription costs $0.10 per hour of audio, while real-time WebSocket transcription costs $0.20 per hour. xAI states that diarization, timestamps and key terms are included without an additional charge.


Sources 4
Provider

About xAI