Grok Voice Transcribe

Grok Voice Transcribe 1.0

by xAI · Currently accessible; original speech-to-text model; scheduled for deprecation in the coming weeks

xAI’s original dedicated speech-to-text model for batch and real-time audio transcription. It supports multiple audio formats, multilingual recognition, interim streaming results, key-term prompting, timestamps, speaker diarization, multichannel transcription, and formatting controls. Batch transcription costs $0.10 per audio hour and streaming costs $0.20 per hour. The model remains accessible through its pinned identifier but is scheduled for deprecation as Grok Voice Transcribe 2.0 becomes the default.

Text
Grok Voice Transcribe 1.0 is a specialized speech-recognition model from xAI rather than a general conversational model. It accepts audio and returns text transcripts through the Speech to Text API, either from uploaded files or URLs or through a real-time WebSocket connection. Its main practical advantage is focused transcription at usage-based audio pricing, although developers starting a new integration should account for the announced deprecation of version 1.0 and evaluate Grok Voice Transcribe 2.0 instead.
Outputs

What Grok Voice Transcribe 1.0 can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming Batch API
Model profile

Performance characteristics

8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Grok Voice Transcribe
Model type Speech Recognition
Release date April 15, 2026
Status Currently accessible; original speech-to-text model; scheduled for deprecation in the coming weeks
Knowledge cutoff notes

No official knowledge cutoff is published. As a dedicated speech-to-text model, its primary behavior is audio recognition rather than knowledge-based text generation.

Model notes

The canonical model identifier is grok-voice-transcribe-1.0. It supports REST file or URL transcription and WebSocket streaming. Current documentation lists support for multiple audio formats, multilingual transcription, interim streaming results, key-term prompting, word-level timestamps, speaker diarization, multichannel transcription, and smart turn controls. Grok Voice Transcribe 2.0 is now the default model and xAI has announced that version 1.0 will be deprecated in the coming weeks, but no exact shutdown date has been published. Pricing is charged by audio duration rather than input and output tokens. The model's knowledge cutoff is not published and is not materially applicable to a dedicated transcription model.

Cost

Model pricing

Input $0.10 per hour of audio for REST batch transcription; $0.20 per hour of audio for streaming
Output Included in the transcription service price; output is text transcript data
Model guide

Grok Voice Transcribe 1.0: xAI’s Original Speech-to-Text Model

Grok Voice Transcribe 1.0 is xAI’s dedicated audio-to-text model for converting recorded or live speech into written transcripts. It supports REST batch transcription and WebSocket streaming, multilingual recognition, interim results, key-term prompting, timestamps, speaker-oriented features, and multiple audio formats. The model remains selectable through its pinned identifier, but xAI now presents Grok Voice Transcribe 2.0 as the default and has announced that version 1.0 will be deprecated in the coming weeks.

What is Grok Voice Transcribe 1.0?

Grok Voice Transcribe 1.0 is xAI’s original dedicated speech-to-text model. Its canonical model identifier is grok-voice-transcribe-1.0. The model analyzes spoken audio and produces a written transcript, making it intended for transcription workflows rather than open-ended text generation or conversation.

xAI exposes the model through its Speech to Text API in two main ways. REST requests are suited to recorded audio, while a WebSocket connection supports real-time or low-latency transcription as audio arrives. In practical terms, this lets an application process an existing meeting recording, an uploaded voice memo, an audio URL, or a live microphone stream.

The model can be used for dictation, meeting notes, voice assistants, accessibility features, customer-support recordings, and other applications in which speech must be converted into searchable or editable text.

Where it fits in xAI’s current lineup

Grok Voice Transcribe 1.0 is a specialized member of xAI’s audio model offering. It is not a general-purpose Grok language model with speech recognition added as a secondary feature; its primary job is recognizing and transcribing speech.

Its current status is important. xAI’s recent Speech to Text documentation identifies Grok Voice Transcribe 2.0 as the default model, while version 1.0 remains accessible when explicitly selected by its pinned identifier. xAI has announced that version 1.0 will be deprecated in the coming weeks, but no exact shutdown date is published in the supplied documentation.

That makes version 1.0 most relevant for existing integrations that need reproducible behavior or have already been tested against this model. New projects should treat migration planning as part of the implementation rather than assuming that the older model will remain the long-term default.

Core transcription capabilities

The model supports both batch and streaming use cases. Batch transcription sends an audio file or supported audio location to the REST service and receives transcript data after processing. Streaming transcription uses a WebSocket connection and can return interim results before the complete utterance or session has finished.

  • Multilingual recognition: The service supports multiple languages listed in xAI’s current Speech to Text documentation.
  • Interim results: Streaming applications can display provisional text while speech is still being processed.
  • Key-term prompting: Applications can provide important names, product terms, technical vocabulary, or other words that may otherwise be difficult to recognize.
  • Word-level timestamps: Timestamps are available where supported by the transcription response, which is useful for captions, search, and synchronized playback.
  • Speaker-oriented transcription: API options include speaker diarization and multichannel transcription. Diarization separates speech by inferred speaker, while multichannel processing can use separate audio channels when a recording provides them.
  • Formatting controls: The API supports formatting behavior for numbers, currencies, units, phone numbers, and similar spoken forms.

These features make the model more suitable for structured transcription pipelines than a basic audio-to-text endpoint that returns only an undifferentiated block of text.

Supported inputs and outputs

The input modality is audio. Supported containers and raw audio formats include WAV, MP3, WebM, OGG, M4A, PCM, mu-law, A-law, and Opus in supported request configurations. Exact encoding and request requirements depend on the API mode and format being submitted, so production clients should validate their audio pipeline against xAI’s API reference.

The output is text transcript data, potentially accompanied by metadata such as interim results, timestamps, or speaker information when the relevant options are enabled. Grok Voice Transcribe 1.0 does not natively generate speech, music, images, video, embeddings, or other non-text outputs.

CapabilityGrok Voice Transcribe 1.0
Audio inputYes
Text outputYes, as transcript data
Real-time streamingYes, through WebSocket
Batch transcriptionYes, through REST
Speech synthesisNo
Image or video generationNo
Tool or function callingNot supported as a model capability

Pricing and performance trade-offs

xAI lists Speech to Text pricing by audio duration rather than by language-model input and output tokens. REST batch transcription is priced at $0.10 per hour of audio, while streaming transcription is priced at $0.20 per hour of audio. The supplied research states that these prices apply to both Grok Voice Transcribe 1.0 and Grok Voice Transcribe 2.0.

The difference between the two rates reflects the service mode: batch processing is generally appropriate when results can wait until an audio file has been processed, while streaming is intended for applications that need text during the conversation or recording. Streaming therefore costs more per audio hour but can reduce the delay between spoken words and visible transcript text.

The research data gives the model an editorial speed score of 8 and cost score of 9. These are catalog evaluations, not xAI-published benchmark results, and they should not be interpreted as a formal accuracy or latency guarantee. The more concrete provider-published distinction is the separate batch and streaming price structure.

Limits and unspecified capabilities

Grok Voice Transcribe 1.0 is specified primarily by audio duration, request format, and transcription behavior rather than by the context-window and output-token limits commonly published for text-generation models. No published token context length or maximum text-output token count is supplied for this model.

This does not mean that requests are unlimited. Audio formats, API request rules, service limits, connection behavior, and any duration restrictions documented by xAI still apply. It means that a token-based context figure is not an appropriate way to estimate the model’s supported recording length from the supplied information.

Reasoning and coding capabilities are not relevant model functions here. The model is designed to recognize speech, not to reason over instructions, write software, answer general questions, or transform a transcript into an analyzed report. A separate text-generation model or application-processing step is more appropriate when the workflow requires summarization, classification, extraction, or code generation after transcription.

Best use cases

Choose Grok Voice Transcribe 1.0 when the central requirement is dependable access to xAI’s original speech-to-text behavior and the integration benefits from explicit model pinning. Suitable applications include:

  • Transcribing uploaded interviews, meetings, calls, lectures, and voice notes through REST.
  • Building live captions or voice interfaces that need interim transcript updates.
  • Processing multilingual audio supported by the Speech to Text service.
  • Improving recognition of specialized names, products, organizations, or technical terms through key-term prompting.
  • Creating searchable recordings with timestamps and, where appropriate, speaker separation.
  • Handling recordings that use common formats such as WAV, MP3, WebM, OGG, M4A, PCM, mu-law, A-law, or Opus.

The model’s audio-duration pricing can also be easier to estimate than token-based pricing for teams that know the approximate number of recording hours they process each month.

When another option may be more appropriate

Grok Voice Transcribe 1.0 is not the best choice for every speech workflow. If an application needs the provider’s current default transcription model, the newer Grok Voice Transcribe 2.0 should be evaluated because xAI has positioned it as the successor and has announced the coming deprecation of version 1.0.

A general-purpose language model is more appropriate when the primary task is conversation, reasoning, document generation, code production, or analysis of transcript content. Similarly, a speech-synthesis model is required when the application must speak responses aloud; Grok Voice Transcribe 1.0 only converts audio into text.

For recorded audio where immediate feedback is unnecessary, REST batch transcription is the more economical mode at $0.10 per audio hour. For live captions, voice controls, or interactive assistants, WebSocket streaming is the relevant option despite its higher listed rate of $0.20 per audio hour.

Implementation guidance and status

Applications that intentionally use this model should specify grok-voice-transcribe-1.0 rather than relying on xAI’s service default. Pinning the identifier helps prevent an automatic switch to the newer default while the application is being tested or while transcript behavior is being compared across versions.

At the same time, pinning version 1.0 should not be treated as a substitute for migration planning. Because xAI has announced deprecation in the coming weeks without publishing an exact shutdown date, teams should test their audio formats, prompting terms, timestamp handling, speaker options, and downstream transcript processing with Grok Voice Transcribe 2.0.

In summary, Grok Voice Transcribe 1.0 remains a focused and relatively inexpensive transcription endpoint with both batch and real-time access. Its clearest strengths are audio-format coverage, streaming support, multilingual transcription, and transcription-oriented controls. Its decisive limitation is lifecycle status: it is an older model that remains accessible today but is scheduled to be replaced.


Answers to Frequently Asked Questions

Is Grok Voice Transcribe 1.0 still available, and should new projects use it?
Grok Voice Transcribe 1.0 remains accessible when explicitly selected by its pinned model identifier, but xAI has announced that it will be deprecated in the coming weeks. New projects should evaluate Grok Voice Transcribe 2.0, which xAI identifies as the current default, while existing users should plan and test a migration.
How much does Grok Voice Transcribe 1.0 cost?
xAI lists batch REST transcription at $0.10 per hour of audio and streaming WebSocket transcription at $0.20 per hour. The streaming rate is higher because it supports real-time processing and interim transcript updates.
What audio formats and transcription features does Grok Voice Transcribe 1.0 support?
Supported inputs include formats such as WAV, MP3, WebM, OGG, M4A, PCM, mu-law, A-law, and Opus, depending on the request configuration. Features include multilingual recognition, key-term prompting, word-level timestamps, speaker diarization, multichannel transcription, interim results, and formatting controls.
What is Grok Voice Transcribe 1.0?
Grok Voice Transcribe 1.0 is xAI’s dedicated speech-to-text model, identified as grok-voice-transcribe-1.0. It converts spoken audio into written transcripts for use cases such as meetings, voice notes, live captions, accessibility tools, and voice assistants.
How can Grok Voice Transcribe 1.0 be accessed?
The model is available through xAI’s Speech to Text API. REST requests support batch transcription of recorded audio, while WebSocket connections support real-time or low-latency transcription with interim results.


Sources 5
Provider

About xAI