What is Grok Voice Transcribe 1.0?
Grok Voice Transcribe 1.0 is xAI’s original dedicated speech-to-text model. Its canonical model identifier is grok-voice-transcribe-1.0. The model analyzes spoken audio and produces a written transcript, making it intended for transcription workflows rather than open-ended text generation or conversation.
xAI exposes the model through its Speech to Text API in two main ways. REST requests are suited to recorded audio, while a WebSocket connection supports real-time or low-latency transcription as audio arrives. In practical terms, this lets an application process an existing meeting recording, an uploaded voice memo, an audio URL, or a live microphone stream.
The model can be used for dictation, meeting notes, voice assistants, accessibility features, customer-support recordings, and other applications in which speech must be converted into searchable or editable text.
Where it fits in xAI’s current lineup
Grok Voice Transcribe 1.0 is a specialized member of xAI’s audio model offering. It is not a general-purpose Grok language model with speech recognition added as a secondary feature; its primary job is recognizing and transcribing speech.
Its current status is important. xAI’s recent Speech to Text documentation identifies Grok Voice Transcribe 2.0 as the default model, while version 1.0 remains accessible when explicitly selected by its pinned identifier. xAI has announced that version 1.0 will be deprecated in the coming weeks, but no exact shutdown date is published in the supplied documentation.
That makes version 1.0 most relevant for existing integrations that need reproducible behavior or have already been tested against this model. New projects should treat migration planning as part of the implementation rather than assuming that the older model will remain the long-term default.
Core transcription capabilities
The model supports both batch and streaming use cases. Batch transcription sends an audio file or supported audio location to the REST service and receives transcript data after processing. Streaming transcription uses a WebSocket connection and can return interim results before the complete utterance or session has finished.
- Multilingual recognition: The service supports multiple languages listed in xAI’s current Speech to Text documentation.
- Interim results: Streaming applications can display provisional text while speech is still being processed.
- Key-term prompting: Applications can provide important names, product terms, technical vocabulary, or other words that may otherwise be difficult to recognize.
- Word-level timestamps: Timestamps are available where supported by the transcription response, which is useful for captions, search, and synchronized playback.
- Speaker-oriented transcription: API options include speaker diarization and multichannel transcription. Diarization separates speech by inferred speaker, while multichannel processing can use separate audio channels when a recording provides them.
- Formatting controls: The API supports formatting behavior for numbers, currencies, units, phone numbers, and similar spoken forms.
These features make the model more suitable for structured transcription pipelines than a basic audio-to-text endpoint that returns only an undifferentiated block of text.
Supported inputs and outputs
The input modality is audio. Supported containers and raw audio formats include WAV, MP3, WebM, OGG, M4A, PCM, mu-law, A-law, and Opus in supported request configurations. Exact encoding and request requirements depend on the API mode and format being submitted, so production clients should validate their audio pipeline against xAI’s API reference.
The output is text transcript data, potentially accompanied by metadata such as interim results, timestamps, or speaker information when the relevant options are enabled. Grok Voice Transcribe 1.0 does not natively generate speech, music, images, video, embeddings, or other non-text outputs.
| Capability | Grok Voice Transcribe 1.0 |
|---|---|
| Audio input | Yes |
| Text output | Yes, as transcript data |
| Real-time streaming | Yes, through WebSocket |
| Batch transcription | Yes, through REST |
| Speech synthesis | No |
| Image or video generation | No |
| Tool or function calling | Not supported as a model capability |
Pricing and performance trade-offs
xAI lists Speech to Text pricing by audio duration rather than by language-model input and output tokens. REST batch transcription is priced at $0.10 per hour of audio, while streaming transcription is priced at $0.20 per hour of audio. The supplied research states that these prices apply to both Grok Voice Transcribe 1.0 and Grok Voice Transcribe 2.0.
The difference between the two rates reflects the service mode: batch processing is generally appropriate when results can wait until an audio file has been processed, while streaming is intended for applications that need text during the conversation or recording. Streaming therefore costs more per audio hour but can reduce the delay between spoken words and visible transcript text.
The research data gives the model an editorial speed score of 8 and cost score of 9. These are catalog evaluations, not xAI-published benchmark results, and they should not be interpreted as a formal accuracy or latency guarantee. The more concrete provider-published distinction is the separate batch and streaming price structure.
Limits and unspecified capabilities
Grok Voice Transcribe 1.0 is specified primarily by audio duration, request format, and transcription behavior rather than by the context-window and output-token limits commonly published for text-generation models. No published token context length or maximum text-output token count is supplied for this model.
This does not mean that requests are unlimited. Audio formats, API request rules, service limits, connection behavior, and any duration restrictions documented by xAI still apply. It means that a token-based context figure is not an appropriate way to estimate the model’s supported recording length from the supplied information.
Reasoning and coding capabilities are not relevant model functions here. The model is designed to recognize speech, not to reason over instructions, write software, answer general questions, or transform a transcript into an analyzed report. A separate text-generation model or application-processing step is more appropriate when the workflow requires summarization, classification, extraction, or code generation after transcription.
Best use cases
Choose Grok Voice Transcribe 1.0 when the central requirement is dependable access to xAI’s original speech-to-text behavior and the integration benefits from explicit model pinning. Suitable applications include:
- Transcribing uploaded interviews, meetings, calls, lectures, and voice notes through REST.
- Building live captions or voice interfaces that need interim transcript updates.
- Processing multilingual audio supported by the Speech to Text service.
- Improving recognition of specialized names, products, organizations, or technical terms through key-term prompting.
- Creating searchable recordings with timestamps and, where appropriate, speaker separation.
- Handling recordings that use common formats such as WAV, MP3, WebM, OGG, M4A, PCM, mu-law, A-law, or Opus.
The model’s audio-duration pricing can also be easier to estimate than token-based pricing for teams that know the approximate number of recording hours they process each month.
When another option may be more appropriate
Grok Voice Transcribe 1.0 is not the best choice for every speech workflow. If an application needs the provider’s current default transcription model, the newer Grok Voice Transcribe 2.0 should be evaluated because xAI has positioned it as the successor and has announced the coming deprecation of version 1.0.
A general-purpose language model is more appropriate when the primary task is conversation, reasoning, document generation, code production, or analysis of transcript content. Similarly, a speech-synthesis model is required when the application must speak responses aloud; Grok Voice Transcribe 1.0 only converts audio into text.
For recorded audio where immediate feedback is unnecessary, REST batch transcription is the more economical mode at $0.10 per audio hour. For live captions, voice controls, or interactive assistants, WebSocket streaming is the relevant option despite its higher listed rate of $0.20 per audio hour.
Implementation guidance and status
Applications that intentionally use this model should specify grok-voice-transcribe-1.0 rather than relying on xAI’s service default. Pinning the identifier helps prevent an automatic switch to the newer default while the application is being tested or while transcript behavior is being compared across versions.
At the same time, pinning version 1.0 should not be treated as a substitute for migration planning. Because xAI has announced deprecation in the coming weeks without publishing an exact shutdown date, teams should test their audio formats, prompting terms, timestamp handling, speaker options, and downstream transcript processing with Grok Voice Transcribe 2.0.
In summary, Grok Voice Transcribe 1.0 remains a focused and relatively inexpensive transcription endpoint with both batch and real-time access. Its clearest strengths are audio-format coverage, streaming support, multilingual transcription, and transcription-oriented controls. Its decisive limitation is lifecycle status: it is an older model that remains accessible today but is scheduled to be replaced.

