Grok TTS

grok-tts

by xAI · Current

Grok TTS is xAI’s specialist text-to-speech model for expressive spoken audio. It supports REST, WebSocket streaming, batch generation, built-in and custom voices, multiple languages, expressive speech tags, and MP3, WAV, PCM, μ-law, and A-law output. xAI lists pricing at $15 per one million input characters, with a 60,000-character REST request limit.

Speech Reasoning Coding
Grok TTS is xAI’s standalone text-to-speech model for applications that need generated speech rather than text responses. It converts supplied text into spoken audio using expressive voices and can deliver results through standard REST requests, WebSocket streaming, or batch workflows. The model supports multiple languages, built-in and custom voices, several audio formats, and expressive cues such as pauses, laughter, sighs, and breaths. xAI lists pricing at $15 per one million input characters.
Outputs

What grok-tts can produce

Speech
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Batch API Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Grok TTS
Model type Other
Release date 2026-04-17
Status Current
Knowledge cutoff notes

xAI does not publish a conventional knowledge cutoff for this speech-generation model. Its behavior depends on the text supplied at generation time rather than on a documented factual knowledge boundary.

Model notes

The canonical xAI model identifier is grok-tts. The model accepts text and returns spoken audio. xAI documents REST, WebSocket streaming, and batch workflows, with built-in and custom voices, multilingual support, and MP3, WAV, PCM, μ-law, and A-law formats. The API documentation states a maximum of 60,000 input characters per REST request. Pricing is character-based rather than token-based. The documented service region is us-east-1, with 50 requests per second and 100 concurrent sessions per team listed for the Text to Speech service. Editorial scores are comparative estimates for a specialist text-to-speech model and do not represent vendor benchmarks.

Cost

Model pricing

Input $15.00 per 1 million input characters
Output Included in character-based pricing; no separate output charge documented
Model guide

Grok TTS: xAI’s Dedicated Model for Expressive Speech Generation

Grok TTS is xAI’s dedicated text-to-speech model for turning written text into expressive spoken audio. It supports REST generation, WebSocket streaming, batch workflows, multiple languages, built-in and custom voices, expressive speech tags, and several audio formats. xAI lists the model at $15 per one million input characters, with a documented limit of 60,000 input characters per REST request.

What is Grok TTS?

Grok TTS is xAI’s dedicated text-to-speech model, identified in the developer documentation by the model name grok-tts. Its job is narrowly defined: it accepts written text and generates spoken audio. Unlike a general-purpose language model, it is not intended to answer questions, write software, search the web, create images, or perform broad reasoning tasks.

The model is part of xAI’s standalone speech API offering. It is designed for developers building applications that need a generated voice, including voice agents, narration systems, podcasts, audiobooks, accessibility features, customer-support experiences, educational material, and interactive audio applications.

For a beginner, the simplest way to understand Grok TTS is as a conversion layer between text and sound. An application first creates or receives text, then sends that text to Grok TTS with voice and output-format settings. The service returns audio that can be played, stored, or passed to another part of an application.

How Grok TTS fits into xAI’s catalog

Grok TTS is a specialist audio-generation model rather than a general Grok chat model. xAI documents it alongside a standalone Speech to Text API announced on April 17, 2026. The two capabilities address opposite parts of a voice workflow: speech-to-text transcribes audio, while Grok TTS turns text back into speech.

This positioning matters when selecting a model. Grok TTS can provide the spoken output for a voice assistant, but it does not independently provide the conversational reasoning or transcription needed for a complete voice assistant. A production system may therefore combine it with a separate language model and, where necessary, a speech-recognition service.

Inputs, outputs, and supported formats

Grok TTS uses text input and spoken audio output. The supplied model information identifies text as its input modality and audio as its output modality. It does not document image, video, or audio input for this model.

xAI documents several output formats:

  • MP3
  • WAV
  • PCM
  • μ-law
  • A-law

The choice of format depends on how the audio will be used. A compressed format such as MP3 can be convenient for distribution, while uncompressed formats such as WAV or PCM may be more appropriate when an application needs to control playback or audio processing. The supplied documentation does not specify a single default format or provide detailed sampling and bitrate tables, so those settings should be checked against the current API reference during implementation.

The model supports expressive built-in voices and, where available, custom voices. It also accepts inline expressive speech tags for cues such as pauses, laughter, sighs, breaths, and other delivery effects. These controls are useful when plain pronunciation is not enough—for example, when producing character dialogue, narration, or a support response that should sound less mechanically read.

REST, streaming, and batch generation

xAI documents three main ways to use Grok TTS. REST generation is suitable for a conventional request-and-response workflow: an application submits text and receives generated audio. The REST API reference states a maximum of 60,000 input characters per request.

For interactive applications, Grok TTS also supports bidirectional WebSocket streaming. Streaming allows audio chunks to be delivered while the text is being processed, so playback can begin before the entire response has finished. This can reduce the perceived waiting time in a voice interface, although the supplied research does not provide a measured latency benchmark.

Batch workflows are better suited to pre-generating large collections of audio. Examples include preparing audiobook chapters, producing a library of educational clips, or rendering many podcast and video-narration segments before publication. Batch processing can separate audio production from the user-facing request path, but the supplied information does not specify batch pricing or a separate batch limit.

The Text to Speech service documentation lists a limit of 50 requests per second and up to 100 concurrent sessions per team. These are service-level limits rather than a guarantee that every application will receive the same real-world throughput. Systems with high traffic should design for rate limiting, retries, and controlled concurrency.

Grok TTS pricing

xAI lists Grok TTS at $15 per one million input characters. The model uses character-based pricing rather than token-based pricing, and the supplied documentation does not identify a separate output-audio charge. This makes the amount of text sent to the service the central usage variable.

Character billing is easy to estimate for fixed scripts, but the cost can change with generated content volume. Applications that repeatedly regenerate the same text, stream long conversations, or produce large amounts of narration should measure their character usage. The 60,000-character REST request limit also means that very long material may need to be divided into multiple requests.

The documented service region is us-east-1. Organizations with regional processing or data-residency requirements should verify whether that deployment arrangement meets their requirements before adopting the service.

Capabilities and limitations

Grok TTS’s main strength is specialization. It is built for expressive speech rather than trying to make one model handle unrelated text, image, and audio tasks. Built-in voices, custom voices, multilingual support, expressive tags, multiple output formats, streaming, and batch processing cover a broad range of speech-production workflows.

Its main limitation is the same specialization. Grok TTS does not replace a reasoning model, a chatbot, a transcription system, an image model, or an embedding service. It receives text that has already been prepared by an application or another model. If the application needs to understand a user’s spoken request, it needs a separate speech-to-text component; if it needs to decide what to say, it needs a language or reasoning component.

xAI does not publish a conventional context-window size or maximum output-token limit for Grok TTS. Those concepts are primarily associated with text-generation models. The documented practical input limit is 60,000 characters per REST request. Output duration and file size will depend on the amount of text, language, voice, speaking rate, and audio format, but the supplied research does not provide a fixed maximum duration.

The model does not provide documented tool or function-calling support. It also does not return text as its primary output. Any application requiring structured data, external actions, or multi-step reasoning should perform those operations outside the TTS request and then send the final text to Grok TTS.

Reasoning, coding, speed, and cost

Grok TTS is not a reasoning model. It can render text that contains an explanation, dialogue, or program-related content, but that does not mean it understands or evaluates the content in the way a general-purpose language model does. It should not be selected for planning, analysis, factual question answering, or autonomous decision-making.

It is similarly not a coding model. It can read code-shaped text aloud for narration or accessibility, but it is not intended to generate, debug, review, or execute software.

The supplied editorial dataset assigns Grok TTS a speed score of 8 and a cost score of 7, while assigning reasoning and coding scores of 1. These are comparative editorial estimates for a specialist speech model, not provider-published benchmark results. They express the practical expectation that a focused TTS service can be a good fit for speech generation and can offer a relatively favorable speed and cost profile, but they should not be treated as measured latency, quality, or price benchmarks against every competing service.

Best use cases for Grok TTS

Grok TTS is a strong candidate when an application already has text and needs natural, expressive audio. Suitable use cases include:

  • Voice agents: deliver spoken replies after a separate system generates the response.
  • Customer support: read scripted or dynamically generated answers aloud.
  • Narration: produce audio for videos, presentations, courses, and explainers.
  • Podcasts and audiobooks: generate longer-form spoken content in batch workflows.
  • Accessibility: provide an audio representation of written content.
  • Interactive characters: use voices and expressive tags for games, entertainment, or conversational experiences.
  • Low-latency playback: use WebSocket streaming when audio should begin before the complete response is ready.

Custom voices can be useful when an application needs a consistent identity, although the supplied research does not define the availability requirements, approval process, or limitations for voice cloning.

When to choose Grok TTS

Choose Grok TTS when the primary requirement is expressive text-to-speech from xAI, especially when WebSocket streaming, batch generation, multiple audio formats, or expressive delivery controls are important. Its character-based pricing can also be straightforward for teams that can estimate their text volume.

Another speech service may be more appropriate if an application requires a documented context or duration limit that differs from Grok TTS’s published specifications, a different deployment region, specialized voice or language coverage, or independently verified audio-quality benchmarks. A general-purpose language model is more appropriate when the main task is reasoning or text generation. A speech-recognition model is required when the application must transcribe user audio. For a complete voice assistant, Grok TTS is best viewed as the speech-output component rather than the entire conversational stack.

Bottom line

Grok TTS is xAI’s focused text-to-speech model for producing expressive spoken audio from text. Its documented offering combines built-in and custom voices, multilingual generation, expressive speech tags, REST requests, WebSocket streaming, batch workflows, and MP3, WAV, PCM, μ-law, and A-law output. The key published constraints are a 60,000-character REST request limit, 50 requests per second, and 100 concurrent sessions per team. At $15 per one million input characters, it is primarily a practical speech-generation component for applications that already have a text-generation or dialogue layer.


Answers to Frequently Asked Questions

What is Grok TTS used for?
Grok TTS is xAI’s dedicated text-to-speech model, identified as grok-tts. It converts written text into expressive spoken audio for voice agents, narration, podcasts, audiobooks, accessibility features, customer support, education, and interactive audio applications.
What audio formats does Grok TTS support?
Grok TTS supports MP3, WAV, PCM, ?-law, and A-law output formats. The best choice depends on whether the application prioritizes distribution efficiency, playback control, or further audio processing.
Does Grok TTS support streaming and long text input?
Yes. Grok TTS supports bidirectional WebSocket streaming, which can deliver audio chunks before the full response is complete, as well as REST and batch generation. REST requests support up to 60,000 input characters. The documented service limits are 50 requests per second and up to 100 concurrent sessions per team.
How much does Grok TTS cost?
xAI lists Grok TTS at $15 per one million input characters. Pricing is character-based, and the supplied documentation does not identify a separate charge for output audio.
Is Grok TTS a complete voice assistant or reasoning model?
No. Grok TTS only generates spoken audio from text. A complete voice assistant generally requires a separate language or reasoning model to decide what to say and, when users speak to the system, a speech-to-text service to transcribe their audio.


Sources 5
Provider

About xAI