GPT-Realtime

GPT-Realtime-Translate

by OpenAI · Current

OpenAI's GPT-Realtime-Translate is a dedicated streaming speech-to-speech translation model for live multilingual audio. It returns translated speech and transcript deltas, supports WebRTC and WebSocket integrations, and costs $0.034 per minute of realtime audio.

Text Speech Reasoning Coding
GPT-Realtime-Translate is designed for applications that need spoken language translation with minimal delay. It accepts incoming audio through OpenAI's dedicated Realtime translation API and produces translated audio plus transcript text as the conversation progresses. OpenAI lists support for more than 70 input languages and 13 output languages, with pricing of $0.034 per minute of realtime audio.
Outputs

What GPT-Realtime-Translate can produce

Text Speech
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

3/10 Reasoning
1/10 Coding
9/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-Realtime
Model type Other
Context window 16K tokens
Maximum output 2K tokens
Knowledge cutoff 2024-09-30
Release date 2026-05-07
Status Current
Knowledge cutoff notes

The official model documentation lists September 30, 2024 as the model's knowledge cutoff. This cutoff is separate from the model's ability to process live audio supplied during a translation session.

Model notes

GPT-Realtime-Translate is a specialized interpreter model rather than a conversational assistant. It uses the dedicated /v1/realtime/translations endpoint and continuously translates incoming audio without a normal response.create lifecycle. OpenAI states that it supports more than 70 input languages and 13 output languages. The model returns translated audio and transcript deltas while source audio is still arriving. WebRTC is recommended for browser audio, while WebSockets support server-side raw audio integrations. Editorial scores reflect its specialization: very high speed and cost efficiency for live translation, but low coding and general reasoning suitability. Pricing is based on realtime audio duration, not text-token input and output rates.

Cost

Model pricing

Input $0.034 per minute of realtime audio
Output $0.034 per minute of realtime audio
Model guide

GPT-Realtime-Translate: Live Speech Translation, Pricing and API Support

GPT-Realtime-Translate is OpenAI's dedicated realtime speech-to-speech translation model. It continuously accepts spoken audio, returns translated speech while the source speaker is still talking, and provides transcript deltas for captions and downstream applications. It is designed for low-latency multilingual interpretation rather than general voice-agent conversations.

What is GPT-Realtime-Translate?

GPT-Realtime-Translate is OpenAI's specialized streaming speech-to-speech translation model. Instead of waiting for a complete recording, transcribing it, translating the text, and then synthesizing speech, an application can stream source audio into a translation session and receive translated audio while the speaker continues talking.

This makes the model an interpreter rather than a general-purpose voice assistant. Its primary job is to preserve the meaning of spoken communication across languages with low delay. It can also return transcript deltas, meaning incremental pieces of text, for captions, subtitles, accessibility features, logs, or monitoring.

The model is currently positioned as a focused member of OpenAI's realtime model lineup. It is intended for live multilingual audio products, not for broad conversational workflows involving research, coding, tool calls, or assistant actions.

How the live translation workflow works

An application opens a dedicated translation session and sends audio as it becomes available. GPT-Realtime-Translate processes the incoming speech continuously and emits translated audio and transcript events. The result is closer to a live interpreter than to a traditional batch translation pipeline.

This approach can reduce the delay introduced by recording an entire sentence or conversation before processing it. It is useful when participants need to follow one another in near real time, although actual perceived latency and translation quality will still depend on the audio connection, pronunciation, background noise, language pair, and application design.

OpenAI describes the model as capable of handling natural speech, regional pronunciation, changing context, and domain-specific language. These are provider claims rather than independent benchmark results, so applications should test the model with representative speakers, terminology, and acoustic conditions before deployment.

Supported input and output types

GPT-Realtime-Translate is built around audio. It accepts audio input and produces translated audio output. It also provides text transcripts, but its core output is spoken translation rather than ordinary text generation.

  • Audio input: Supported.
  • Audio output: Supported; translated speech is returned during the session.
  • Transcript output: Supported through transcript deltas.
  • Text input: Not listed as a supported input modality for this model.
  • Image and video input: Not supported.
  • Streaming: Supported.
  • Music, image, and video generation: Not supported.

The transcript stream is especially useful when a product needs both spoken translation and visible text. For example, a meeting application could play the translated audio to one participant while displaying the same translation as captions for another.

Languages and translation scope

OpenAI states that GPT-Realtime-Translate supports more than 70 input languages and 13 output languages. The supplied documentation does not provide a complete language-by-language table, so teams should confirm that their specific source and target languages are supported before building a production workflow.

The model is most appropriate when the required operation is translation between spoken languages. It is less appropriate when the application must answer questions, maintain a broad assistant persona, invoke external tools, or take actions based on the conversation. Those requirements call for a general realtime voice model or a separate orchestration layer rather than this dedicated interpreter model.

API access and integration options

GPT-Realtime-Translate is available through OpenAI's dedicated Realtime translation endpoint: /v1/realtime/translations. OpenAI documents two principal connection methods.

  • WebRTC: Suitable for browser-based applications. A browser can send a microphone track and receive translated speech as a remote audio track.
  • WebSockets: Suitable for server-side systems that receive and process raw audio, including telephony, SIP, broadcast, and media-processing pipelines. The documented integration uses base64-encoded 24 kHz PCM16 audio and handles returned audio deltas directly.

The dedicated endpoint is an important distinction from ordinary assistant-style realtime APIs. GPT-Realtime-Translate continuously translates incoming audio and does not use the usual assistant response lifecycle described for general conversational sessions.

The supplied research does not identify tool calling, function calling, structured outputs, or fine-tuning support. Each is listed as unsupported. This limits the model's role to translation and transcription-related output rather than acting as an autonomous application controller.

Context window and technical limits

The documented context window is 16,000 tokens, and the maximum output limit is 2,000 tokens. Tokens are units used to represent pieces of text; in this model's primary workload, they are more relevant to transcript and session context than to a conventional long-form text response.

These limits should not be interpreted as a promise that a continuous meeting can be translated indefinitely without application-side session management. Developers should test how their chosen session design handles long conversations, transcript retention, interruptions, and reconnects. The available research does not specify a maximum audio-session duration or a separate retention policy for conversation history.

The model's listed knowledge cutoff is September 30, 2024. That cutoff concerns learned background knowledge. It does not prevent the model from processing live audio supplied during a translation session, but it does mean that the model should not be treated as a source of current world information.

Pricing and cost trade-offs

GPT-Realtime-Translate is priced at $0.034 per minute of realtime audio. The supplied model information describes this as duration-based pricing rather than separate text-token input and output rates. This makes the cost model relatively straightforward for products that can estimate the number of audio minutes they will process.

Actual spending will depend on total audio duration, concurrent sessions, retries, and the amount of traffic routed through the service. A product that keeps microphones open continuously may consume more billable audio time than one that sends audio only during active speech, although the research does not specify billing rules for silence or connection time. Confirm those details in the current OpenAI pricing documentation before forecasting production costs.

The price is a trade-off between specialized realtime behavior and the need to process audio continuously. A batch workflow that can wait for a recording to finish may have different cost and latency characteristics, while a general voice assistant may offer broader capabilities at the expense of being less focused on direct translation. The editorial assessment supplied for this model rates its speed and cost efficiency highly for live translation, but those are comparative evaluations rather than OpenAI-published benchmarks.

Reasoning, coding, and tool capabilities

GPT-Realtime-Translate is not intended to be a reasoning or coding model. Its editorial reasoning score is low, and its coding score is also low because the model is optimized for interpreting speech rather than solving programming problems or performing extended analysis. These scores are editorial judgments, not standardized provider measurements.

Tool use is listed as unsupported. The model should therefore not be selected when the translation system must directly search the web, query a database, call business functions, schedule appointments, or execute code. An application can still place the model inside a larger system, but any orchestration, business logic, or external actions would need to be handled outside the model.

There is also no listed support for structured outputs. Transcript deltas and translated audio can be consumed by software, but they should not be confused with a provider-guaranteed JSON schema or function-call format.

Main strengths and limitations

Strengths

  • Low-latency design: It translates streaming audio instead of requiring a completed recording before processing.
  • Speech-to-speech output: Applications receive translated speech directly, which is useful for conversations, calls, rooms, and broadcasts.
  • Caption support: Transcript deltas can support subtitles, accessibility features, and operational monitoring.
  • Dedicated integration paths: WebRTC supports browser audio, while WebSockets support server-side raw-audio systems.
  • Focused cost model: The published price is based on realtime audio duration at $0.034 per minute.
  • Broad stated language coverage: OpenAI reports more than 70 input languages and 13 output languages.

Limitations

  • Specialized purpose: It is an interpreter model, not a general voice assistant.
  • No listed tools or function calling: It cannot directly perform external actions through supported model tools.
  • No structured outputs: Applications needing guaranteed schema-constrained responses should use another design.
  • No image or video support: Visual inputs are outside the documented modality set.
  • No fine-tuning support: The supplied specifications list fine-tuning as unsupported.
  • Quality depends on real audio conditions: Pronunciation, background noise, specialist terms, turn-taking, and language pair can affect results.
  • Session planning is still required: The documented 16,000-token context window and 2,000-token maximum output do not define unlimited continuous-session behavior.

When to choose GPT-Realtime-Translate

Choose GPT-Realtime-Translate when the central requirement is fast spoken translation between people who use different languages. Good candidates include multilingual meetings, live lessons, translated customer-support calls, cross-border sales conversations, video rooms, broadcasts, and applications that need both translated audio and captions.

It is particularly suitable when waiting for a full recording would make the experience impractical. WebRTC is the natural starting point for a browser microphone experience, while WebSockets are better suited to backend services, telephony, SIP, broadcast, or other systems that already handle raw audio.

Choose a different option when translation is only a minor feature inside a broader assistant workflow. If the system must answer follow-up questions, maintain a complex conversation, call tools, browse for current information, write code, or take actions, a general realtime voice model or a multi-model architecture is more appropriate. Likewise, a batch speech pipeline may be preferable when latency is unimportant and the application needs a different processing or cost model.

Bottom line

GPT-Realtime-Translate is a narrowly focused OpenAI model for live spoken translation. Its defining advantages are streaming behavior, translated audio output, transcript deltas, dedicated WebRTC and WebSocket access, and a clear duration-based price. Its limitations are equally important: it is not a general assistant, does not provide listed tool or structured-output support, and is not designed for coding, visual understanding, or broad reasoning.

For a product whose main challenge is making multilingual speech understandable with minimal delay, this specialization is the reason to consider it. For products that need translation plus autonomous assistance, external actions, or extensive analysis, it should be treated as one component of a larger system rather than the complete solution.


Answers to Frequently Asked Questions

Is GPT-Realtime-Translate a general voice assistant with tool calling?
No. GPT-Realtime-Translate is designed primarily for live speech translation and transcription. The supplied specifications list tool calling, function calling, structured outputs, fine-tuning, image input, and video input as unsupported, so broader assistant actions require another model or an external orchestration layer.
How can developers integrate GPT-Realtime-Translate?
Developers can use the dedicated /v1/realtime/translations endpoint through WebRTC or WebSockets. WebRTC is suited to browser microphone applications, while WebSockets are designed for server-side raw-audio systems such as telephony, SIP, broadcast, and media-processing pipelines.
Which languages does GPT-Realtime-Translate support?
OpenAI states that GPT-Realtime-Translate supports more than 70 input languages and 13 output languages. The complete language list should be confirmed in the current documentation before production use.
What is GPT-Realtime-Translate?
GPT-Realtime-Translate is OpenAI’s specialized streaming speech-to-speech translation model. It processes incoming audio continuously and returns translated audio along with transcript deltas for captions, subtitles, accessibility features, or monitoring.
How much does GPT-Realtime-Translate cost?
GPT-Realtime-Translate is priced at $0.034 per minute of realtime audio. Actual spending may vary depending on audio duration, concurrent sessions, retries, and current OpenAI billing rules for silence or connection time.


Sources 3
Provider

About OpenAI