What is GPT-Audio-1.5?
GPT-Audio-1.5 is OpenAI's generally available model for applications that need both audio understanding and spoken responses through the Chat Completions API. A developer can send a spoken prompt, optionally include text instructions, and receive a response as text, audio, or both. This makes the model suitable for voice-enabled chat interfaces, audio question answering, and assistants that need to connect spoken interaction with application tools.
The model was released to the Chat Completions API on February 23, 2026. Its canonical model ID is gpt-audio-1.5. It belongs to OpenAI's GPT-Audio family and is positioned as an audio-in, audio-out model rather than as a general-purpose image, video, or persistent realtime voice model.
In practical terms, GPT-Audio-1.5 handles a request-and-response workflow. The application sends an input, the model processes it, and the API returns a generated result. That differs from a persistent realtime session, where audio can flow continuously, users can interrupt the assistant, and the system is designed for rapid turn-taking. OpenAI's audio documentation positions realtime or live voice offerings as more appropriate for those continuous conversational experiences.
Input, output, and core capabilities
GPT-Audio-1.5 accepts text and audio input. It produces text and audio output, including spoken responses selected through the request's audio settings. Image and video input are not supported, and the model does not generate images or video.
The audio capability is useful when an application needs more than separate speech-to-text and text-to-speech steps. For example, a voice assistant can receive a user's spoken question, use the model to interpret the request and formulate an answer, and return that answer as speech while also retaining a text response for a transcript or user interface. Audio files or encoded audio content can be supplied alongside text instructions.
The model supports streaming, allowing output to be delivered progressively rather than waiting for the entire response to finish. It also supports function calling. Function calling lets the model request an application-defined operation, such as looking up an account record or creating a service ticket; the application remains responsible for executing the function and returning the result. This is particularly useful for voice interfaces that need to trigger backend actions.
Structured outputs are not supported, and fine-tuning is not supported. The supplied model information does not establish native first-party web-search support for GPT-Audio-1.5, so current information should be supplied through application context or an appropriate external tool when needed.
Technical specifications
| Specification | GPT-Audio-1.5 |
|---|---|
| Provider | OpenAI |
| Model family | GPT-Audio |
| Model ID | gpt-audio-1.5 |
| Primary API | Chat Completions |
| Context window | 128,000 tokens |
| Maximum output | 16,384 tokens |
| Knowledge cutoff | September 30, 2024 |
| Input | Text and audio |
| Output | Text and audio |
| Streaming | Supported |
| Function calling | Supported |
| Structured outputs | Not supported |
| Fine-tuning | Not supported |
| Batch API | Available according to the supplied model information |
The 128,000-token context window is the maximum amount of model context available for an interaction, including the relevant input and conversation history. It is not an audio duration limit. Audio is billed in audio tokens, and the supplied research does not specify a fixed maximum duration for an individual audio file or request.
The September 30, 2024 knowledge cutoff is important for applications that ask about current events, changing prices, live schedules, or other recently updated information. The model can still use information supplied in the request or returned by an application tool, but that external context does not change the underlying cutoff.
Pricing and API access
GPT-Audio-1.5 uses separate rates for text and audio tokens. Text input costs $2.50 per 1 million tokens, while text output costs $10.00 per 1 million tokens. Audio input costs $32.00 per 1 million audio tokens, and audio output costs $64.00 per 1 million audio tokens.
| Usage type | Price |
|---|---|
| Text input | $2.50 per 1 million tokens |
| Text output | $10.00 per 1 million tokens |
| Audio input | $32.00 per 1 million audio tokens |
| Audio output | $64.00 per 1 million audio tokens |
These are usage-based API prices, not a recurring consumer subscription price. Text and audio use different billing units, so the final cost depends on the mixture of spoken input, spoken output, text instructions, generated text, and conversation history. A voice application that returns both audio and text can incur charges for both output types.
GPT-Audio-1.5 is available through the Chat Completions API. Audio requests can specify the audio modality, a voice, and an output format. Streaming and function calling are supported, while structured outputs and fine-tuning are not. Batch API availability is listed in the supplied model information.
Reasoning, coding, speed, and cost trade-offs
OpenAI's supplied model information does not provide a public benchmark score for reasoning or coding on GPT-Audio-1.5. The editorial assessment assigns a reasoning score of 5 out of 10 and a coding score of 4 out of 10. These are subjective positioning scores, not provider-published measurements. They reflect that GPT-Audio-1.5 is primarily designed for audio conversation rather than for advanced text reasoning or software development.
The same editorial assessment gives the model a speed score of 7 out of 10 and a cost score of 5 out of 10. These scores are also editorial rather than official specifications. The practical cost trade-off is clearer from the published prices: audio tokens are substantially more expensive than text tokens, particularly for audio output. Applications should avoid generating speech when text is sufficient, and should manage conversation history carefully when usage volume matters.
Compared with a text-only model, GPT-Audio-1.5 adds direct audio input and spoken output but may be a less economical choice for workloads that never use audio. Compared with a dedicated realtime voice option, it offers a standard Chat Completions workflow with streaming and function calling, but it is not the documented choice for persistent sessions, interruption handling, or continuous low-latency speech interaction.
Best use cases
- Voice-enabled chat: Build an assistant that accepts spoken questions and answers aloud while optionally displaying a text transcript.
- Audio question answering: Let users submit an audio recording with instructions and receive a spoken or written response.
- Tool-enabled voice interfaces: Use function calling to connect spoken requests to application actions such as account lookups, scheduling, or support workflows.
- Spoken-content analysis: Process audio together with text instructions when an application needs an audio-aware conversational response.
- Hybrid voice and text experiences: Return audio for listening and text for captions, search, logging, or accessibility features.
- Request-and-response audio workflows: Use the model when each interaction can be handled as an API request rather than as an always-open conversation.
Function calling is especially useful when the assistant must do something beyond generating speech. For example, a customer-support application could receive a spoken question, ask the model to identify the required backend operation, execute the approved function, and then provide the result in both text and audio. The application should validate tool arguments and enforce permissions rather than treating the model's function request as an authorization decision.
Limitations to consider
GPT-Audio-1.5 does not accept images or video and cannot generate visual or video content. It also lacks structured outputs, so it is not the best fit for workflows that require the model to return data conforming reliably to a predefined JSON schema. Fine-tuning is unavailable, which limits model customization for organizations that need to train behavior on their own examples.
Audio pricing is another important limitation. At $32.00 per 1 million audio input tokens and $64.00 per 1 million audio output tokens, audio-heavy workloads require more careful cost management than text-only interactions. The model's 16,384-token maximum output is ample for ordinary conversational replies, but applications should still constrain responses when long spoken output would be expensive or inconvenient for users.
The model is not documented as a persistent realtime voice system. If an application depends on immediate turn-taking, user interruptions, continuous audio exchange, or session-oriented voice behavior, OpenAI's realtime or live voice offerings may be more appropriate. GPT-Audio-1.5 is better understood as an audio-capable Chat Completions model than as a replacement for every realtime speech architecture.
Finally, the September 30, 2024 knowledge cutoff means that the model should not be relied on for current facts without supplied context or an external retrieval tool. The research also does not establish native web search for this exact model.
When to choose GPT-Audio-1.5
Choose GPT-Audio-1.5 when the central requirement is a straightforward API interaction that can understand audio and return natural-language speech, especially when the application also needs text responses, streaming, or function calling. It is a strong fit for voice assistants, audio-aware support tools, and other applications where the model's audio input and output are core features rather than optional extras.
Choose a text-focused model instead when the workload has no audio, requires lower text-processing costs, depends heavily on structured JSON output, or is primarily advanced reasoning and coding. Choose a realtime voice-oriented option when the product needs an ongoing low-latency conversation with interruptions and continuous session behavior. The decision therefore depends less on whether GPT-Audio-1.5 can produce a spoken answer and more on whether Chat Completions request-and-response behavior matches the application's interaction design.
For teams that do choose it, the main implementation decisions are whether to request text, audio, or both; how to stream the response; which voice and audio format to use; and where function calling should be permitted. Keeping spoken responses concise, limiting unnecessary audio generation, and providing current information through controlled application tools can improve both cost and reliability.

