What is GPT-5 Chat?
GPT-5 Chat is an OpenAI API model alias that points to the GPT-5 snapshot previously used in ChatGPT. Its purpose is to provide ChatGPT-aligned conversational behavior through OpenAI's developer APIs, rather than to serve as a separate consumer ChatGPT subscription or a general-purpose product category.
The canonical API identifier is gpt-5-chat-latest. OpenAI lists the model for both the Chat Completions API and the Responses API. That makes it suitable for applications that need ordinary text generation as well as applications that connect the model to tools or structured application workflows.
The most important qualification is its status: OpenAI currently marks GPT-5 Chat as deprecated. The supplied model information does not provide a specific shutdown date. Developers beginning a new project should therefore treat it as a compatibility target or an evaluation reference, not as the default choice for a long-lived integration.
Inputs, outputs, and supported capabilities
GPT-5 Chat accepts text and image inputs and returns text. Image input allows the model to interpret visual material alongside written instructions, which is useful for tasks such as explaining a diagram, answering questions about an image, or extracting meaning from a visual reference. The model does not natively accept audio or video according to the supplied specifications.
Its output is text-only. GPT-5 Chat does not generate images, audio, video, music, embeddings, or speech. This distinction matters when selecting an architecture: an application can provide an image to GPT-5 Chat for analysis, but it cannot use the model itself as an image-generation or voice-output engine.
- Text input: Supported
- Image input: Supported
- Audio input: Not supported
- Video input: Not supported
- Text output: Supported
- Image, audio, video, music, embedding, and speech output: Not supported
- Streaming: Supported
- Function calling: Supported
- Structured outputs: Supported
Context window and maximum output
The model has a 128,000-token context window. A token is a unit of text used by the model; it may represent a whole short word, part of a longer word, punctuation, or other text. The context window covers the material the model can consider during a request, including the conversation, instructions, supplied documents, and other input content.
A 128,000-token window is large enough for substantial conversations and document-based tasks, but it is not unlimited. Applications that routinely process very large archives or long-running histories may need to summarize, retrieve only relevant passages, or choose a newer model with a larger context capacity if one is suitable for the workload.
The maximum output is 16,384 tokens. This is the upper limit for the generated response, not a promise that every request will produce an output of that length. In practical use, applications should still set sensible output limits to control latency, cost, and response size.
GPT-5 Chat pricing
OpenAI lists GPT-5 Chat at $1.25 per 1 million input tokens and $10 per 1 million output tokens. Cached input is priced at $0.125 per 1 million cached input tokens. These are API usage prices; they are not ChatGPT subscription prices.
| Usage type | Price per 1 million tokens |
|---|---|
| Input | $1.25 |
| Cached input | $0.125 |
| Output | $10.00 |
Input and output are charged separately, so an application that generates long answers can spend substantially more on output than on input. Prompt caching can reduce the cost of repeated input content when the same or similar prompt material is reused, but the supplied information does not specify the precise cache-duration or eligibility rules. Batch processing is also listed as supported, which can be useful for workloads that do not require immediate responses.
Reasoning, coding, and tool support
GPT-5 Chat supports function calling, also known as tool calling. This allows an application to describe an available function, receive the model's requested arguments, execute the function in application code, and return the result to the model. Typical uses include retrieving records, calling business systems, validating information, or triggering controlled application actions. The model does not independently perform those external actions; the surrounding application remains responsible for execution, permissions, and validation.
Structured outputs are supported as well. In a structured-output workflow, the application asks for a response that follows a defined schema rather than relying on free-form prose. This can make the model more useful for extracting fields, producing workflow data, or passing consistent results to another program. Structured outputs should not automatically be treated as proof of a separate legacy JSON mode: the supplied research leaves the standalone json_mode capability unverified.
The supplied editorial evaluation rates GPT-5 Chat's reasoning and coding capabilities at 8 out of 10. Those scores are editorial assessments, not OpenAI-published benchmark results. They indicate that the model is considered suitable for general reasoning, text-based problem solving, and coding assistance, but they should not be interpreted as a formal performance guarantee or as evidence that it is the best option for every programming or reasoning task.
Likewise, the supplied editorial ratings give the model a speed score of 8 out of 10 and a cost score of 7 out of 10. These are comparative editorial judgments. The verified pricing and feature information supports a more concrete conclusion: GPT-5 Chat combines moderate input pricing with considerably higher output pricing, while its deprecated status may be a more important operational concern than small differences in speed or cost.
Main strengths and trade-offs
GPT-5 Chat's primary strength is its alignment with the GPT-5 behavior previously used in ChatGPT. That makes it relevant when an existing application was built around that behavior and changing models could affect prompts, output style, tool interactions, or evaluation results.
It also combines several useful application features in one text-generating model:
- Text generation for conversational interfaces and content workflows.
- Image understanding for visual question answering and image-aware assistance.
- A 128,000-token context window for long conversations and substantial source material.
- Function calling for applications that need controlled interaction with external tools.
- Structured outputs for machine-readable responses.
- Streaming for displaying partial responses as they are generated.
- Prompt caching and batch processing for selected cost or throughput patterns.
Its trade-offs are equally important. The model is deprecated, does not support native audio or video, produces text only, and does not support fine-tuning. Its 128,000-token context window may also be less suitable than newer alternatives for applications built around very large inputs. The available research does not establish an exact retirement date, so teams using it should monitor OpenAI's model catalog and prepare a migration path.
Best use cases for GPT-5 Chat
GPT-5 Chat is most appropriate when compatibility with the GPT-5 Chat snapshot matters more than adopting the newest available model. Examples include maintaining an existing ChatGPT-aligned assistant, reproducing historical evaluations, or testing whether a migration changes conversational behavior.
It can also fit applications that need a combination of text and image understanding without native audio or video processing. Suitable tasks include image-aware customer support, document or diagram explanation, structured extraction from visual material, text drafting, question answering, and tool-enabled conversational workflows.
For a production integration, the model's function-calling and structured-output support can help separate natural-language interaction from application logic. The model can propose a tool call or produce schema-conforming data, while the application verifies the request and performs any consequential operation.
When another option may be more appropriate
A current OpenAI model is generally a better starting point for a new deployment because GPT-5 Chat is deprecated. The supplied research does not identify a specific replacement model, so it would be inappropriate to claim that one named successor has identical behavior or capabilities. Instead, teams should compare current ChatGPT-aligned models against their own prompts, tool definitions, latency requirements, and evaluation set.
Another option may also be preferable when the application needs native audio or video input, image or audio generation, fine-tuning, or a context window larger than 128,000 tokens. GPT-5 Chat cannot provide those functions directly. A multimodel architecture may be necessary when a workflow combines text reasoning with speech, video, or non-text generation.
Cost-sensitive applications should examine the balance between input and output usage. The $1.25 per million input-token price is much lower than the $10 per million output-token price, so concise responses, caching of repeated prompts, and batch processing where appropriate may matter more than the headline model price. These optimizations do not resolve the deprecation issue, however; they only improve the economics of an integration that still depends on this model.
Bottom line
GPT-5 Chat is a capable, ChatGPT-aligned OpenAI API model for text generation, image understanding, structured responses, streaming, and tool-enabled applications. Its verified limits are a 128,000-token context window and a 16,384-token maximum output, with API pricing of $1.25 per million input tokens, $0.125 per million cached input tokens, and $10 per million output tokens.
Its defining practical limitation is not a missing feature but its catalog status. OpenAI marks GPT-5 Chat as deprecated and does not provide an exact shutdown date in the supplied information. Use it when preserving GPT-5 Chat compatibility is important; for new systems, evaluate a current model instead and avoid making this deprecated alias a long-term dependency.

