GPT-4

GPT-4 Turbo

by OpenAI · Deprecated but currently accessible; scheduled for shutdown on October 23, 2026

GPT-4 Turbo is OpenAI’s deprecated multimodal model for text generation and image understanding. It offers a 128,000-token context window, 4,096-token maximum output, function calling, streaming, JSON mode, and API compatibility for existing applications, but its scheduled October 23, 2026 shutdown makes it unsuitable for most new long-term deployments.

Text Reasoning Coding
GPT-4 Turbo remains a relevant choice for existing applications that need a 128,000-token context window, image understanding, function calling, or compatibility with established GPT-4 Turbo integrations. However, it is an older model: the current alias points to gpt-4-turbo-2024-04-09, OpenAI recommends newer models such as GPT-4o for new applications, and access is scheduled to end on October 23, 2026.
Outputs

What GPT-4 Turbo can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming JSON mode Batch API
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
6/10 Speed
4/10 Cost efficiency
Specifications

Technical details

Model family GPT-4
Model type Multimodal
Context window 128K tokens
Maximum output 4K tokens
Knowledge cutoff December 1, 2023
Release date 2024-04-09
Status Deprecated but currently accessible; scheduled for shutdown on October 23, 2026
Deprecation date 2026-04-22
Shutdown date 2026-10-23
Knowledge cutoff notes

The current official model page lists December 1, 2023 as the knowledge cutoff. Earlier launch materials described the preview as containing knowledge through April 2023; that earlier date applies to the preview-era model rather than the current gpt-4-turbo-2024-04-09 snapshot.

Model notes

The gpt-4-turbo alias currently points to gpt-4-turbo-2024-04-09. The model page lists it as an older high-intelligence model and recommends newer models such as GPT-4o for new applications. It accepts text and image input and returns text. Function calling, streaming, Chat Completions, Responses, and Batch API usage are supported. The current model page lists Structured Outputs as unsupported, while OpenAI's GPT-4 Turbo launch documentation separately confirms JSON mode. OpenAI announced a shutdown date of October 23, 2026 for gpt-4-turbo and its associated snapshot.

Cost

Model pricing

Input $10 per 1 million input tokens
Output $30 per 1 million output tokens
Model guide

GPT-4 Turbo: OpenAI’s Legacy 128K-Context Vision Model

GPT-4 Turbo is OpenAI’s older multimodal model for high-context text generation and image understanding. It accepts text and images, produces text, supports function calling, streaming, Chat Completions, Responses, Batch API usage, and JSON mode, and provides a 128,000-token context window. The model is deprecated and scheduled for shutdown on October 23, 2026.

What is GPT-4 Turbo?

GPT-4 Turbo is an OpenAI model for general-purpose text generation and image understanding. It accepts text and image inputs and returns text. In practical terms, this means it can answer questions, summarize documents, follow instructions, analyze visual material, generate structured responses, and interact with application tools through function calling.

The current gpt-4-turbo alias points to the gpt-4-turbo-2024-04-09 snapshot. It belongs to the GPT-4 model family and is best understood today as a legacy model rather than OpenAI’s preferred option for new development. Its continuing value comes from its established API behavior, large context window, and support for image input and common application features.

Where GPT-4 Turbo fits in OpenAI’s lineup

GPT-4 Turbo was introduced as a more capable and less expensive successor to the original GPT-4. OpenAI now describes it as an older high-intelligence model and recommends newer models, including GPT-4o, for new applications. That positioning creates an important distinction: GPT-4 Turbo may still be technically suitable for an existing system, but its scheduled shutdown makes it a poor foundation for a new long-lived deployment.

The model can therefore make sense when compatibility is more important than adopting the newest model. An application already tuned to GPT-4 Turbo’s responses, tool calls, JSON-mode behavior, or image-analysis workflow may prefer a controlled migration rather than an immediate replacement. New projects should account for the model’s lifecycle status from the beginning.

Input, output, and supported capabilities

GPT-4 Turbo is multimodal on input but text-only on output. It can combine written instructions with images, but it does not natively generate images, audio, or video.

  • Text input: Supported.
  • Image input: Supported.
  • Text output: Supported.
  • Audio input or output: Not supported as a native modality.
  • Video input or output: Not supported as a native modality.
  • Image output: Not supported.
  • Function calling: Supported.
  • Streaming: Supported.
  • JSON mode: Supported.

Image understanding is useful for tasks such as describing visual content, extracting meaning from documents with figures, and answering questions about an image alongside a text prompt. The available research does not establish native image generation or media generation, so those capabilities should not be assumed from the model’s multimodal input support.

Context window and maximum output

GPT-4 Turbo has a context window of 128,000 tokens. A context window is the amount of text and other supported input that the model can consider in a request, including the conversation, instructions, and application-provided documents. This makes the model suitable for high-context tasks such as reviewing long documents or maintaining a substantial working prompt.

The maximum output is 4,096 tokens. The context window and output limit are separate: a request may contain a large amount of input, but the model’s generated response is still limited to 4,096 tokens. Applications that need very long responses may need to divide the task into multiple calls or use a model with a more suitable output limit.

The current model documentation lists a knowledge cutoff of December 1, 2023. Retrieval systems, external tools, and documents supplied by an application can provide newer information during a request, but they do not change the model’s underlying training-data cutoff.

GPT-4 Turbo pricing

OpenAI’s listed pricing is $10 per 1 million input tokens and $30 per 1 million output tokens. Input tokens are the text or other request content sent to the model, while output tokens are the content it generates. Because output tokens cost three times as much as input tokens, applications that produce lengthy responses should monitor response size as well as prompt volume.

These prices make GPT-4 Turbo a relatively expensive option compared with newer, smaller models. Its cost may still be justified when an application depends on its context capacity, image understanding, function-calling behavior, or existing GPT-4 Turbo compatibility. For straightforward, high-volume text tasks where the same compatibility is not required, a newer lower-cost option may be more appropriate.

API and structured-data support

GPT-4 Turbo supports Chat Completions, Responses, and Batch API usage. It also supports streaming, which allows an application to receive generated output progressively rather than waiting for the complete response. Function calling allows the model to request an application-defined function, such as retrieving data or initiating an operation; the application remains responsible for implementing and executing that function.

The model supports JSON mode, which can be used when an application needs the response formatted as a valid JSON object. JSON mode should not be confused with Structured Outputs. The current model documentation does not list Structured Outputs support, so developers should not assume that GPT-4 Turbo can enforce a particular JSON schema simply because it can produce JSON-formatted responses.

For an existing integration, this distinction matters. JSON mode can help with machine-readable output, but applications may still need validation, error handling, and recovery logic. If strict schema conformance is a core requirement, the model’s documented lack of Structured Outputs support is a limitation that should be considered before deployment.

Reasoning, coding, speed, and cost trade-offs

GPT-4 Turbo is positioned as a high-intelligence general-purpose model, but the supplied research does not provide a provider-published benchmark for reasoning or coding quality. Editorial evaluations in the supplied data rate its reasoning and coding capabilities at 7 out of 10; those are comparative editorial scores, not OpenAI specifications or benchmark results.

The same editorial data rates speed at 6 out of 10 and cost at 4 out of 10. These scores indicate a middle-ground assessment rather than a measured latency guarantee. In practical terms, GPT-4 Turbo’s trade-off is straightforward: it offers broad capability, image input, and a large context window, but it is not positioned as the newest, fastest, or least expensive option.

For coding, the model can be useful in established systems that need code-related discussion, document analysis, function calling, or structured responses. However, the supplied research does not establish a special coding mode or a particular software-development benchmark. Its coding suitability should therefore be evaluated against the requirements of the specific application rather than inferred from the GPT-4 name alone.

Best use cases for GPT-4 Turbo

GPT-4 Turbo is most defensible when an application needs several of its documented features together:

  • Legacy application maintenance: Existing services built around the GPT-4 Turbo API contract may continue using it while migration work is planned.
  • Long-document analysis: The 128,000-token context window can accommodate substantial prompts and application-provided documents.
  • Image-understanding workflows: The model can interpret images alongside text instructions while returning a textual analysis.
  • Function-calling systems: Applications can connect the model to external functions and services.
  • JSON-mode workflows: Applications that need JSON-formatted responses can use the documented JSON mode, with their own validation.
  • Established Chat Completions integrations: Teams that have already tested the model’s behavior may value compatibility over switching immediately.

These use cases are strongest when the application has a clear reason to remain on GPT-4 Turbo. They are weaker when the only requirement is general text generation, because newer models may offer a better cost, speed, capability, or lifecycle profile.

When should you choose GPT-4 Turbo?

Choose GPT-4 Turbo when you are maintaining an existing integration that depends on its documented behavior, needs a 128,000-token context window, uses image input, or relies on its function-calling and JSON-mode workflows. It can also be a temporary choice when a migration requires time for regression testing and output comparison.

Choose another option when you are starting a new project, need long-term availability beyond October 23, 2026, want the lowest practical cost, require newer reasoning capabilities, or need native audio, video, or image generation. OpenAI specifically recommends newer models such as GPT-4o for new applications. The appropriate replacement depends on the target system, but GPT-4 Turbo’s scheduled shutdown should be treated as a migration requirement rather than a distant theoretical risk.

Limitations and scheduled shutdown

GPT-4 Turbo has several limitations that affect current adoption. It is deprecated, its knowledge cutoff is December 1, 2023, its maximum output is 4,096 tokens, and it does not provide native audio, video, or image output. The current model documentation also does not list Structured Outputs support.

OpenAI has announced that access to gpt-4-turbo and the gpt-4-turbo-2024-04-09 snapshot will shut down on October 23, 2026. Existing users should inventory model calls, compare candidate replacements, test tool-call and JSON behavior, and verify image-analysis results before switching. Because model changes can affect formatting and application logic, migration should include output validation rather than only changing the model identifier.

Bottom line

GPT-4 Turbo is a capable legacy model whose strongest practical differentiators are its 128,000-token context window, text-and-image input, text output, function calling, streaming, and established API support. Its $10-per-million input-token and $30-per-million output-token pricing is no longer especially cost-focused, and its scheduled October 23, 2026 shutdown makes it unsuitable as a long-term default for most new projects.

For an existing system that depends on GPT-4 Turbo, continued use may be reasonable during a carefully managed transition. For new development, the model’s lifecycle status should outweigh familiarity unless a specific compatibility requirement makes it necessary.


Answers to Frequently Asked Questions

When will GPT-4 Turbo shut down, and should new projects use it?
OpenAI has announced that access to gpt-4-turbo and the gpt-4-turbo-2024-04-09 snapshot will shut down on October 23, 2026. It may remain useful for maintaining existing integrations, but OpenAI recommends newer models such as GPT-4o for new applications.
Does GPT-4 Turbo support image generation, audio, or video?
No. GPT-4 Turbo supports text and image inputs and returns text, but it does not natively generate images, audio, or video, and it does not support audio or video as native input or output modalities.
How much does GPT-4 Turbo cost?
GPT-4 Turbo costs $10 per 1 million input tokens and $30 per 1 million output tokens. Output tokens cost three times as much as input tokens, so applications should monitor response length as well as prompt size.
What is GPT-4 Turbo?
GPT-4 Turbo is an OpenAI legacy model for general-purpose text generation and image understanding. It accepts text and image inputs, produces text responses, and supports features such as function calling, streaming, and JSON mode.
What is the context window and maximum output of GPT-4 Turbo?
GPT-4 Turbo has a 128,000-token context window and a maximum output of 4,096 tokens. The context window determines how much input the model can process, while the output limit restricts the length of its generated response.


Sources 4
Provider

About OpenAI