What GPT-4 Turbo Preview was
GPT-4 Turbo Preview was OpenAI’s preview release of the GPT-4 Turbo model family. OpenAI introduced the model for developers on November 6, 2023, initially using the gpt-4-1106-preview identifier. The later gpt-4-turbo-preview alias pointed to the gpt-4-0125-preview snapshot.
Its main practical distinction was the combination of GPT-4-level general-purpose text generation with a 128,000-token context window. A context window is the amount of text a model can consider in one request, including the prompt and the conversation or documents supplied to it. That capacity made the preview model suitable for long documents and extended conversations that would have exceeded the limits of earlier GPT-4 versions.
GPT-4 Turbo Preview is no longer an available production option. OpenAI deprecated the relevant preview snapshot, and the model was shut down on March 26, 2026. The specifications below are therefore historical rather than instructions for starting a new integration.
Specifications and limits
| Specification | Documented value |
|---|---|
| Provider | OpenAI |
| Model family | GPT-4 Turbo |
| Initial release | November 6, 2023 |
| Context window | 128,000 tokens |
| Maximum output | 4,096 tokens |
| Knowledge cutoff | December 1, 2023 |
| Input and output | Text input and text output |
| Current status | Retired; shut down March 26, 2026 |
The 128,000-token context limit describes the full request capacity, not a guaranteed amount of generated text. The maximum output was 4,096 tokens, so a long prompt could leave relatively little room for the response even when the overall context window was large. The December 1, 2023 knowledge cutoff also means the model should not be treated as up to date on events or information after that date.
Historical capabilities
At launch, OpenAI described GPT-4 Turbo Preview as supporting improved instruction following, JSON mode, parallel function calling, reproducible outputs through a seed parameter, and improved log-probability support. JSON mode was intended to make responses conform to valid JSON, which was useful when an application needed to pass model output into another program.
The preview model also supported fine-tuning according to its model documentation. Fine-tuning adapts a model to examples supplied by a developer, potentially making its responses more consistent for a specific style or task. In this case, that capability was relevant to specialized text-generation workflows rather than to image, audio, or video production.
There is an important distinction between launch-era descriptions and the later model profile. The supplied documentation records the current profile as not supporting streaming, function calling, or structured outputs, even though launch materials described JSON mode and function-calling features. These capabilities should therefore be treated as historical preview-era functionality, not as a dependable specification for a current service.
Historical pricing and cost trade-offs
GPT-4 Turbo Preview was priced at $10 per 1 million input tokens and $30 per 1 million output tokens. Input tokens are the text sent to the model; output tokens are the text it generates. The different rates meant that long generated answers could cost substantially more than short ones, even when the input document was the same size.
Those prices were substantially below the original GPT-4 pricing available around the model’s launch. The preview therefore targeted applications that needed a high-capability GPT-4 model and a large context window without paying the earlier GPT-4 rates. However, price alone is no longer a reason to select it: the model is retired and cannot accept new API requests.
Modalities and technical profile
GPT-4 Turbo Preview was a text model. It accepted text input and returned text output. It did not natively generate images, audio, or video, and the documented profile does not identify image, audio, or video input. It was consequently aimed at language tasks such as summarization, extraction, question answering, and code-related text generation rather than media creation or voice interaction.
Its reasoning and coding scores in the supplied model data are editorial evaluations, not OpenAI-published benchmark results. Both are recorded as 7 out of 10, while the speed score is also 7 out of 10, and the cost score is 4 out of 10. These values can provide a rough catalog-level comparison, but they should not be interpreted as standardized benchmark measurements or guarantees of application performance.
The model’s documented fine-tuning and historical JSON-mode support made it useful for structured text workflows. Tool and function support is more complicated: launch materials described function calling, while the later profile marked function calling as unsupported. Developers evaluating the historical model should not assume that every preview-era capability applied to every snapshot or API configuration.
Best use cases when it was available
GPT-4 Turbo Preview was most appropriate for workloads where a large amount of text had to be considered at once. Examples included:
- Summarizing long reports, manuals, transcripts, or collections of documents.
- Answering questions about substantial text supplied in the prompt.
- Generating structured text or JSON for downstream processing.
- Building general-purpose assistants with a larger conversational history.
- Supporting software-development workflows that benefited from GPT-4-class text generation.
- Fine-tuning text behavior for a specialized application, where the documented fine-tuning support was relevant.
The large context window was its clearest practical advantage. A user could provide more source material in one request than many earlier models allowed, reducing the need to split a document into many separate prompts. That did not eliminate the need to manage context carefully: longer requests still consumed tokens, increased processing requirements, and could leave less room for the answer.
Limitations and alternatives
The most important limitation is availability. GPT-4 Turbo Preview was a preview release rather than a current production model, and its underlying snapshot was shut down. New applications should use an actively supported model instead of relying on its identifier or attempting to reproduce its historical behavior.
It was also limited to text input and output. Applications requiring image generation, audio processing, video processing, or real-time voice interaction needed a different type of model. Similarly, developers who require currently supported structured outputs, dependable tool calling, or a maintained API contract should choose an available production option rather than treating the preview documentation as authoritative.
In capability-versus-cost terms, GPT-4 Turbo Preview occupied a historical middle ground: it offered a large context window and GPT-4-level general-purpose text generation at lower historical prices than the original GPT-4, but it was not a low-cost specialist model and was not optimized for media tasks. A smaller, actively supported model may be more appropriate for simple classification or high-volume generation, while a current high-capability model may be preferable for demanding reasoning, coding, or tool-driven applications. The supplied research does not identify a specific current replacement, so no particular successor should be assumed here.
When to choose GPT-4 Turbo Preview
For a new project, the answer is effectively never: the model is retired and unavailable. It may still be worth studying when maintaining historical code, interpreting archived evaluation results, documenting an old system, or comparing the design goals of early GPT-4 Turbo releases with current models.
When it was available, a developer would have chosen it for long-context text generation, document analysis, structured text responses, or a general assistant where the 128,000-token window mattered. A developer would have avoided it for new production systems, multimodal generation, real-time use, or any deployment that requires a supported and stable model lifecycle.
Bottom line
GPT-4 Turbo Preview was an important preview-era OpenAI model because it brought a 128,000-token context window and lower historical GPT-4 pricing to text-generation applications. Its strengths were long-context handling, general-purpose language generation, and documented historical support for JSON mode and fine-tuning. Its limitations included text-only operation, preview-era uncertainty around some capabilities, a 4,096-token output ceiling, and eventual retirement. Today it is best understood as a historical model specification, not as an option for a new API integration.

