GPT-3.5

GPT-3.5 Turbo

by OpenAI · Deprecated; still available through the OpenAI API

OpenAI's GPT-3.5 Turbo is a deprecated but still accessible text-only API model for inexpensive chat, summarization, classification, extraction, rewriting, and basic coding tasks. It offers a 16,385-token context window, 4,096-token maximum output, September 2021 knowledge cutoff, low input and output token prices, and fine-tuning support, but newer models are generally better for new applications.

Text Reasoning Coding
GPT-3.5 Turbo helped make chat-oriented language-model APIs affordable for developers. The current gpt-3.5-turbo alias is still available through OpenAI's API, but OpenAI classifies the model as deprecated and recommends newer options for most new projects. It is best understood as a low-cost, text-only model for straightforward workloads rather than a choice for advanced reasoning, multimodal input, or long-term systems that need the latest capabilities.
Outputs

What GPT-3.5 Turbo can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning Batch API
Model profile

Performance characteristics

4/10 Reasoning
4/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family GPT-3.5
Model type General Purpose
Context window 16K tokens
Maximum output 4K tokens
Knowledge cutoff 2021-09-01
Release date 2023-03-01
Status Deprecated; still available through the OpenAI API
Knowledge cutoff notes

OpenAI's current model documentation explicitly lists September 1, 2021 as the knowledge cutoff. Web search or external retrieval can provide newer information during use but does not change the underlying cutoff.

Model notes

The canonical gpt-3.5-turbo alias remains API-accessible as of September 23, 2026, but OpenAI classifies GPT-3.5 Turbo as a deprecated legacy model and recommends gpt-4o-mini for new use. The current model page lists text-only input and output, a 16,385-token context window, a 4,096-token maximum output, streaming unsupported, function calling unsupported, structured outputs unsupported, and fine-tuning supported. Deprecated snapshots include gpt-3.5-turbo-0125 and gpt-3.5-turbo-1106. Pricing is per 1 million tokens. Editorial scores are comparative estimates rather than provider benchmarks.

Cost

Model pricing

Input $0.50 per 1 million input tokens
Output $1.50 per 1 million output tokens
Model guide

GPT-3.5 Turbo: An Affordable Legacy Model for Text-Only API Workloads

GPT-3.5 Turbo is OpenAI's deprecated, text-only language model for inexpensive conversational responses, summarization, classification, extraction, rewriting, basic coding assistance, and other high-volume API workloads. Its current alias remains accessible through the OpenAI API and offers a 16,385-token context window, a 4,096-token maximum output, low per-token pricing, and fine-tuning support, but newer models are generally more appropriate for new applications.

What is GPT-3.5 Turbo?

GPT-3.5 Turbo is a general-purpose language model from OpenAI. It receives text and generates text, making it suitable for conversational applications and routine language-processing tasks. Developers have used it for chatbots, summarization, rewriting, classification, information extraction, lightweight content generation, and basic coding assistance.

The model was introduced for API access on March 1, 2023. Its current API alias is gpt-3.5-turbo, and its primary interface is OpenAI's Chat Completions endpoint. Although the model remains accessible, OpenAI now places GPT-3.5 Turbo in its deprecated legacy category. This status matters because the unpinned alias can point to the provider's currently designated GPT-3.5 Turbo version, while older snapshots may be retired or replaced.

Current status and positioning

GPT-3.5 Turbo is not OpenAI's recommended default for new applications. OpenAI recommends GPT-4o mini instead for new projects, describing that model as more capable, multimodal, and cheaper. That recommendation does not make GPT-3.5 Turbo unusable: the model can still be practical when an existing integration is already tuned for its behavior, when a workload is strictly text-based, or when a simple task benefits from a low token price.

The main trade-off is straightforward. GPT-3.5 Turbo can provide inexpensive and generally fast text processing, but it has an older knowledge cutoff, a smaller context window than some newer alternatives, fewer documented capabilities on its current model page, and weaker suitability for demanding reasoning or agentic workflows. The comparative reasoning, coding, speed, and cost scores in the supplied research are editorial estimates, not provider-published benchmarks.

Verified technical specifications

SpecificationGPT-3.5 Turbo
ProviderOpenAI
Model IDgpt-3.5-turbo
Model familyGPT-3.5
StatusDeprecated, but still available through the OpenAI API
Context window16,385 tokens
Maximum output4,096 tokens
Knowledge cutoffSeptember 1, 2021
Input and outputText only
Primary interfaceChat Completions
Fine-tuningSupported, subject to current eligibility and policy requirements

A token is a unit of text used for processing and billing; it may represent part of a word, a whole short word, punctuation, or whitespace. The 16,385-token context window is the combined space available for the conversation or other supplied input and the generated response. The output limit is separate in practical use: GPT-3.5 Turbo can generate up to 4,096 output tokens according to the supplied model documentation.

The September 1, 2021 knowledge cutoff means the model should not be treated as a reliable source for events or facts that emerged after that date. External retrieval can provide current information to an application, but retrieval does not update the model's underlying knowledge.

Pricing and cost trade-offs

OpenAI's listed pricing is $0.50 per 1 million input tokens and $1.50 per 1 million output tokens. Input tokens are the text sent to the model, while output tokens are the text it generates. Because output tokens cost more than input tokens, applications that request long responses can spend more even when their prompts are relatively small.

This pricing makes GPT-3.5 Turbo attractive for high-volume text operations such as short classifications, extraction, rewriting, and routine support responses. Cost alone should not determine the choice, however. A cheaper model can require more retries, stricter prompt engineering, or additional validation if it performs less reliably on a particular task. GPT-3.5 Turbo's age, 16,385-token context limit, and text-only design should be included in any total-cost comparison.

Supported modalities and tool capabilities

GPT-3.5 Turbo is text-only. The current model documentation identifies image, audio, and video input as unsupported, and the model does not directly produce images, audio, video, or other non-text media. It is therefore unsuitable as the core model for applications that need users to submit photographs, recordings, or video for native model analysis.

The supplied current model information lists streaming, function calling, and structured outputs as unsupported for this canonical model page. This is important for developers who need dependable machine-readable responses or tool-driven workflows. Historical GPT-3.5 Turbo snapshots, including gpt-3.5-turbo-0613, supported function calling before retirement, but that historical capability should not be assumed for the current canonical alias.

Fine-tuning is listed as supported, subject to OpenAI's current platform eligibility and policy restrictions. Fine-tuning can be relevant when a legacy application needs more consistent behavior for a narrow task, but it does not change the model's text-only modality, knowledge cutoff, context limit, or deprecated status.

Main strengths and limitations

Where GPT-3.5 Turbo works well

  • Low operating cost: Its per-token prices suit large volumes of relatively simple text requests.
  • Routine language processing: It can handle summarization, rewriting, classification, extraction, and straightforward generation.
  • Conversational applications: It remains suitable for simple chatbots and legacy support interfaces that do not need advanced reasoning.
  • Established integrations: Existing applications may already be tuned to its response style, token behavior, and prompt format.
  • Fine-tuning availability: Eligible users can consider fine-tuning for supported specialized workloads.

Where it falls short

  • Deprecated status: It is a legacy model rather than a forward-looking default for new systems.
  • Limited knowledge currency: Its knowledge cutoff is September 1, 2021, so current-information tasks require external retrieval.
  • Restricted context and output: The context window is 16,385 tokens and the maximum output is 4,096 tokens.
  • No native multimodal processing: It cannot directly accept images, audio, or video.
  • Limited current tool support: The supplied model page lists function calling, structured outputs, and streaming as unsupported.
  • Not intended for demanding reasoning: Advanced analysis, complex agentic workflows, and difficult coding tasks may require a more capable current model.

GPT-3.5 Turbo is a reasonable choice when the task is primarily text-based, repetitive, cost-sensitive, and not dependent on current facts. Examples include:

  • High-volume text classification and routing
  • Extracting fields from short, predictable documents
  • Summarizing or rewriting user-provided text
  • Generating routine descriptions, drafts, or support responses
  • Simple customer-service chatbots
  • Basic coding assistance where advanced reasoning is not required
  • Existing applications that have already been tested and tuned around GPT-3.5 Turbo

Applications should validate generated output when incorrect extraction, classification, or wording could cause operational or financial harm. The model's low price does not remove the need for application-level testing and safeguards.

When to choose GPT-3.5 Turbo

Choose GPT-3.5 Turbo when a legacy integration already depends on it, the workload is text-only, the prompts and outputs fit within its limits, and low token cost is more important than access to newer capabilities. It can also be sensible for simple tasks where a more advanced model would add cost without providing a meaningful benefit.

Choose a newer small model instead when starting a new project, especially if you need stronger overall capability, multimodal input, current platform features, more dependable structured responses, or a better long-term support position. OpenAI specifically recommends GPT-4o mini for new projects in the supplied research. A current alternative may also be preferable when the application requires advanced reasoning, current information, image or audio handling, complex tool use, or a longer-lived model contract.

Migration considerations

Moving from GPT-3.5 Turbo to another model can change output style, instruction-following behavior, token usage, response length, and supported features. Do not assume that a replacement will produce identical wording or formatting. Test representative prompts, edge cases, latency expectations, token costs, and any downstream parser that consumes the response.

Migration deserves particular attention when an application informally asks for JSON. The supplied current model information does not verify structured-output support for GPT-3.5 Turbo, so applications that need guaranteed machine-readable output should not treat prompt instructions alone as an equivalent feature. Existing systems should also review any dependence on the older 4,096-token output limit, legacy snapshots, or historical function-calling behavior.

Bottom line

GPT-3.5 Turbo remains a low-cost option for straightforward, text-only API work, particularly in established applications and high-volume processing pipelines. Its 16,385-token context window, 4,096-token output ceiling, September 2021 knowledge cutoff, and deprecated status define its practical boundaries. For new systems that need current capabilities or a longer-term foundation, the provider's recommended newer models deserve evaluation before committing to GPT-3.5 Turbo.


Answers to Frequently Asked Questions

Should new projects use GPT-3.5 Turbo or a newer model?
New projects should generally evaluate a newer model, especially when they need multimodal input, stronger reasoning, current platform features, dependable structured outputs, complex tool use, or a longer-term support position. OpenAI specifically recommends GPT-4o mini for new projects in the supplied research. GPT-3.5 Turbo can still make sense for simple text-only tasks or established integrations where low cost is the priority.
How much does GPT-3.5 Turbo cost?
The listed price is $0.50 per 1 million input tokens and $1.50 per 1 million output tokens. Output tokens cost more than input tokens, so applications that generate long responses may incur higher costs.
What are GPT-3.5 Turbo's context window, output limit, and knowledge cutoff?
GPT-3.5 Turbo has a 16,385-token context window and a maximum output of 4,096 tokens. Its knowledge cutoff is September 1, 2021, so applications requiring current information should use external retrieval or a newer model.
What is GPT-3.5 Turbo best used for?
GPT-3.5 Turbo is best suited to straightforward, text-only workloads such as classification, summarization, rewriting, information extraction, routine content generation, simple chatbots, and basic coding assistance. It is particularly practical for high-volume, cost-sensitive applications and existing integrations.
Is GPT-3.5 Turbo still available through the OpenAI API?
Yes. GPT-3.5 Turbo remains available through the OpenAI API, but OpenAI categorizes it as a deprecated legacy model rather than a recommended default for new applications. Its current API alias is gpt-3.5-turbo.


Sources 6
Provider

About OpenAI