What is GPT-3.5 Turbo?
GPT-3.5 Turbo is a general-purpose language model from OpenAI. It receives text and generates text, making it suitable for conversational applications and routine language-processing tasks. Developers have used it for chatbots, summarization, rewriting, classification, information extraction, lightweight content generation, and basic coding assistance.
The model was introduced for API access on March 1, 2023. Its current API alias is gpt-3.5-turbo, and its primary interface is OpenAI's Chat Completions endpoint. Although the model remains accessible, OpenAI now places GPT-3.5 Turbo in its deprecated legacy category. This status matters because the unpinned alias can point to the provider's currently designated GPT-3.5 Turbo version, while older snapshots may be retired or replaced.
Current status and positioning
GPT-3.5 Turbo is not OpenAI's recommended default for new applications. OpenAI recommends GPT-4o mini instead for new projects, describing that model as more capable, multimodal, and cheaper. That recommendation does not make GPT-3.5 Turbo unusable: the model can still be practical when an existing integration is already tuned for its behavior, when a workload is strictly text-based, or when a simple task benefits from a low token price.
The main trade-off is straightforward. GPT-3.5 Turbo can provide inexpensive and generally fast text processing, but it has an older knowledge cutoff, a smaller context window than some newer alternatives, fewer documented capabilities on its current model page, and weaker suitability for demanding reasoning or agentic workflows. The comparative reasoning, coding, speed, and cost scores in the supplied research are editorial estimates, not provider-published benchmarks.
Verified technical specifications
| Specification | GPT-3.5 Turbo |
|---|---|
| Provider | OpenAI |
| Model ID | gpt-3.5-turbo |
| Model family | GPT-3.5 |
| Status | Deprecated, but still available through the OpenAI API |
| Context window | 16,385 tokens |
| Maximum output | 4,096 tokens |
| Knowledge cutoff | September 1, 2021 |
| Input and output | Text only |
| Primary interface | Chat Completions |
| Fine-tuning | Supported, subject to current eligibility and policy requirements |
A token is a unit of text used for processing and billing; it may represent part of a word, a whole short word, punctuation, or whitespace. The 16,385-token context window is the combined space available for the conversation or other supplied input and the generated response. The output limit is separate in practical use: GPT-3.5 Turbo can generate up to 4,096 output tokens according to the supplied model documentation.
The September 1, 2021 knowledge cutoff means the model should not be treated as a reliable source for events or facts that emerged after that date. External retrieval can provide current information to an application, but retrieval does not update the model's underlying knowledge.
Pricing and cost trade-offs
OpenAI's listed pricing is $0.50 per 1 million input tokens and $1.50 per 1 million output tokens. Input tokens are the text sent to the model, while output tokens are the text it generates. Because output tokens cost more than input tokens, applications that request long responses can spend more even when their prompts are relatively small.
This pricing makes GPT-3.5 Turbo attractive for high-volume text operations such as short classifications, extraction, rewriting, and routine support responses. Cost alone should not determine the choice, however. A cheaper model can require more retries, stricter prompt engineering, or additional validation if it performs less reliably on a particular task. GPT-3.5 Turbo's age, 16,385-token context limit, and text-only design should be included in any total-cost comparison.
Supported modalities and tool capabilities
GPT-3.5 Turbo is text-only. The current model documentation identifies image, audio, and video input as unsupported, and the model does not directly produce images, audio, video, or other non-text media. It is therefore unsuitable as the core model for applications that need users to submit photographs, recordings, or video for native model analysis.
The supplied current model information lists streaming, function calling, and structured outputs as unsupported for this canonical model page. This is important for developers who need dependable machine-readable responses or tool-driven workflows. Historical GPT-3.5 Turbo snapshots, including gpt-3.5-turbo-0613, supported function calling before retirement, but that historical capability should not be assumed for the current canonical alias.
Fine-tuning is listed as supported, subject to OpenAI's current platform eligibility and policy restrictions. Fine-tuning can be relevant when a legacy application needs more consistent behavior for a narrow task, but it does not change the model's text-only modality, knowledge cutoff, context limit, or deprecated status.
Main strengths and limitations
Where GPT-3.5 Turbo works well
- Low operating cost: Its per-token prices suit large volumes of relatively simple text requests.
- Routine language processing: It can handle summarization, rewriting, classification, extraction, and straightforward generation.
- Conversational applications: It remains suitable for simple chatbots and legacy support interfaces that do not need advanced reasoning.
- Established integrations: Existing applications may already be tuned to its response style, token behavior, and prompt format.
- Fine-tuning availability: Eligible users can consider fine-tuning for supported specialized workloads.
Where it falls short
- Deprecated status: It is a legacy model rather than a forward-looking default for new systems.
- Limited knowledge currency: Its knowledge cutoff is September 1, 2021, so current-information tasks require external retrieval.
- Restricted context and output: The context window is 16,385 tokens and the maximum output is 4,096 tokens.
- No native multimodal processing: It cannot directly accept images, audio, or video.
- Limited current tool support: The supplied model page lists function calling, structured outputs, and streaming as unsupported.
- Not intended for demanding reasoning: Advanced analysis, complex agentic workflows, and difficult coding tasks may require a more capable current model.
Recommended use cases
GPT-3.5 Turbo is a reasonable choice when the task is primarily text-based, repetitive, cost-sensitive, and not dependent on current facts. Examples include:
- High-volume text classification and routing
- Extracting fields from short, predictable documents
- Summarizing or rewriting user-provided text
- Generating routine descriptions, drafts, or support responses
- Simple customer-service chatbots
- Basic coding assistance where advanced reasoning is not required
- Existing applications that have already been tested and tuned around GPT-3.5 Turbo
Applications should validate generated output when incorrect extraction, classification, or wording could cause operational or financial harm. The model's low price does not remove the need for application-level testing and safeguards.
When to choose GPT-3.5 Turbo
Choose GPT-3.5 Turbo when a legacy integration already depends on it, the workload is text-only, the prompts and outputs fit within its limits, and low token cost is more important than access to newer capabilities. It can also be sensible for simple tasks where a more advanced model would add cost without providing a meaningful benefit.
Choose a newer small model instead when starting a new project, especially if you need stronger overall capability, multimodal input, current platform features, more dependable structured responses, or a better long-term support position. OpenAI specifically recommends GPT-4o mini for new projects in the supplied research. A current alternative may also be preferable when the application requires advanced reasoning, current information, image or audio handling, complex tool use, or a longer-lived model contract.
Migration considerations
Moving from GPT-3.5 Turbo to another model can change output style, instruction-following behavior, token usage, response length, and supported features. Do not assume that a replacement will produce identical wording or formatting. Test representative prompts, edge cases, latency expectations, token costs, and any downstream parser that consumes the response.
Migration deserves particular attention when an application informally asks for JSON. The supplied current model information does not verify structured-output support for GPT-3.5 Turbo, so applications that need guaranteed machine-readable output should not treat prompt instructions alone as an equivalent feature. Existing systems should also review any dependence on the older 4,096-token output limit, legacy snapshots, or historical function-calling behavior.
Bottom line
GPT-3.5 Turbo remains a low-cost option for straightforward, text-only API work, particularly in established applications and high-volume processing pipelines. Its 16,385-token context window, 4,096-token output ceiling, September 2021 knowledge cutoff, and deprecated status define its practical boundaries. For new systems that need current capabilities or a longer-term foundation, the provider's recommended newer models deserve evaluation before committing to GPT-3.5 Turbo.

