What is GPT-5.5?
GPT-5.5 is a general-purpose model from OpenAI aimed at complex professional work. It is designed for tasks such as software engineering, long-form research, technical analysis, data work, document generation, spreadsheet workflows, and agents that need to plan, call tools, inspect results, and continue working toward an objective.
The model is available through OpenAI’s API and is also used in ChatGPT and Codex experiences. In the API, the current model alias is gpt-5.5, while the documented dated snapshot is gpt-5.5-2026-04-23. The alias is suitable when OpenAI’s current version is preferred; the dated snapshot is more appropriate when an application needs a stable model reference for testing or reproducibility.
GPT-5.5 belongs at the high-capability end of OpenAI’s current model lineup. That positioning matters because the model is not primarily optimized for the cheapest short answers or the lowest possible latency. Its value is strongest when a task involves multiple steps, substantial context, difficult reasoning, code changes, external information, or several tool calls.
GPT-5.5 specifications at a glance
| Specification | GPT-5.5 |
|---|---|
| Provider | OpenAI |
| API model identifier | gpt-5.5 |
| Dated snapshot | gpt-5.5-2026-04-23 |
| Context window | Approximately 1,050,000 tokens |
| Maximum output | 128,000 tokens |
| Native input | Text and images |
| Native output | Text, including structured responses and tool calls |
| Reasoning effort | None, low, medium, high, and xhigh |
| Fine-tuning | Not supported |
| Standard input price | $5 per 1 million input tokens |
| Cached input price | $0.50 per 1 million tokens |
| Standard output price | $30 per 1 million output tokens |
These are model and API specifications supplied by OpenAI’s documentation. Statements about where GPT-5.5 performs best, how much it can reduce retries, or whether its additional capability is worth the price are practical evaluations rather than guarantees.
Input and output modalities
GPT-5.5 accepts text and images. Image input allows an application to provide visual material alongside written instructions, such as screenshots, diagrams, scanned pages, charts, or images of a software interface. The supplied model information does not list audio or video input as supported modalities.
The model produces text. That text can be an ordinary answer, source code, a structured machine-readable response, or a request to call a tool. GPT-5.5 does not natively generate images, audio, or video. A workflow may still use tools that access other modalities—for example, an image-generation tool or a computer-use environment—but the model’s own output remains text and tool instructions rather than a native media file.
This distinction is important when selecting GPT-5.5. It is suitable for an assistant that interprets an image and then explains or acts on it, but it is not a direct replacement for a dedicated image, speech, music, or video generation model.
Context window and output limits
GPT-5.5 has an approximately 1.05-million-token context window. A context window is the amount of information the model can consider in one request and its surrounding interaction, including instructions, conversation history, uploaded material, tool results, and generated content. This unusually large limit is useful for code repositories, long research collections, extensive specifications, and document sets that would otherwise need to be divided into many smaller requests.
The maximum output is 128,000 tokens. That does not mean every response should be that long. Large outputs consume more time and cost more, and a focused response is often easier to review. The limit is most useful for tasks such as generating substantial code changes, producing detailed reports, or completing long structured transformations.
Prompts exceeding 272,000 input tokens receive different pricing for the full session: twice the standard input price and 1.5 times the standard output price. Applications that routinely send very large contexts should account for this threshold rather than estimating cost from the standard rates alone.
Reasoning and coding capabilities
GPT-5.5 provides configurable reasoning effort levels: none, low, medium, high, and xhigh, with medium identified as the default. Reasoning effort controls how much deliberate processing the model applies before producing an answer. Lower settings can be appropriate for simpler or more latency-sensitive work, while higher settings are intended for problems that benefit from more extensive analysis. More reasoning is not automatically better for every request: it can increase latency and cost without improving a simple task.
OpenAI positions GPT-5.5 for high-complexity reasoning and professional coding. In practical terms, that includes understanding unfamiliar codebases, debugging failures, planning multi-file changes, refactoring, writing tests, analyzing technical requirements, and using tools to inspect or modify a working environment.
OpenAI reports improvements over GPT-5.4 on coding, computer-use, research, and knowledge-work evaluations. Those are provider-reported comparative results, not a guarantee that GPT-5.5 will outperform every alternative on every application. Real-world results still depend on the prompt, available tools, context quality, evaluation criteria, and the amount of human review.
Tools and agentic workflows
GPT-5.5 is designed for workflows in which the model does more than answer one question. It can plan a sequence, call a tool, inspect the returned information, revise its approach, and continue until it reaches a goal. This pattern is often called an agentic workflow, although the surrounding application remains responsible for permissions, validation, error handling, and safety controls.
In the Responses API, the supported tool ecosystem includes web search, file search, image generation tools, code interpreter, hosted shell, apply patch, Skills, computer use, MCP, and tool search. It also supports function calling, which lets an application expose its own operations—such as querying a database or creating a ticket—in a defined format.
Structured outputs are useful when the result must follow a specified schema rather than being free-form prose. For example, an application could request fields for a bug report, a research record, or an extraction result. Structured output improves consistency, but it does not guarantee that the underlying facts are correct; validation and, where necessary, human review are still needed.
Streaming allows partial output to be delivered as it is generated, which can make interactive applications feel more responsive. Batch processing is available for workloads that can be submitted for asynchronous processing rather than requiring an immediate answer. Prompt caching can reduce the cost of repeatedly sending reusable context.
GPT-5.5 pricing
GPT-5.5 costs $5 per 1 million input tokens and $30 per 1 million output tokens. Cached input is priced at $0.50 per 1 million tokens. Input tokens generally include the instructions, conversation material, documents, images or other request content represented for processing; output tokens are generated content and can include lengthy reasoning-related responses, code, structured data, or tool calls as applicable to the API workflow.
The output price is six times the standard input price, so applications should avoid requesting unnecessarily long answers. A design that retrieves only relevant documents, limits unneeded history, uses concise output schemas, and caches stable instructions can control costs more effectively than simply shortening user prompts.
Batch pricing is available through OpenAI’s Batch API. Prompts above 272,000 input tokens are charged at twice the standard input price and 1.5 times the standard output price for the full session. Regional processing endpoints carry a 10 percent uplift. These pricing rules make GPT-5.5 most economical when its stronger reasoning or reduced need for retries produces meaningful value, rather than for routine classification or short, repetitive responses.
Main strengths
- Long-context work: The approximately 1.05-million-token context window can accommodate large codebases, document collections, and extended tool results.
- Complex coding: The model is intended for debugging, refactoring, testing, code generation, and multi-step software engineering.
- Reasoning control: Multiple reasoning effort levels let applications trade depth against speed and cost.
- Tool integration: Web search, file search, code execution, computer use, shell access, MCP, and other tools support workflows that require actions or external information.
- Structured responses: Function calling and structured outputs make it easier to integrate results into software.
- Multimodal understanding: Image input allows visual information to be analyzed together with text.
- Production features: Streaming, prompt caching, and batch processing support different application architectures and workload patterns.
Limitations and trade-offs
GPT-5.5 is not a native image, audio, or video generation model. It also does not list audio or video input as supported. Applications requiring those capabilities may need separate specialized models or tools.
Fine-tuning is not supported. Teams that need a model customized through a formal fine-tuning process should consider an alternative that offers that capability, or use prompting, retrieval, structured outputs, and application-level controls where appropriate.
The model can produce inaccurate or overconfident answers. Web search and file search can improve grounding, but they do not eliminate incorrect interpretation, unreliable sources, or tool failures. Important legal, financial, medical, security, and operational decisions require appropriate verification.
GPT-5.5 can also be excessive for simple tasks. A smaller or lower-cost model may be more suitable for short classification, straightforward extraction, basic rewriting, high-volume automation, or applications where the fastest possible response matters more than extended reasoning. GPT-5.5’s higher output price is easier to justify when the task is difficult enough that better planning, coding, or tool use reduces retries and manual work.
Best use cases
- Maintaining and extending large software repositories
- Debugging, testing, and reviewing complex code changes
- Long-context research across many documents
- Technical analysis and professional report generation
- Spreadsheet and data-analysis workflows using code execution
- Grounded assistants that combine web search and file search
- Computer-use workflows involving browser or desktop interfaces
- Multi-step agents that call business tools and inspect their results
- Image-and-text analysis, such as interpreting diagrams or screenshots
- Large structured transformations where a high output limit is useful
When to choose GPT-5.5
Choose GPT-5.5 when the task is difficult, context-heavy, or tool-dependent and the cost of an incorrect or incomplete first attempt is significant. It is a strong candidate for an engineering agent that must inspect a repository, modify several files, run tests, and explain the result; a research assistant that must synthesize a large document set; or a professional workflow that combines reasoning, web access, file analysis, and structured output.
Choose a smaller or faster model when the task is predictable and narrow, such as labeling text, extracting a few fields, answering simple questions, or producing short routine summaries. The lower-cost option may deliver better economics at scale even if GPT-5.5 has higher peak capability.
Choose a specialized media model when the requirement is native image, audio, video, music, or speech generation. GPT-5.5 can participate in a broader tool-based workflow, but its native output is text. Choose an alternative with fine-tuning when changing model behavior through trained examples is a core requirement.
For applications that need reproducible behavior during evaluation or deployment, the dated snapshot gpt-5.5-2026-04-23 is preferable to relying only on the moving gpt-5.5 alias. For applications that want OpenAI’s current version without managing snapshot changes, the canonical alias is more convenient.
Bottom line
GPT-5.5 is best understood as a high-capability text model for complex reasoning, coding, long-context analysis, and tool-driven work. Its one-million-token-scale context, configurable reasoning, image input, structured outputs, and extensive tool support make it suitable for demanding professional workflows. The trade-offs are equally clear: it is relatively expensive, can be slower than simpler alternatives when using deeper reasoning, does not natively produce media, and still requires verification. Its strongest justification is not ordinary chat, but difficult work where sustained reasoning and reliable tool orchestration are worth the additional cost.

