Claude Opus

Claude Opus 5.5

by Claude · Active (latest)

Anthropic's Claude Opus 5.5 is an Opus-tier model for long-running agentic coding and knowledge work. It provides a 1-million-token context window, 128,000-token standard output limit, always-on adaptive thinking, text-and-image input, text output, tool support, structured outputs, prompt caching, batch processing, and standard pricing of $4 per million input tokens and $20 per million output tokens.

Text Reasoning Coding
Claude Opus 5.5 is Anthropic's latest Opus model, released on September 22, 2026. It is designed for extended software-engineering workflows, repository-scale analysis, agentic tasks, code review, and other knowledge work. Its defining practical advantages are a 1-million-token context window, 128,000-token maximum standard output, and adaptive thinking that is always enabled and controlled through an effort setting rather than a manually assigned thinking budget.
Outputs

What Claude Opus 5.5 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
7/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Claude Opus
Model type Reasoning
Context window 1M tokens
Maximum output 128K tokens
Knowledge cutoff June 2026
Release date 2026-09-22
Status Active (latest)
Shutdown date 2027-09-22
Knowledge cutoff notes

Anthropic lists both the reliable knowledge cutoff and training data cutoff as June 2026 for Claude Opus 5.5. Web search or external tool results can provide newer information during use but do not change the underlying cutoff.

Model notes

Canonical Claude API model ID is claude-opus-5-5. Amazon Bedrock uses anthropic.claude-opus-5-5; Google Cloud, Microsoft Foundry, and Claude Platform on AWS use platform-specific claude-opus-5-5 identifiers. Adaptive thinking is always on and cannot be disabled. The default effort is medium. Standard output is limited to 128K tokens; Message Batches API supports up to 300K output tokens in beta with the applicable beta header. Five-minute cache writes cost $5 per million tokens, one-hour cache writes cost $8 per million tokens, and cache reads cost $0.20 per million tokens. Batch input and output pricing is discounted by 50%. Fast mode is available only as a Claude API research preview and is priced separately. On the Claude API and Google Cloud, the older computer_20251124 tool is not supported; use computer_toolset_20260801 instead. Editorial scores are comparative estimates, not vendor-provided ratings.

Cost

Model pricing

Input $4 per million input tokens
Output $20 per million output tokens
Model guide

Claude Opus 5.5 for Long-Running Agentic Coding: Pricing, Limits, and Trade-offs

Claude Opus 5.5 is Anthropic's current Opus model for long-running agentic coding and knowledge work. It combines a 1-million-token context window, a 128,000-token standard output limit, always-on adaptive thinking, text-and-image input, tool use, structured outputs, prompt caching, batch processing, and standard pricing of $4 per million input tokens and $20 per million output tokens.

What Claude Opus 5.5 is

Claude Opus 5.5 is Anthropic's current Opus-tier model for tasks that require sustained reasoning across large amounts of information. It is aimed primarily at long-running software engineering and knowledge-work workflows rather than short, low-latency requests.

In practical terms, the model can work through large repositories, lengthy technical documents, screenshots, diagrams, charts, and other files while maintaining a large working context. It can also interact with tools, which allows an application to connect it to actions such as retrieving information, inspecting files, or carrying out steps in an agentic workflow.

The canonical Claude API model ID is claude-opus-5-5. The model is also available through Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, although those services may use platform-specific identifiers. Amazon Bedrock uses anthropic.claude-opus-5-5.

Where it fits in Anthropic's lineup

Claude Opus 5.5 sits in Anthropic's Opus tier, which is positioned for demanding reasoning, coding, and knowledge-work tasks. Anthropic classifies its latency as moderate relative to the current Claude lineup. That positioning makes it a middle ground between maximum capability and operational efficiency: it is intended for deeper, longer-running work, but it is more expensive and slower than Anthropic's Sonnet and Haiku tiers.

Compared with Claude Opus 5, Opus 5.5 has lower standard token pricing, but it also changes important implementation details. Thinking is always enabled, and applications cannot disable it or manually provide a fixed thinking-token budget. Tool-selection behavior also differs, so existing Opus 5 integrations may need updates rather than simply changing the model ID.

Key specifications and limits

SpecificationClaude Opus 5.5
ProviderAnthropic
Release dateSeptember 22, 2026
Model IDclaude-opus-5-5
Context window1,000,000 tokens
Standard maximum output128,000 tokens
Batch maximum outputUp to 300,000 tokens in beta with the applicable batch header
InputText and images
OutputText
ReasoningAlways-on adaptive thinking
Default effortMedium
StatusActive

The 1-million-token context window is especially relevant for repository-scale coding, large document collections, long conversations, and workflows that need to retain substantial intermediate material. The context limit is separate from the maximum output limit: the model can accept a very large working context, while a normal individual response remains limited to 128,000 tokens.

Anthropic documents a higher output ceiling of up to 300,000 tokens for the Message Batches API in beta when the applicable beta header is used. That batch limit should not be treated as the normal output allowance for standard interactive requests.

Input and output modalities

Claude Opus 5.5 accepts text and images and produces text. Its image input supports analysis of materials such as screenshots, diagrams, charts, and visual documents. The supplied specifications also identify PDF and file workflows as supported.

The model does not natively produce images, audio, or video. This distinction matters when selecting it for a multimodal application: it can inspect visual content, but it is not an image, speech, music, or video-generation model. Its multimodal capability is therefore primarily analytical rather than generative.

How adaptive thinking works

Claude Opus 5.5 uses adaptive thinking continuously. Rather than allowing an application to switch thinking off or assign a fixed thinking-token budget, Anthropic provides an effort parameter that controls the depth of the model's thinking. The documented default effort is medium. Adjusting effort can change the balance between reasoning depth, latency, and cost.

This design is useful for tasks where the model may need to spend different amounts of effort on different steps. A code review or complex repository change may justify deeper reasoning, while a simpler transformation may benefit from a lower effort setting. However, applications migrating from a model with explicit thinking controls must remove unsupported settings and recalibrate their effort configuration.

Thinking blocks are associated with the model and conversation. When a tool-using application passes thinking blocks between turns, it should preserve them unmodified and account for compatibility if it switches models during the same conversation. Text produced between tool calls may also appear in thinking blocks rather than ordinary text blocks, so user interfaces should not assume that all progress updates arrive as standard text.

Claude Opus 5.5 pricing

Standard Claude API pricing is:

  • Input: $4 per million tokens.
  • Output: $20 per million tokens.
  • Five-minute prompt-cache writes: $5 per million tokens.
  • One-hour prompt-cache writes: $8 per million tokens.
  • Prompt-cache reads: $0.20 per million tokens.

Batch processing receives a 50% discount on standard input and output pricing. That produces an effective batch price of $2 per million input tokens and $10 per million output tokens. Batch processing is therefore more attractive for workloads that do not require an immediate response.

The minimum cacheable prompt length is 512 tokens. Prompt caching can reduce the cost of repeatedly sending a large, stable prefix, such as repository instructions, reference documentation, or a persistent system prompt. Cache duration affects the write price: five-minute and one-hour cache writes have different rates, while cache reads are substantially cheaper than ordinary input.

Anthropic also offers Fast mode as a Claude API research preview. Fast mode is priced separately and is not available on the listed partner cloud platforms, so its pricing should not be assumed to match the standard rates above.

Coding, tools, and structured workflows

Claude Opus 5.5 is particularly suited to software engineering tasks that extend beyond generating a short code snippet. Appropriate examples include repository-scale code understanding, refactoring, debugging, code review, test-oriented development, and multi-step changes that require inspecting many files.

The model supports server-side and client-side tools, streaming, prompt caching, batch processing, and structured outputs. Structured outputs are useful when an application needs responses that follow a defined schema instead of relying on loosely formatted prose.

There is an important tool-use restriction: forced tool choices using tool_choice values of any or tool are not supported. Applications that previously depended on those settings should use automatic tool selection with strict tool use or use structured outputs when a schema-constrained result is required.

Computer-use integrations also need attention during migration. On the Claude API and Google Cloud, the earlier computer_20251124 computer-use tool is not accepted; those integrations must use the newer computer_toolset_20260801 toolset. Anthropic's model-specific documentation states that Amazon Bedrock continues to support the earlier computer-use tool.

Main strengths

  • Large working context: The 1-million-token window is suitable for large repositories, extensive document sets, and long-running conversations.
  • Long-form output: The 128,000-token standard output limit supports substantial code, analysis, and document-synthesis tasks.
  • Adaptive reasoning: Always-on thinking and effort controls allow applications to adjust reasoning depth without managing a fixed thinking budget.
  • Agentic workflows: Tool support, streaming, structured outputs, caching, and batch processing cover many components needed for multi-step applications.
  • Visual analysis: Text-and-image input makes the model useful for screenshots, diagrams, charts, PDFs, and other visual material alongside text.
  • Lower Opus pricing than its predecessor: Its standard $4 input and $20 output rates are lower than the supplied research's description of Claude Opus 5 pricing, although Opus 5.5 remains more expensive than lower Anthropic tiers.

Limitations and trade-offs

Claude Opus 5.5 is not the best fit for every workload. Its moderate latency and higher price make it less suitable for applications that prioritize the fastest possible response or the lowest cost per request. Sonnet and Haiku tiers may be more appropriate for those requirements, although the supplied research does not provide their specific prices or limits.

The model's output is text-only, so applications requiring native image, audio, video, music, or speech generation need a different model or an additional specialized service. Its always-on thinking can also increase latency or cost compared with a workflow that allows reasoning to be disabled.

Tool behavior is another compatibility limitation. Integrations that require forced tool selection, manually disabled thinking, or a fixed thinking budget need to be redesigned. Developers should also handle thinking blocks correctly, particularly when displaying progress or preserving tool-use turns.

Best use cases

  • Long-running agentic software development and repository-scale engineering.
  • Complex code generation, refactoring, debugging, and code review.
  • Research, analysis, document synthesis, and other knowledge-work tasks involving large source material.
  • Multimodal analysis of screenshots, diagrams, charts, PDFs, and technical documents.
  • Tool-using applications that need adaptive reasoning and schema-constrained responses.
  • Asynchronous or high-volume processing where the Batch API discount can reduce cost.
  • Repeated workflows with large stable prompts that can benefit from prompt caching.

When to choose Claude Opus 5.5

Choose Claude Opus 5.5 when the task benefits from a large context, sustained reasoning, and the ability to combine text or image analysis with tools. It is a strong candidate for an agent that must inspect a substantial codebase, maintain context over many steps, review technical evidence, or produce a lengthy structured result.

Choose a faster or less expensive model tier when the task is routine, latency-sensitive, or performed at very high volume and does not need Opus-level reasoning or a million-token context. Choose a specialized generation model when the required output is an image, audio, video, music, or speech rather than text.

For existing Claude Opus 5 applications, migration is most appropriate when the lower standard pricing, larger context, or updated capabilities justify integration changes. Before switching, remove unsupported thinking controls, replace forced tool selection, review computer-use tool requirements, and test how the application handles thinking blocks.

Bottom line

Claude Opus 5.5 is designed for demanding, extended workflows rather than simple chat or lowest-cost inference. Its combination of a 1-million-token context window, 128,000-token standard output, always-on adaptive thinking, multimodal input, and tool support makes it especially relevant to agentic coding and large-scale knowledge work. The trade-off is higher cost and moderate latency, along with migration requirements for applications that depend on manually controlled thinking or forced tool selection.


Answers to Frequently Asked Questions

What are the main limitations of Claude Opus 5.5?
Claude Opus 5.5 has moderate latency and is more expensive than lower Anthropic tiers. Its thinking cannot be disabled or assigned a fixed token budget, its output is text-only, and existing integrations may require changes to thinking controls, forced tool selection, computer-use tools, and thinking-block handling.
Does Claude Opus 5.5 support adaptive thinking and tool use?
Yes. Claude Opus 5.5 uses always-on adaptive thinking, with an effort parameter controlling reasoning depth; the default effort is medium. It supports server-side and client-side tools, streaming, prompt caching, batch processing, and structured outputs, but it does not support forced tool choices using tool_choice values of any or tool.
What are Claude Opus 5.5's context window and output limits?
Claude Opus 5.5 has a 1-million-token context window and a standard maximum output of 128,000 tokens. The Message Batches API supports up to 300,000 output tokens in beta when the applicable batch header is used.
What is Claude Opus 5.5 best used for?
Claude Opus 5.5 is best suited to long-running agentic software development, repository-scale code understanding, refactoring, debugging, code review, research, document synthesis, and multimodal analysis involving screenshots, diagrams, charts, and PDFs.
How much does Claude Opus 5.5 cost?
Standard Claude API pricing is $4 per million input tokens and $20 per million output tokens. Five-minute prompt-cache writes cost $5 per million tokens, one-hour cache writes cost $8 per million tokens, and cache reads cost $0.20 per million tokens. Batch processing applies a 50% discount to standard input and output rates.


Sources 8
Provider

About Claude