What is Claude Opus 4.8?
Claude Opus 4.8 is an Anthropic large language model designed for demanding reasoning and coding tasks. Its canonical API identifier is claude-opus-4-8. The model sits in Anthropic’s Opus tier, which is intended for workloads that benefit from deeper analysis and sustained task execution rather than simply the lowest cost or fastest response.
Anthropic released Claude Opus 4.8 on May 28, 2026. The supplied model documentation describes it as an active legacy model: it remains accessible, but Anthropic recommends migrating new deployments to newer Opus releases. The documentation lists no retirement date earlier than May 28, 2027.
In practical terms, Claude Opus 4.8 is aimed at work such as transforming a large codebase, analyzing a long collection of business or legal documents, conducting research with external tools, or operating as part of an agent that must plan and complete several steps. It is not an image, video, or audio generation model. Its output is text, including text that may contain code, structured data, explanations, or tool calls.
Core specifications and supported modalities
The following are specifications reported in the supplied Anthropic model documentation:
| Specification | Claude Opus 4.8 |
|---|---|
| Provider | Anthropic |
| Model ID | claude-opus-4-8 |
| Release date | May 28, 2026 |
| Status | Active legacy model |
| Context window | 1,000,000 tokens |
| Standard maximum output | 128,000 tokens |
| Batch maximum output | Up to 300,000 tokens in beta |
| Input | Text and images |
| Output | Text |
| Knowledge cutoff | January 2026 |
| Thinking | Adaptive, with high default effort |
A token is a unit used to measure text for model processing and billing. A 1-million-token context window allows an application to provide a very large collection of material in one request or across a long interaction, subject to the platform’s request and implementation limits. This is useful for code repositories, extensive document sets, long transcripts, and multi-stage tasks where losing earlier context would be a problem.
The model can inspect images as input, including visual information in supported documents and PDFs, but the supplied specifications do not describe native image creation. It also does not provide native audio or video input or output. Applications requiring those modalities need a different model or an additional processing service.
Reasoning, coding, and tool use
Claude Opus 4.8 supports adaptive thinking. This allows the application to use more or less internal reasoning effort depending on the task and configured settings. Higher effort can be useful for difficult planning, debugging, mathematical or analytical work, and tasks involving multiple dependent decisions. The trade-off is that deeper reasoning can consume more tokens and increase latency.
The model is particularly suited to software engineering that extends beyond generating an isolated function. Anthropic positions it for agentic coding, codebase-scale transformations, long-running development tasks, and workflows in which the model must inspect files, plan changes, call tools, and revise its work. Examples include proposing a coordinated change across many modules, explaining the impact of a migration, or helping an automated coding agent work through a multi-step issue.
Tool use lets an application give the model access to external functions or services. Claude Opus 4.8 can participate in tool-calling workflows, where it requests an operation such as retrieving information, examining a file, running a permitted action, or interacting with another system and then uses the returned result in its next response. The model documentation also references compatibility with web search, computer use, tool search, code execution, and browser-oriented tools. Exact availability depends on the interface and deployment, so a feature available in one Anthropic or cloud environment should not automatically be assumed to be available in every integration.
Structured outputs are supported. This is useful when an application needs the response to follow a defined data structure instead of receiving free-form prose. Structured output support should not be confused with native non-text generation: the model still produces text, even when that text is constrained to a machine-readable format.
Pricing and API economics
Anthropic lists Claude Opus 4.8 at $5 per million input tokens and $25 per million output tokens. Input tokens include material sent to the model, while output tokens include the response it generates. A request that uses substantial reasoning, produces long code, or repeatedly processes large context can therefore cost more than a short question-and-answer exchange.
Prompt caching is supported. The supplied pricing information lists cache-write rates of $6.25 per million tokens for five-minute storage and $10 per million tokens for one-hour storage. Cache reads are priced at $0.50 per million tokens. Caching can reduce the cost of repeatedly sending the same large instructions, documents, or project context, although the application must be designed to use cached content correctly.
The Batch API provides a 50% discount on input and output token pricing. Batch processing is better suited to jobs that do not require an immediate response, such as analyzing a large queue of documents or running an offline evaluation. Standard streaming responses are also supported for applications that want to display generated text as it becomes available.
The model supports long outputs of up to 128,000 tokens under the standard limit. Batch requests can support up to 300,000 output tokens in beta according to the supplied documentation. These are maximum documented limits, not a guarantee that every prompt will produce an output of that size. Long responses also increase processing time and cost.
Strengths and practical use cases
Claude Opus 4.8’s main advantage is the combination of high-end reasoning, coding, long context, vision input, and tool use in one model. That combination is valuable when a task is too complicated for a short prompt but does not justify splitting the work across many specialized systems.
- Large-scale software engineering: Reviewing or modifying a substantial codebase, planning changes across multiple files, and assisting with long-running coding agents.
- Enterprise knowledge work: Comparing policies, contracts, reports, or internal documentation while preserving relationships between distant passages.
- Research and analysis: Combining supplied documents with web search or other retrieval tools, then explaining the evidence and reasoning behind a conclusion.
- Vision-assisted document work: Interpreting images or visual PDF content alongside text.
- Agentic workflows: Planning several actions, calling external tools, checking returned results, and continuing until a larger task is complete.
- Long-form generation: Producing extensive technical documentation, analyses, or code when the application can justify the additional output cost.
Anthropic’s supplied release and product materials emphasize improved reliability, tool calling, long-horizon task completion, computer use, and fewer unsupported claims compared with earlier Opus models. These are provider claims rather than independent benchmark results. The supplied editorial assessment rates its reasoning and coding capability very highly, but those scores are comparative estimates and should not be treated as Anthropic-published measurements.
Limitations and trade-offs
The main limitation is economic and operational. Claude Opus 4.8 costs substantially more than a lower-tier model would for the same token volume, and its high-capability configuration may be slower. That makes it a poor default for simple classification, routine extraction, short summaries, or very high-volume generation when a less expensive model can meet the quality requirement.
The model’s knowledge cutoff is January 2026. It cannot be assumed to know events or changes after that date from its underlying training. Web search, retrieval, or another connected tool can provide newer information during a session, but external retrieval does not change the model’s knowledge cutoff.
Image understanding does not mean image generation. Claude Opus 4.8 can analyze supported visual inputs, but it does not natively create images, audio, or video. Applications built around those outputs need a different model or a separate generation service.
Feature availability also varies by deployment. Claude Opus 4.8 is available through the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud, and Microsoft Foundry, but platform-specific identifiers, tool integrations, quotas, and rollout status may differ. Developers should verify the target platform’s documentation before depending on a particular tool or output feature.
Finally, a large context window is not a substitute for good task design. Sending a million tokens can increase cost and may make it harder to identify the most important information. Applications should still organize documents, use retrieval where appropriate, and limit the context to material relevant to the current decision.
When to choose Claude Opus 4.8
Choose Claude Opus 4.8 when the work benefits from sustained reasoning, large context, advanced coding, image understanding, or multi-step tool use, and when the additional cost and latency are acceptable. It is a strong candidate for a coding agent that must understand a broad repository, an enterprise workflow involving many related documents, or a research system that combines analysis with external tools.
Choose a faster or less expensive model when the task is repetitive, short, predictable, or cost-sensitive. A smaller model may be more appropriate for high-volume summarization, straightforward extraction, simple customer-support responses, or applications where response time matters more than maximum reasoning depth. A dedicated image, audio, or video model is more appropriate when the required output is not text.
Claude Opus 4.8 is also less attractive for a new deployment if the application specifically benefits from the provider’s newest Opus generation. Anthropic classifies this model as legacy and recommends newer Opus releases for new projects. It can still make sense when its documented behavior, existing integration, pricing, or compatibility fits an established system, but teams should include migration planning in their evaluation.
Bottom line
Claude Opus 4.8 is a high-capability, text-output model built for difficult reasoning, agentic coding, long-context analysis, vision input, and tool-assisted workflows. Its 1-million-token context window, 128,000-token standard output limit, adaptive thinking, structured outputs, prompt caching, and batch support make it suitable for demanding applications. Its $5-per-million input and $25-per-million output pricing, relatively high resource requirements, January 2026 knowledge cutoff, and legacy status mean it should be selected for a specific capability need rather than used automatically for every request.

