Claude Sonnet

Claude Sonnet 5

by Claude · Current and generally available

Claude Sonnet 5 is Anthropic’s general-purpose model for coding, agentic automation, long-context analysis, tool use, and high-volume assistants. It accepts text and images, produces text, supports adaptive thinking and browser or computer-use workflows, and offers a 1-million-token context window with up to 128,000 output tokens. Standard pricing is $2 per million input tokens and $10 per million output tokens.

Text Actions Reasoning Coding
Claude Sonnet 5 is Anthropic’s fast, agent-focused model for software engineering, multi-step automation, document analysis, and other workloads that need a balance of capability, latency, and operating cost. It supports text and image input, text output, adaptive thinking, tool use, browser and computer-use workflows, and a 1-million-token context window. Its standard API price is $2 per million input tokens and $10 per million output tokens.
Outputs

What Claude Sonnet 5 can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

9/10 Reasoning
9/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Claude Sonnet
Model type General Purpose
Context window 1M tokens
Maximum output 128K tokens
Knowledge cutoff January 2026
Release date 2026-06-30
Status Current and generally available
Knowledge cutoff notes

Anthropic distinguishes the reliable knowledge cutoff from the broader training-data cutoff. The current model overview lists Claude Sonnet 5's reliable knowledge cutoff and training-data cutoff as January 2026.

Model notes

Canonical Claude API model ID is claude-sonnet-5. The model launched on June 30, 2026. It is available through the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. Adaptive thinking is enabled by default. Manual extended thinking with thinking.type enabled and budget_tokens is unsupported and returns a 400 error. Non-default temperature, top_p, and top_k values are rejected. Sonnet 5 uses a new tokenizer that produces approximately 30% more tokens for equivalent text than Sonnet 4.6. Prompt caching can provide up to 90% cost savings for eligible input, and batch processing provides a 50% discount. Current standard pricing is $2 per million input tokens and $10 per million output tokens. The model supports browser use and computer use on the Claude API and Google Cloud. Anthropic's model overview lists a reliable knowledge cutoff of January 2026 and indicates retirement no sooner than June 30, 2027; no exact shutdown date has been announced.

Cost

Model pricing

Input $2 per million input tokens
Output $10 per million output tokens
Model guide

Claude Sonnet 5: Anthropic’s Fast Model for Coding and AI Agents

Claude Sonnet 5 is Anthropic’s general-purpose model for coding, agentic workflows, tool use, long-context analysis, and high-volume applications. Released on June 30, 2026, it accepts text and images, returns text, supports adaptive thinking, offers a 1-million-token context window and up to 128,000 output tokens, and is priced at $2 per million input tokens and $10 per million output tokens.

What is Claude Sonnet 5?

Claude Sonnet 5 is a general-purpose multimodal language model provided by Anthropic. It is positioned as the current Sonnet model for applications that need substantial reasoning and coding ability without using Anthropic’s more expensive Opus tier. The model is designed particularly for software engineering, agentic workflows, knowledge work, tool-driven automation, and applications that process large amounts of context.

Anthropic released Claude Sonnet 5 on June 30, 2026. Its canonical Claude API model identifier is claude-sonnet-5. The model is generally available through the Claude API and is also offered through Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry, subject to each platform’s availability and integration terms.

The specifications and availability described here come from Anthropic’s supplied model documentation. Assessments such as whether Sonnet 5 is a good fit for a particular workload are practical evaluations based on those capabilities, not independent benchmark results or claims that Anthropic has published as formal scores.

Where Claude Sonnet 5 fits in Anthropic’s lineup

Sonnet 5 is intended to occupy the middle ground between maximum capability and maximum efficiency. Anthropic describes it as a combination of speed and intelligence for demanding everyday workloads. In practical terms, it is aimed at teams that need strong coding, reasoning, and agent performance at higher throughput and lower cost than higher-priced Opus models.

That positioning makes Sonnet 5 different from a small, efficiency-first model that is selected mainly for the lowest possible latency or price. It is also different from choosing the highest-capability model for every task regardless of cost. Sonnet 5 is most relevant when the system must complete multi-step work reliably enough to justify a capable model, while still handling substantial request volume.

Modalities and core capabilities

Claude Sonnet 5 accepts text and image input and produces text output. Image input enables visual question answering and analysis of images or supported visual document content, but the model does not natively generate images, audio, or video. Computer-use integrations can cause software interfaces to be manipulated through tool calls, yet the model’s own response remains primarily text-based.

Supported capabilities include vision, multilingual use, structured outputs, streaming, tool calling, prompt caching, batch processing, PDF handling, and the Files API. The model also supports server-side and client-side tools. Browser use and computer use are supported on the Claude API and Google Cloud according to the supplied documentation.

Tool use means that an application can give the model access to defined functions or external systems. For example, a software agent could ask Sonnet 5 to inspect repository files, call a business API, or interact with a browser through an approved tool. The external action is performed by the integration; Sonnet 5 decides what tool call to make and generates the accompanying text or structured arguments.

Context window and output limits

Claude Sonnet 5 has a 1-million-token context window. The supplied documentation identifies this as both the default and maximum context size, rather than offering a smaller context variant for this model. A token is a piece of text used by the model, so the limit covers the material supplied in the request together with the conversation and other context the application includes.

The synchronous Messages API supports up to 128,000 output tokens. This is a ceiling, not a requirement: most responses should use substantially fewer tokens. The large limit is useful for long code-generation tasks, extensive analyses, and agent workflows that need to return a sizeable result in one request.

Applications migrating from Sonnet 4.6 should review token budgeting carefully. Sonnet 5 uses a new tokenizer that can produce approximately 30% more tokens for equivalent input text, depending on the content. Even with a lower per-token price, a workload may not see a proportional reduction in its total bill if the same material is represented by more tokens.

Reasoning and response behavior

Sonnet 5 supports adaptive thinking, which is enabled by default. Adaptive thinking allows the model to vary how much internal reasoning it uses according to the task rather than requiring an application to assign a manual reasoning budget for every request. This is useful for workflows that mix short routine requests with more complex planning or debugging.

Manual extended thinking with thinking.type: "enabled" and a budget_tokens value is not supported and returns an error. Developers should therefore use the documented adaptive-thinking controls and effort setting instead of carrying forward configuration designed for an earlier model.

Applications should not assume that the first content block returned by the API is text. Adaptive thinking can produce thinking blocks before text blocks, so client code should inspect content-block types explicitly. This is an implementation detail with practical consequences: code that blindly reads the first block as a text response may fail or discard useful output.

Pricing, caching, and throughput

Standard Claude Sonnet 5 pricing is $2 per million input tokens and $10 per million output tokens. Input and output are billed separately, so applications that generate lengthy responses can incur substantially more output cost than short-answer workloads.

Prompt caching can reduce eligible input costs by up to 90%, according to the supplied research. Caching is most relevant when an application repeatedly sends the same large system instructions, reference material, or project context. Batch processing provides a 50% discount and is more appropriate when requests do not need immediate synchronous responses.

The cost trade-off is therefore workload-dependent. Sonnet 5 can be economical for high-volume applications when prompts are designed efficiently, repeated context is cached, or non-urgent work is batched. However, developers should calculate costs using the model’s actual tokenization and expected output length rather than relying only on the headline per-token prices.

Coding and agentic workloads

Coding is one of Sonnet 5’s clearest intended uses. The model is designed for software engineering agents, code generation, repository analysis, debugging, refactoring, and multi-step development workflows. Its long context can help an application provide substantial codebases, documentation, test output, and task history in one interaction.

For agentic work, the combination of adaptive thinking, tool calling, browser use, computer use, and long context is more important than any single feature. A coding agent might inspect a repository, plan a change, edit several files, run tests through a tool, interpret failures, and revise the implementation. Sonnet 5 can coordinate those steps, while the surrounding application remains responsible for permissions, execution, validation, and safety controls.

Computer-use support should not be interpreted as unrestricted autonomous access. The application must decide which tools are available and what actions are permitted. Production systems should review generated actions before allowing changes to files, accounts, infrastructure, or other consequential systems.

Best use cases for Claude Sonnet 5

  • Software engineering agents: repository analysis, implementation, debugging, code review, and test-driven workflows.
  • Long-context analysis: large document collections, lengthy specifications, source code, and visual PDFs.
  • Tool-driven automation: multi-step business processes that require function calls, browser interaction, or computer-use operations.
  • Research and knowledge work: structured investigation, synthesis, drafting, and analysis supported by external tools.
  • High-volume assistants: customer-facing or internal applications that need a capable model with comparatively moderate per-token pricing.
  • Structured automation: applications that need constrained machine-readable responses through Anthropic’s structured-output mechanism.

Limitations and migration caveats

Sonnet 5 is not a native image, audio, or video generation model. If the primary requirement is to create media rather than understand text or images, another specialized option is more appropriate. It also does not provide a separate legacy JSON mode merely because it supports structured outputs; applications should use the documented structured-output feature when they need machine-readable responses.

Several compatibility changes matter when migrating from Sonnet 4.6. Adaptive thinking is enabled by default, manual extended thinking is unsupported, and non-default values for temperature, top_p, or top_k are rejected. Assistant-message prefilling is also unsupported. Existing integrations should remove incompatible sampling settings, replace prefilling with system instructions or structured outputs, inspect content blocks by type, and retest output-length budgets.

Long context does not guarantee factual accuracy or correct tool actions. The model’s reliable knowledge cutoff is listed as January 2026, so current information should be supplied through approved retrieval or web-search tools and independently verified when accuracy matters. Anthropic’s documentation indicates retirement no sooner than June 30, 2027, but no exact shutdown date is supplied.

When to choose Claude Sonnet 5

Choose Claude Sonnet 5 when the workload needs a balance of reasoning, coding, tool use, context capacity, speed, and cost. It is a strong candidate for applications that must complete multi-step tasks, understand large inputs, or operate as a software or research agent, but that do not need the highest-priced model for every request.

Consider a different option when the task is extremely latency-sensitive and simple enough for a smaller model, when native image or video generation is required, or when the application depends on manual extended-thinking budgets, assistant-message prefilling, or unrestricted sampling parameters. A higher-capability Opus model may be more suitable for tasks where maximum reasoning performance outweighs throughput and price, while a more specialized model may be preferable for media generation or other non-text outputs.

For teams evaluating Sonnet 5, the most useful test is an end-to-end workload trial: measure task success, tool-call reliability, response latency, actual token counts, cache effectiveness, and the amount of human review required. Those results will show whether its intended speed-and-capability balance matches the application better than a smaller, cheaper model or a more capable, more expensive alternative.


Answers to Frequently Asked Questions

What should developers consider when migrating to Claude Sonnet 5?
Developers migrating from Sonnet 4.6 should account for adaptive thinking being enabled by default, the lack of support for manual extended thinking budgets, rejection of non-default temperature, top_p, and top_k values, and the removal of assistant-message prefilling. Client code should also inspect API content blocks by type because thinking blocks may appear before text blocks, and token usage may increase because Sonnet 5 uses a new tokenizer.
What is Claude Sonnet 5 best used for?
Claude Sonnet 5 is well suited to software engineering agents, repository analysis, code generation, debugging, refactoring, long-context document analysis, research, structured automation, and multi-step workflows that use tools, browser interaction, or computer-use integrations.
How much does Claude Sonnet 5 cost?
Standard pricing is $2 per million input tokens and $10 per million output tokens. Prompt caching can reduce eligible input costs by up to 90%, while batch processing provides a 50% discount for requests that do not require immediate responses. Actual costs depend on tokenization, output length, caching, and workload design.
What is Claude Sonnet 5?
Claude Sonnet 5 is Anthropic’s general-purpose multimodal language model for coding, reasoning, knowledge work, tool-driven automation, and AI agents. It is positioned between smaller efficiency-focused models and higher-priced Opus models, offering a balance of capability, speed, context capacity, and cost.
What are Claude Sonnet 5’s context window and output limits?
Claude Sonnet 5 supports a 1-million-token context window and up to 128,000 output tokens through the synchronous Messages API. The large context is useful for codebases, lengthy documents, research materials, and multi-step agent workflows.


Sources 7
Provider

About Claude