Claude 4.5

Claude Haiku 4.5

by Claude · Active (latest); retirement not sooner than October 15, 2026

Claude Haiku 4.5 is Anthropic's fastest Claude 4.5 model for low-latency assistants, coding support, image understanding, customer-service agents, and high-volume processing. It supports text and image input, text output, extended thinking, tools, structured outputs, prompt caching, and batch processing.

Text Reasoning Coding
Claude Haiku 4.5 is Anthropic's lightweight Claude 4.5 model for applications where response speed and inference cost matter alongside strong reasoning and coding. Released on October 15, 2025, it accepts text and images, produces text, and is available through the Claude API and supported cloud platforms. Its $1-per-million input-token and $5-per-million output-token pricing makes it suited to interactive products and large volumes of automated work.
Outputs

What Claude Haiku 4.5 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
10/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Claude 4.5
Model type Lightweight
Context window 200K tokens
Maximum output 64K tokens
Knowledge cutoff February 2025
Release date 2025-10-15
Status Active (latest); retirement not sooner than October 15, 2026
Knowledge cutoff notes

Anthropic distinguishes the reliable knowledge cutoff of February 2025 from the broader training data cutoff of July 2025. Web search or external tools can provide newer information during use but do not change the underlying knowledge cutoff.

Model notes

The canonical Claude API snapshot is claude-haiku-4-5-20251001. The claude-haiku-4-5 identifier is a convenience alias resolving to that fixed snapshot. Anthropic lists text and image input with text output, manual extended thinking, a 200K context window, and up to 64K output tokens. Structured outputs are supported, but a separate legacy JSON-mode capability was not independently verified. The model is active as of September 24, 2026, with retirement not sooner than October 15, 2026; Anthropic has not published an exact shutdown date. Training data cutoff is July 2025, while the reliable knowledge cutoff is February 2025. Editorial scores are comparative estimates, not vendor benchmarks.

Cost

Model pricing

Input $1 per million input tokens; cache write $1.25 per million tokens for 5 minutes or $2 per million tokens for 1 hour; cache read $0.10 per million tokens
Output $5 per million output tokens; batch API output pricing is 50% lower
Model guide

Claude Haiku 4.5: Near-Frontier Performance for Fast, High-Volume AI

Claude Haiku 4.5 is Anthropic's fastest Claude 4.5 model, designed for low-latency assistants, coding support, customer-service agents, image understanding, and high-volume processing. It combines text and image input, text output, a 200,000-token context window, up to 64,000 output tokens, manual extended thinking, tool use, structured outputs, prompt caching, and relatively low API pricing.

What is Claude Haiku 4.5?

Claude Haiku 4.5 is Anthropic's fastest model in the Claude 4.5 family. The model is intended to handle tasks that need more than simple text generation but still require quick responses and controlled operating costs. Anthropic positions it as a near-frontier model for interactive applications, coding assistance, customer-service agents, and high-volume processing.

It was released on October 15, 2025. The model is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Its canonical Claude API snapshot is claude-haiku-4-5-20251001. Developers can also use the convenience alias claude-haiku-4-5, which resolves to that fixed snapshot.

In practical terms, Haiku 4.5 is a speed-and-cost-oriented choice within Anthropic's lineup. It is not designed to generate images, audio, or video. Its multimodal capability refers to what it can receive: text and images. Its responses are text, including generated code, structured JSON, explanations, classifications, and tool-related content.

Capabilities, context, and output limits

Claude Haiku 4.5 has a 200,000-token context window and supports up to 64,000 output tokens. A context window is the amount of text, code, images, and conversation material the model can consider during a request. This gives the model room to process long documents, sizeable codebases, detailed instructions, and extended conversations, although actual usable capacity can depend on the request format and platform.

The model supports manual extended thinking through the Messages API. Extended thinking gives the model additional internal reasoning space before it produces an answer. It can be useful for multi-step coding, analysis, planning, and difficult decisions, but applications should account for the additional latency and output-token usage that deeper reasoning can introduce.

Anthropic's documentation also lists tool use, streaming responses, prompt caching, batch processing, and structured outputs. Streaming allows an application to display a response as it is generated rather than waiting for the complete answer. Tool use allows the model to request actions from application-defined tools, while structured outputs can constrain responses to a specified JSON schema.

Structured outputs should not automatically be treated as the same thing as a separate legacy JSON mode. The supplied documentation verifies schema-constrained JSON and strict tool-use support for Haiku 4.5, but does not independently verify a distinct legacy JSON-mode capability.

Input and output modalities

Haiku 4.5 accepts text and image input and returns text output. Image understanding can support tasks such as examining visual material supplied with a prompt, interpreting diagrams, or analyzing supported visual documents. The model does not natively generate images, video, audio, music, or speech.

This distinction matters when selecting the model. An application that needs to read an image and produce a written explanation can use Haiku 4.5. An application that must create an image or return spoken audio needs a separate generation or speech system.

CapabilityClaude Haiku 4.5
Text inputSupported
Image inputSupported
Text outputSupported
Image, video, audio, or speech outputNot supported
Context window200,000 tokens
Maximum output64,000 tokens

Pricing and cost control

Standard Claude API pricing is $1 per million input tokens and $5 per million output tokens. Input tokens are the material sent to the model, while output tokens are the generated response. The difference between the two prices makes response length an important cost consideration, particularly when extended thinking or very long generated documents are enabled.

Prompt caching can reduce the cost of repeatedly sending the same context. Haiku 4.5 cache writes cost $1.25 per million tokens for a five-minute cache and $2 per million tokens for a one-hour cache. Cache reads cost $0.10 per million tokens. These prices are separate from ordinary input processing, so caching is most relevant when an application reuses large system instructions, reference documents, or other stable context.

The Message Batches API applies a 50% discount to standard input and output pricing. Batch processing is better suited to work that does not need an immediate response, such as bulk classification, document transformation, or offline evaluation. Interactive requests generally need the standard API path and its normal pricing.

Reasoning and coding performance

Anthropic describes Haiku 4.5 as offering near-frontier reasoning and coding performance while remaining the fastest Claude 4.5 model. The provider's positioning makes it a candidate for tasks such as debugging, code generation, test writing, repository questions, structured extraction, and multi-step analysis.

Its manual extended-thinking support is particularly relevant when a fast model needs to work through a more complicated problem. However, Haiku 4.5 should not be assumed to deliver the maximum reasoning quality available from every larger or more specialized model. The supplied research supports a near-frontier positioning, not a universal claim that it is the best model for every difficult reasoning benchmark or domain.

For coding assistants, the model's advantages are response speed, low token pricing, long context, and tool support. These characteristics can help with interactive pair programming, code review, error explanation, and parallel coding tasks. Applications should still validate generated code, especially when the model is connected to tools that can change files, call services, or affect production systems.

Best use cases

  • Interactive assistants: Fast responses make Haiku 4.5 appropriate for chat interfaces where users are waiting for each turn.
  • Customer-service agents: The model can classify requests, draft replies, summarize conversations, and use application tools to retrieve relevant information.
  • High-volume text processing: Its relatively low input and output prices are useful for classification, extraction, rewriting, and document triage across large datasets.
  • Coding assistance: Developers can use it for pair programming, code explanation, test generation, and debugging.
  • Image understanding: It can combine image input with text instructions when an application needs written analysis of visual material.
  • Parallel subagents: A larger model can delegate smaller research, coding, classification, or transformation tasks to multiple Haiku 4.5 instances.

When should you choose Claude Haiku 4.5?

Choose Claude Haiku 4.5 when latency, throughput, and cost are important but the task still benefits from strong reasoning, coding, image understanding, and tool use. It is a good fit for applications that generate many responses, serve users interactively, or need to process long inputs without paying the price of a larger model for every request.

It is also a practical option when the workload has mixed difficulty. Routine requests can use the fast model directly, while a larger model can handle only the cases that need deeper reasoning. This type of routing can reduce average cost without forcing every task onto the most expensive model.

Another model may be more appropriate when the application needs maximum frontier reasoning quality, a larger context window, or native image, video, audio, or speech generation. Haiku 4.5 is also not the best choice for unrestricted usage assumptions: account limits, rate limits, platform availability, and conversation size can affect the experience. The supplied research specifically identifies workloads requiring a 1-million-token context window and native audio output as cases where another option should be considered.

Implementation considerations

Applications migrating from earlier Haiku models should use the current Claude 4.5 model identifier and review Anthropic's migration guidance. Haiku 4.5 uses manual extended thinking rather than adaptive thinking. Developers should also use only one of temperature or top_p, update legacy tool versions where applicable, handle the refusal stop reason, and review rate limits before deploying a workload at scale.

Tool-enabled applications should treat model-generated tool requests as untrusted instructions that require application-side validation. Structured outputs can make downstream parsing more predictable, but schema validation does not remove the need to check values, permissions, and business rules.

Anthropic lists Haiku 4.5 as active, with retirement not sooner than October 15, 2026. This is a minimum retirement commitment rather than a confirmed shutdown date. The model's reliable knowledge cutoff is February 2025, while its broader training data cutoff is July 2025. Web search or external tools can provide newer information during use, but they do not change the underlying knowledge cutoff.

Bottom line

Claude Haiku 4.5 is built for fast, economical inference without limiting the model to simple text completion. Its strongest practical combination is a 200,000-token context window, 64,000-token maximum output, text and image input, manual extended thinking, coding capability, tool use, and low standard API pricing. It is most compelling for interactive and high-volume workloads that need capable text reasoning but do not require media generation or the maximum possible frontier-model performance.


Answers to Frequently Asked Questions

When should you choose Claude Haiku 4.5 instead of a larger model?
Choose Claude Haiku 4.5 when low latency, high throughput, and controlled costs matter while the task still requires strong reasoning, coding, image understanding, or tool use. A larger model may be preferable for maximum frontier reasoning quality, a 1-million-token context window, or native media generation.
Does Claude Haiku 4.5 support image generation, audio, or video?
No. Claude Haiku 4.5 can accept image input and provide text-based analysis, but it does not natively generate images, video, audio, music, or speech.
How much does Claude Haiku 4.5 cost?
Standard Claude API pricing is $1 per million input tokens and $5 per million output tokens. Prompt caching and the Message Batches API can reduce costs for repeated or non-urgent workloads.
What is Claude Haiku 4.5 best used for?
Claude Haiku 4.5 is best suited to fast, high-volume workloads such as interactive assistants, customer-service agents, coding support, document classification, structured extraction, image understanding, and parallel subagent tasks.
What are the context window and output limits of Claude Haiku 4.5?
Claude Haiku 4.5 has a 200,000-token context window and supports up to 64,000 output tokens. It also supports text and image input, while its output is text.


Sources 9
Provider

About Claude