What Claude Fable 5.1 is
Claude Fable 5.1 is an active Anthropic model designed for tasks where sustained reasoning quality matters more than minimum latency or cost. Anthropic positions it for demanding knowledge work, including long-running coding agents, multistep research, complex document analysis, and work involving spreadsheets or presentations.
The model is the high-capability option in the Claude Fable 5 family. Anthropic recommends it when a high-effort use of Claude Opus 5.5 does not achieve the required quality level. That positioning is important: Fable 5.1 is not intended to be the default choice for every prompt. It is aimed at difficult tasks that benefit from extended context, careful reasoning, and repeated interaction with tools or source material.
Claude Fable 5.1 was released on September 1, 2026, and is listed as active and generally available. Anthropic states that retirement will not occur sooner than September 1, 2027.
Core specifications and context limits
The model has a 1,000,000-token context window and can generate up to 128,000 tokens in a single response. A context window is the amount of conversation, source material, tool output, and other information the model can consider during a request. In practical terms, the large limit is useful for keeping substantial codebases, long research collections, or large business documents available within one workflow.
| Specification | Claude Fable 5.1 |
|---|---|
| Provider | Anthropic |
| Model ID | claude-fable-5-1 |
| Release date | September 1, 2026 |
| Status | Active |
| Context window | 1,000,000 tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | June 2026 |
| Input | Text and images |
| Output | Text, including supported structured formats |
The June 2026 knowledge cutoff is separate from the model’s ability to use current information through tools. Anthropic’s web-search tool can retrieve newer information during a request, but web access does not update the model’s underlying training-data cutoff.
Reasoning and coding capabilities
Claude Fable 5.1 uses always-on adaptive thinking. This means the model can spend additional effort working through a difficult request rather than treating every prompt as a short, direct generation task. Its documented default effort level is high, with configurable levels of low, medium, high, xhigh, and max. These settings allow developers to trade off reasoning effort against response time and cost.
For coding, the model is intended for long-running agents and complex software engineering rather than only short code completions. A coding agent can use it to inspect a large project, reason across multiple files, plan a change, use tools, and revise its approach as new results arrive. The large context window is particularly relevant when the task depends on retaining extensive code, requirements, test output, and documentation at the same time.
These capabilities do not mean every coding task requires Fable 5.1. A smaller or faster model may be more efficient for straightforward code generation, routine transformations, or high-volume requests. Fable 5.1 makes the most sense when errors are expensive, the task has many dependencies, or the work extends across multiple reasoning and tool-use steps.
Supported modalities and output types
Claude Fable 5.1 accepts text and image inputs and produces text output. Image input allows the model to analyze visual material alongside written instructions, which can help with documents, diagrams, screenshots, or presentation-related workflows when those inputs are supplied in supported requests.
The model does not natively generate images, audio, video, music, embeddings, or executable actions. Its multimodal capability is therefore input-focused: it can understand text and images, but its documented output modality remains text. Structured text such as JSON is available through constrained output features, but that should not be confused with direct image, audio, or video generation.
Tools, structured outputs, and API features
Claude Fable 5.1 supports tool use, streaming responses, prompt caching, structured outputs, server-side web search, and asynchronous batch processing. Tool use allows an application to provide callable functions or services that the model can request during a workflow. Web search is available through Anthropic’s first-party web-search tool.
Structured outputs can constrain responses to machine-readable formats such as JSON Schema. This is useful when an application needs predictable fields rather than free-form prose. The supplied documentation verifies structured outputs, but does not independently verify a separate legacy JSON-mode capability.
There is an important migration detail for developers moving from Claude Fable 5. Forced tool use with tool_choice set to any or tool returns an error on Claude Fable 5.1. Anthropic instead advises using automatic tool choice with strict tool schemas when schema-conformant tool calls are required. Applications should also account for thinking-block compatibility when passing conversations between model versions: earlier models may not be able to consume this model’s thinking blocks, and editing earlier turns can invalidate them.
Streaming allows partial output to be received as it is generated, which can improve the perceived responsiveness of a long answer even though Fable 5.1 is comparatively slow. Batch processing is intended for asynchronous workloads and receives a 50% discount from the standard input and output prices.
Pricing and prompt caching
Standard Claude API pricing is $10 per million input tokens and $50 per million output tokens. Output is priced substantially higher than input, so applications should avoid requesting unnecessarily long responses and should use concise prompts or staged workflows where appropriate.
| Usage type | Price |
|---|---|
| Standard input | $10 per million tokens |
| Standard output | $50 per million tokens |
| Five-minute prompt-cache write | $12.50 per million tokens |
| One-hour prompt-cache write | $20 per million tokens |
| Prompt-cache read | $0.25 per million tokens |
| Batch input | $5 per million tokens |
| Batch output | $25 per million tokens |
Prompt caching is especially relevant for agents that repeatedly send the same instructions, project files, or reference material. Cache writes cost more than ordinary input tokens, but cache reads are much cheaper. Anthropic states that caching can reduce typical workload costs by approximately 25% and highly agentic workload costs by up to approximately 45%, depending on how effectively the application reuses cached context. These are provider-reported estimates rather than guaranteed savings for every workload.
Availability
Claude Fable 5.1 is available through the Claude API and through several cloud distribution channels, including Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Eligible Claude Pro, Max, Team, and Enterprise users can also access it through the relevant Claude products.
Availability, quotas, and commercial terms can vary by platform. Developers choosing a cloud marketplace should confirm the applicable regional availability, billing arrangement, request limits, and feature support before deploying the model.
Main strengths and limitations
Where it is strongest
- Long-context work: The 1-million-token context window can accommodate unusually large collections of source material, code, or conversation history.
- Extended reasoning: Adaptive thinking and configurable effort levels support tasks that need more deliberate analysis than a short-answer model.
- Agentic coding: The model is designed for long-running software-engineering workflows involving multiple files, tools, tests, and revisions.
- Complex knowledge work: Research, document analysis, spreadsheet work, and presentation tasks can benefit from retaining substantial context.
- Application integration: Tool use, structured outputs, streaming, caching, web search, and batch processing cover common production workflow needs.
Where it is weaker or less suitable
- Speed: Anthropic’s comparison describes Fable 5.1 as slower than its Opus, Sonnet, and Haiku alternatives.
- Cost: At $50 per million output tokens, large responses can become expensive, especially when high effort is used repeatedly.
- Modality limits: It analyzes text and images but does not natively produce images, audio, video, music, embeddings, or executable actions.
- Tool-choice migration: Applications relying on forced
anyortoolselection need to change their tool-calling approach. - Knowledge cutoff: Information after June 2026 requires retrieval or another current-information source.
When to choose Claude Fable 5.1
Choose Claude Fable 5.1 when a task is difficult enough that additional reasoning quality, long context, or sustained tool use justifies higher latency and cost. Good examples include an agent that must work through a large software repository, a research process that combines many documents and search steps, or an analysis that requires keeping extensive source material available while producing a detailed result.
It is also a reasonable choice when a workflow needs both a large context window and structured application output. For example, an application could provide a substantial set of documents, ask the model to analyze them, and require the result to follow a JSON Schema for downstream processing.
A faster or less expensive model is likely more appropriate for simple questions, routine classification, short summaries, high-volume extraction, or low-latency interactive features. Fable 5.1’s advantages are less valuable when the task has a small context, requires little reasoning, or can be completed reliably without extended tool use. A model with native image, audio, or video generation is also more appropriate when producing those media types is the primary requirement.
Bottom line
Claude Fable 5.1 is Anthropic’s high-capability option for demanding, long-running work. Its defining practical advantages are the 1-million-token context window, 128,000-token output ceiling, adaptive thinking, and support for tools, caching, web search, structured outputs, streaming, and batch processing. The trade-off is clear: it is slower and more expensive than smaller alternatives, and it produces text rather than generated media. For complex coding, research, and document workflows where quality and sustained reasoning matter more than speed or minimum cost, those trade-offs may be justified.

