Developer platform

Claude API

Developer overview for Claude, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Messages API
SDK support Official SDKs for Python, TypeScript, C#, Go, Java, PHP, and Ruby; official ant CLI; direct REST and cURL support.
Rate limits Organization-level usage tiers enforce monthly spend caps and per-model RPM, input-token-per-minute, and output-token-per-minute limits. Message Batches, Files API, web search, and Managed Agents have additional limits. 429 responses include retry-after w
Platform

API overview

Endpoints

API access

Base URL https://api.anthropic.com
Primary API Messages API
Pricing

API pricing

Pricing model Usage-based pay-as-you-go token pricing, with separate pricing for some tools, caching, and service tiers

Pricing varies by model and is billed per input and output token. Batch processing is generally 50% lower than standard pricing; prompt caching and web search have separate pricing rules.

Developer experience

SDKs & usability

SDKs Official SDKs for Python, TypeScript, C#, Go, Java, PHP, and Ruby; official ant CLI; direct REST and cURL support.
Ease of use High for standard Messages API integrations; official SDKs, Workbench, typed responses, streaming helpers, retries, and extensive examples reduce setup effort.
Documentation Strong and detailed official documentation with API reference pages, language-specific SDK guides, feature guides, migration notes, and operational documentation.
Latency Latency varies by model, workload, context size, service tier, region, and queue conditions. Standard is best-effort, batch is asynchronous, and qualifying fast-mode or priority options are model- and contract-dependent. Streaming reduces time to first vi
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

The recommended direct integration is POST /v1/messages using the Claude API. The API supports stateless multi-turn messages, system instructions, streaming via server-sent events, client tool use, server tools including web search and web fetch, structured JSON outputs on supported models, image input, PDF and document processing, prompt caching, message batches, token counting, and the Files API. Claude Managed Agents provides reusable versioned agents, sessions, environments, tools, skills, and MCP integrations but is currently beta and uses a beta header. Feature availability varies by model and platform. Direct Anthropic API access is distinct from cloud-hosted access through Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS.

Examples

API examples

curl https://api.anthropic.com/v1/messages \
  --header "Authorization: Bearer ${ANTHROPIC_API_KEY}" \
  --header "anthropic-version: 2023-06-01" \
  --header "content-type: application/json" \
  --data '{
    "model": "claude-sonnet-5",
    "max_tokens": 1024,
    "system": "You are a concise technical assistant.",
    "messages": [
      {
        "role": "user",
        "content": "Explain how HTTP retries should handle a 429 response."
      }
    ]
  }'

curl https://api.anthropic.com/v1/messages \
  --header "Authorization: Bearer ${ANTHROPIC_API_KEY}" \
  --header "anthropic-version: 2023-06-01" \
  --header "content-type: application/json" \
  --data '{
    "model": "claude-sonnet-5",
    "max_tokens": 1024,
    "stream": true,
    "output_config": {
      "format": {
        "type": "json_schema",
        "schema": {
          "type": "object",
          "properties": {
            "answer": {"type": "string"},
            "priority": {"type": "string", "enum": ["low", "medium", "high"]}
          },
          "required": ["answer", "priority"],
          "additionalProperties": false
        }
      }
    },
    "messages": [
      {
        "role": "user",
        "content": "Classify this incident: the API is returning intermittent 429 errors."
      }
    ]
  }' | while IFS= read -r line; do
    case "$line" in
      data:\ {*) printf '%s\n' "$line" ;;
    esac
done
# Install: pip install -U anthropic
import json
import os
import anthropic

client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])

try:
    message = client.messages.create(
        model="claude-sonnet-5",
        max_tokens=1024,
        system="You are a concise technical assistant.",
        messages=[
            {"role": "user", "content": "What is an idempotent HTTP request?"},
            {"role": "assistant", "content": "An idempotent request has the same intended effect when repeated."},
            {"role": "user", "content": "Give two API design examples."},
        ],
    )
    text = "\n".join(block.text for block in message.content if block.type == "text")
    print(text)
except anthropic.RateLimitError as exc:
    print(f"Rate limited: {exc}")
except anthropic.APIError as exc:
    print(f"Anthropic API error: {exc}")

print("Request ID:", getattr(message, "_request_id", None))

# Streaming example
with client.messages.stream(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Summarize the benefits of streaming responses."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    print()

# Structured output example
structured = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=512,
    output_config={
        "format": {
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "summary": {"type": "string"},
                    "severity": {"type": "string", "enum": ["low", "medium", "high"]},
                },
                "required": ["summary", "severity"],
                "additionalProperties": False,
            },
        }
    },
    messages=[{"role": "user", "content": "Classify an API outage that affects 10% of requests."}],
)
json_text = next(block.text for block in structured.content if block.type == "text")
print(json.loads(json_text))
// Install: npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: process.env.ANTHROPIC_API_KEY,
});

try {
  const message = await client.messages.create({
    model: "claude-sonnet-5",
    max_tokens: 1024,
    system: "You are a concise technical assistant.",
    messages: [
      { role: "user", content: "Explain exponential backoff for API clients." },
      { role: "assistant", content: "Exponential backoff increases the delay between retries." },
      { role: "user", content: "Show a simple retry sequence." },
    ],
  });

  const text = message.content
    .filter((block) => block.type === "text")
    .map((block) => block.text)
    .join("\n");
  console.log(text);
} catch (error) {
  if (error instanceof Anthropic.RateLimitError) {
    console.error("Rate limited; retry after the server-provided delay.");
  } else if (error instanceof Anthropic.APIError) {
    console.error(`Anthropic API error: ${error.status} ${error.message}`);
  } else {
    throw error;
  }
}

// Streaming example
const stream = client.messages.stream({
  model: "claude-sonnet-5",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Describe server-sent events in one paragraph." }],
});

stream.on("text", (text) => process.stdout.write(text));
await stream.finalMessage();
process.stdout.write("\n");

// Tool-use example
const toolResponse = await client.messages.create({
  model: "claude-sonnet-5",
  max_tokens: 1024,
  tools: [
    {
      name: "get_weather",
      description: "Get the current weather for a city.",
      input_schema: {
        type: "object",
        properties: { city: { type: "string" } },
        required: ["city"],
      },
    },
  ],
  messages: [{ role: "user", content: "What is the weather in Boston?" }],
});

for (const block of toolResponse.content) {
  if (block.type === "tool_use") {
    console.log(JSON.stringify({ tool: block.name, input: block.input }));
  }
}
<?php

$apiKey = getenv('ANTHROPIC_API_KEY');
$url = 'https://api.anthropic.com/v1/messages';
$payload = [
    'model' => 'claude-sonnet-5',
    'max_tokens' => 1024,
    'system' => 'You are a concise technical assistant.',
    'messages' => [
        [
            'role' => 'user',
            'content' => 'Explain how to handle transient API failures.'
        ]
    ]
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'anthropic-version: 2023-06-01',
        'Content-Type: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 60,
]);

$raw = curl_exec($ch);
$curlError = curl_error($ch);
$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

if ($raw === false || $curlError !== '') {
    throw new RuntimeException('Network error: ' . $curlError);
}

$data = json_decode($raw, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? 'Unknown Anthropic API error';
    throw new RuntimeException("HTTP {$status}: {$message}");
}

$text = '';
foreach ($data['content'] ?? [] as $block) {
    if (($block['type'] ?? null) === 'text') {
        $text .= $block['text'];
    }
}

echo $text . PHP_EOL;

// Tool-use request example
$toolPayload = [
    'model' => 'claude-sonnet-5',
    'max_tokens' => 1024,
    'tools' => [[
        'name' => 'get_weather',
        'description' => 'Get the current weather for a city.',
        'input_schema' => [
            'type' => 'object',
            'properties' => ['city' => ['type' => 'string']],
            'required' => ['city']
        ]
    ]],
    'messages' => [[
        'role' => 'user',
        'content' => 'What is the weather in Boston?'
    ]]
];

// Send $toolPayload to the same endpoint, inspect content blocks with type tool_use,
// execute the named function in your application, then send a tool_result message.
Policies

Data & usage

Data training

For Anthropic commercial offerings, including the Anthropic API, inputs and outputs are not used to train models by default. Anthropic may use material when a customer explicitly opts in, submits feedback or bug reports, or participates in an applicable development or improvement program. Commercial customers retain rights in their inputs and outputs under the applicable commercial terms.

Data retention

Anthropic states that standard Anthropic API inputs and outputs are automatically deleted from its backend within 30 days of receipt or generation, subject to legal, Usage Policy, security, and contractual exceptions. Eligible customers may arrange Zero Data Retention, but ZDR eligibility differs by feature. Files uploaded through the Files API persist until deletion or optional expiration, are workspace-scoped, and have separate lifecycle behavior.

Rate limits

Organization-level usage tiers enforce monthly spend caps and per-model RPM, input-token-per-minute, and output-token-per-minute limits. Message Batches, Files API, web search, and Managed Agents have additional limits. 429 responses include retry-after w

Developer guide

Anthropic Claude API: Beginner's Guide to the Developer Platform

Anthropic's Claude API is a usage-priced developer platform built around the Messages API. It gives applications access to Claude models for text generation, multimodal input, structured JSON responses, streaming, tool calling, file analysis, web search, code execution, MCP connectivity, and agent workflows through official SDKs or HTTPS requests.
Anthropic's Claude Platform lets developers add Claude model capabilities to applications and services. The recommended starting point is the Messages API, which accepts a model, instructions, conversation messages, and optional tools or files, then returns typed content blocks. This guide explains how to obtain access, send a first request, understand responses, estimate usage, use streaming and tools, and evaluate the platform's production trade-offs.
Anthropic's Claude Platform provides direct, usage-priced access to Claude through the Messages API, with SDKs and support for streaming, tools, structured JSON, files, multimodal input, web search, MCP, and agent workflows.

What the Claude API is and when to use it

The Claude API is Anthropic's direct developer interface for sending requests to Claude models from your own application. The main endpoint is the Messages API at https://api.anthropic.com/v1/messages. A request can contain system instructions, a sequence of user and assistant messages, optional images or documents, and tools that Claude may request your application to run.

The API is stateless: your application is responsible for retaining conversation history and sending the relevant messages with each request. This makes it suitable for chat applications, document analysis, coding assistants, structured-data extraction, customer-service workflows, and agents that call external systems.

Use the direct API when you want to control application logic, data storage, tool execution, authentication, and deployment. Anthropic also provides higher-level options such as the Claude Agent SDK and Claude Managed Agents for applications that need more managed agent execution.

Who is it for?

The platform is aimed at developers building production services or prototypes that need Claude model access. Beginners can start with a single HTTPS request or an official SDK. Intermediate developers can add streaming, validated structured output, reusable files, client-side tools, server-side tools, prompt caching, batch processing, or MCP connectivity.

Getting access and obtaining credentials

  1. Create a Claude Console account.
  2. Create an Anthropic API key.
  3. Store the key in the ANTHROPIC_API_KEY environment variable rather than placing it directly in source code.
  4. Install an official SDK or send HTTPS requests to the API.

Direct requests authenticate with the x-api-key header and must include the anthropic-version header. SDKs normally read ANTHROPIC_API_KEY automatically. Keep the key on a server or other protected runtime; do not expose it in browser code or public repositories.

Choosing the API and model

For new direct integrations, use the Messages API rather than older Claude completion-style endpoints. Select an active model identifier from Anthropic's current models catalog. Model availability, capabilities, pricing, and retirement dates can change, so avoid hard-coding a model that has been deprecated and review Anthropic's model lifecycle documentation periodically.

Model choice affects cost, response speed, context handling, and available features. The supplied research does not specify a single universally appropriate model or fixed price, so applications should compare the currently active models against their workload and budget.

Making the first request

The following raw HTTP example uses a placeholder model identifier. Replace active-model-id with an active model from Anthropic's current catalog.

curl https://api.anthropic.com/v1/messages 
  --header "x-api-key: ${ANTHROPIC_API_KEY}" 
  --header "anthropic-version: 2023-06-01" 
  --header "content-type: application/json" 
  --data '{
    "model": "active-model-id",
    "max_tokens": 512,
    "system": "You are a concise technical assistant.",
    "messages": [
      {
        "role": "user",
        "content": "Explain exponential backoff in three sentences."
      }
    ]
  }'

The request specifies a model, an output-token limit, an optional system instruction, and a messages array. The API key is read from the environment, which reduces the risk of accidentally committing it to source control.

The same request with the Python SDK

Anthropic provides official SDKs for Python, TypeScript, C#, Go, Java, PHP, and Ruby. This example uses the current Python client pattern.

pip install anthropic
import os
from anthropic import Anthropic

client = Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])

response = client.messages.create(
    model="active-model-id",
    max_tokens=512,
    system="You are a concise technical assistant.",
    messages=[
        {"role": "user", "content": "Explain exponential backoff in three sentences."}
    ],
)

print("n".join(
    block.text for block in response.content if block.type == "text"
))

Understanding the response

A Messages API response is not simply a plain text string. It contains content blocks, and each block has a type. Text responses use text blocks; tool requests use tool-use blocks; other capabilities can produce other block types. Code should inspect the block type and extract text deliberately.

For a basic text response, applications commonly select blocks where block.type == "text". Tool-enabled applications instead look for a tool-use block, execute the requested function, and send the result back in a subsequent user message. This content-block design lets one response contain more than one kind of output.

How pricing generally works

Claude API pricing is usage-based. Charges generally depend on the selected model's input and output token rates, with different models having different prices. A token is a unit of text processed by the model; the exact number depends on the content and tokenization.

Additional pricing rules may apply to prompt-cache writes and reads, long-context usage, batch processing, and server-side tools. The Message Batches API provides discounted asynchronous processing for eligible workloads. Web search and other server-side tools may also add feature-specific charges.

Because model prices change and the research does not provide a single current rate, check Anthropic's pricing page before estimating costs. In production, record token usage, model selection, tool activity, and request volume so that application-level cost can be monitored.

Important capabilities

Streaming responses

Streaming sends response events as they are generated instead of waiting for the complete message. It is useful for chat interfaces and long responses because the application can display text progressively and reduce the time before users see output.

with client.messages.stream(
    model="active-model-id",
    max_tokens=512,
    messages=[
        {"role": "user", "content": "Explain server-sent events briefly."}
    ],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
print()

At the HTTP level, streaming uses server-sent events. A streaming client should process event types rather than assuming every event contains displayable text.

Tool and function calling

Tool use lets Claude request that your application run a function. You provide a tool name, description, and JSON input schema. Your code performs the operation—such as querying a database or calling a weather service—and sends the result back to Claude. Anthropic also supports strict tool schemas for validating tool inputs.

tools = [
    {
        "name": "get_weather",
        "description": "Return current weather for a city.",
        "input_schema": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
            "additionalProperties": False,
        },
        "strict": True,
    }
]

response = client.messages.create(
    model="active-model-id",
    max_tokens=512,
    tools=tools,
    messages=[
        {"role": "user", "content": "What is the weather in Boston?"}
    ],
)

for block in response.content:
    if block.type == "tool_use":
        print(block.name, block.input)

The application must validate authorization, arguments, and side effects before executing a requested tool. The model does not directly gain permission to access your systems; your application controls whether and how each requested operation runs.

Structured JSON outputs

Structured outputs let an application request JSON text that follows a supplied JSON Schema. This is useful for extraction, classification, and downstream workflows that need predictable fields. Structured output is different from ordinary prompting for JSON because the request specifies a schema for validation.

response = client.messages.create(
    model="active-model-id",
    max_tokens=512,
    output_config={
        "format": {
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "name": {"type": "string"},
                    "priority": {"type": "string"}
                },
                "required": ["name", "priority"],
                "additionalProperties": False
            }
        }
    },
    messages=[
        {"role": "user", "content": "Extract: Alex Kim has high priority."}
    ],
)

json_text = next(
    block.text for block in response.content if block.type == "text"
)

Parse and validate the returned text in your application before using it. Strict tool use validates tool inputs separately from structured response output.

Files, documents, and images

Vision-capable Claude models can accept images through supported base64 data, URLs, or file references. Documents and PDFs can also be supplied directly or through the Files API. The Files API supports uploading, listing, retrieving, downloading, and deleting reusable files.

File expiration can be configured from 3,600 seconds to 7,776,000 seconds, or 90 days. After expiration, file content is no longer retrievable through the API, although metadata may remain readable for up to 30 days. Expiration should not be treated as an immediate guaranteed-deletion control; explicitly delete sensitive files when they are no longer needed.

Server-side tools and agent features

Depending on the model, account, tool version, and release status, Anthropic provides server-side tools such as web search, web fetch, code execution, tool search, memory, computer use, browser use, and MCP connectivity. The platform also includes Claude Agent SDK functionality and Claude Managed Agents for higher-level agent workflows.

These features can reduce the amount of infrastructure an application must build, but they introduce additional permissions, cost, latency, and operational considerations. Confirm availability and current terms for the specific account and model before designing around a feature.

Playground and developer tools

The Claude Console provides a place to create API keys and experiment with requests. Anthropic's official SDKs provide typed or idiomatic clients, streaming, retries, and error handling across Python, TypeScript, C#, Go, Java, PHP, and Ruby.

For a first prototype, the Console and a short SDK script are usually enough. For production, use environment-based secrets, structured logging, request and cost monitoring, retry handling, and tests for tool and schema behavior.

Limits and production considerations

Rate limits are measured using requests per minute, input tokens per minute, and output tokens per minute. Limits vary by model and organization tier. The supplied documentation reports that current maximums for several active model classes can reach 1,000 RPM, 2,000,000 input tokens per minute, and 400,000 output tokens per minute, but those figures are not universal guarantees.

When a request is throttled, a 429 response can include retry-after and rate-limit headers. Production clients should honor the server-provided delay, use exponential backoff, and monitor remaining request and token capacity. Managed Agents and the Files API have separate limits.

Latency depends on the model, prompt size, output length, reasoning settings, tools, and service tier. Streaming can reduce time to the first visible output. Tool loops also add network round trips and should have explicit timeouts, failure handling, and limits on repeated calls.

Privacy and data handling

Anthropic's commercial terms state that Anthropic may not train models on customer content from its commercial services. Organizations should still review current commercial terms, service-specific terms, privacy documentation, supported-region requirements, retention controls, and any enterprise or zero-data-retention agreement before sending sensitive information.

Advantages and limitations

AreaWhat to consider
API designThe Messages API provides a consistent foundation for text, multimodal input, tools, streaming, and structured output.
Developer experienceOfficial SDKs, a Console, streaming support, retries, and typed content blocks support both prototypes and production services.
Agent workflowsTool use, MCP, server-side tools, the Agent SDK, and Managed Agents support increasingly complex workflows.
CostsUsage is token-based, model-specific, and potentially increased by caching, batches, long context, or server-side tools.
OperationsApplications must manage conversation history, tool execution, retries, rate limits, secrets, monitoring, and model lifecycle changes.
Fine-tuningThe direct platform does not currently expose a general-purpose public fine-tuning API, so developers generally rely on prompting, retrieval, tools, caching, structured outputs, and model selection.

When the Claude API is a good or poor choice

The API is a good fit when an application needs Claude's instruction following, long-context processing, reasoning, coding assistance, document analysis, multimodal input, structured extraction, tool use, or agent orchestration. It is also a practical choice when the team wants direct control over the request lifecycle rather than only a prebuilt chat interface.

It may be a poor fit when a workload requires a general-purpose public fine-tuning API, fixed predictable costs independent of usage, capabilities unavailable in the selected model or account, or a fully managed agent system without application-side tool and security work. Confirm model availability, pricing, rate limits, retention terms, and required server-side tools before committing to the design.

For a reliable first implementation, begin with one Messages API request, extract content blocks explicitly, add retries and usage monitoring, and introduce streaming, tools, files, or structured outputs only when the application needs them.

Sources 18