Developer platform

Amazon Bedrock API

Developer overview for Amazon, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Amazon Bedrock Runtime Converse API
SDK support Official AWS SDK support includes Python Boto3, JavaScript or TypeScript AWS SDK v3, Java, Go, .NET, C++, Ruby, PHP and other AWS SDK languages.
Rate limits Quotas are model-, Region-, account- and endpoint-dependent. bedrock-runtime generally uses per-model token quotas and, for some models, requests-per-minute quotas. bedrock-mantle uses separate input-token and output-token quotas for applicable models. Ex
Platform

API overview

Endpoints

API access

Base URL https://bedrock-runtime.{region}.amazonaws.com
Primary API Amazon Bedrock Runtime Converse API
Pricing

API pricing

Pricing model Usage-based AWS pricing determined by model, modality, tokens or other model-specific units, Region, service tier and optional capacity commitments.

Model-specific on-demand pricing is published by AWS. Bedrock also supports Standard, Flex, Priority, Batch and Reserved or Provisioned Throughput options for supported models. There is no single universal API subscription price.

Developer experience

SDKs & usability

SDKs Official AWS SDK support includes Python Boto3, JavaScript or TypeScript AWS SDK v3, Java, Go, .NET, C++, Ruby, PHP and other AWS SDK languages.
Ease of use Moderate. Converse provides a unified interface for supported models, while AWS credentials, IAM policies, Regions, model access and model-specific request formats add operational complexity.
Documentation Extensive and technically detailed, with separate user guides, API references, quotas, model-availability pages, SDK examples and service-specific documentation.
Latency Latency depends on model, Region, request size, service tier, traffic and routing. Converse responses include latency metrics. Standard, Priority, Flex, cross-Region inference, provisioned capacity and Reserved tiers provide different cost, latency or cap
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Fine-tuning
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

Amazon Bedrock provides control-plane and runtime APIs for model discovery, inference, model customization, evaluation, guardrails, knowledge bases, agents, prompt management, flows, data automation and related services. Converse is the recommended common conversational interface when supported by the selected model; ConverseStream provides streaming. InvokeModel and InvokeModelWithResponseStream remain available for model-specific or lower-level inference. Tool use and function calling are supported through tool configuration for compatible models. Structured outputs are supported through JSON Schema output configuration and strict tool definitions for supported models. Image and document inputs are supported only for compatible models and formats. Web Search is currently available as a built-in Bedrock capability for supported models through the Responses API. Bedrock Agents Classic is no longer open to new customers; AWS directs new customers toward Amazon Bedrock AgentCore for comparable agent-runtime capabilities.

Examples

API examples

set -euo pipefail

: "${AWS_BEARER_TOKEN_BEDROCK:?Set AWS_BEARER_TOKEN_BEDROCK to an Amazon Bedrock API key}"
REGION="${AWS_REGION:-us-east-1}"
MODEL_ID="${BEDROCK_MODEL_ID:?Set BEDROCK_MODEL_ID to a model available in your Region}"
URL="https://bedrock-runtime.${REGION}.amazonaws.com/model/${MODEL_ID}/converse"

response="$(curl --fail-with-body --silent --show-error --request POST "$URL" \
  --header 'Content-Type: application/json' \
  --header "Authorization: Bearer ${AWS_BEARER_TOKEN_BEDROCK}" \
  --data @- <<'JSON'
{
  "system": [
    {"text": "You are a concise technical assistant."}
  ],
  "messages": [
    {"role": "user", "content": [{"text": "Explain Amazon Bedrock Converse in one sentence."}]}
  ],
  "inferenceConfig": {
    "maxTokens": 200,
    "temperature": 0.2
  }
}
JSON
)"
printf '%s\n' "$response" | jq -r '.output.message.content[]? | .text? // empty' || {
  printf '%s\n' "$response"
  exit 1
}
# Install: python -m pip install boto3
import json
import os
import boto3
from botocore.exceptions import BotoCoreError, ClientError

region = os.getenv("AWS_REGION", "us-east-1")
model_id = os.environ["BEDROCK_MODEL_ID"]
client = boto3.client("bedrock-runtime", region_name=region)

messages = [
    {
        "role": "user",
        "content": [{"text": "What is Amazon Bedrock?"}],
    }
]

try:
    response = client.converse(
        modelId=model_id,
        system=[{"text": "Answer clearly and briefly for a developer."}],
        messages=messages,
        inferenceConfig={"maxTokens": 300, "temperature": 0.2},
    )
    text = "".join(
        item.get("text", "")
        for item in response["output"]["message"].get("content", [])
    )
    print(text)

    stream = client.converse_stream(
        modelId=model_id,
        messages=[
            {
                "role": "user",
                "content": [{"text": "Give three practical Bedrock API tips."}],
            }
        ],
        inferenceConfig={"maxTokens": 300, "temperature": 0.2},
    )
    for event in stream["stream"]:
        delta = event.get("contentBlockDelta", {}).get("delta", {})
        if "text" in delta:
            print(delta["text"], end="", flush=True)
    print()
except (ClientError, BotoCoreError) as exc:
    print(f"Bedrock request failed: {exc}", file=__import__("sys").stderr)
    raise SystemExit(1)
// Install: npm install @aws-sdk/client-bedrock-runtime
import {
  BedrockRuntimeClient,
  ConverseCommand,
  ConverseStreamCommand,
} from "@aws-sdk/client-bedrock-runtime";

const region = process.env.AWS_REGION ?? "us-east-1";
const modelId = process.env.BEDROCK_MODEL_ID;
if (!modelId) throw new Error("Set BEDROCK_MODEL_ID");

const client = new BedrockRuntimeClient({ region });

try {
  const response = await client.send(new ConverseCommand({
    modelId,
    system: [{ text: "You are a concise technical assistant." }],
    messages: [
      { role: "user", content: [{ text: "Explain Amazon Bedrock in one sentence." }] },
    ],
    inferenceConfig: { maxTokens: 200, temperature: 0.2 },
  }));

  const text = (response.output?.message?.content ?? [])
    .map((part) => part.text ?? "")
    .join("");
  console.log(text);

  const stream = await client.send(new ConverseStreamCommand({
    modelId,
    messages: [
      { role: "user", content: [{ text: "List three Bedrock API best practices." }] },
    ],
    inferenceConfig: { maxTokens: 300, temperature: 0.2 },
  }));

  for await (const event of stream.stream ?? []) {
    const textDelta = event.contentBlockDelta?.delta?.text;
    if (textDelta) process.stdout.write(textDelta);
  }
  process.stdout.write("\n");
} catch (error) {
  console.error("Bedrock request failed:", error);
  process.exitCode = 1;
}
<?php

$region = getenv('AWS_REGION') ?: 'us-east-1';
$modelId = getenv('BEDROCK_MODEL_ID');
$apiKey = getenv('AWS_BEARER_TOKEN_BEDROCK');

if (!$modelId || !$apiKey) {
    fwrite(STDERR, "Set BEDROCK_MODEL_ID and AWS_BEARER_TOKEN_BEDROCK\n");
    exit(1);
}

$url = "https://bedrock-runtime.{$region}.amazonaws.com/model/" . rawurlencode($modelId) . "/converse";
$payload = [
    'system' => [
        ['text' => 'You are a concise technical assistant.']
    ],
    'messages' => [
        [
            'role' => 'user',
            'content' => [
                ['text' => 'Explain Amazon Bedrock in one sentence.']
            ]
        ]
    ],
    'inferenceConfig' => [
        'maxTokens' => 200,
        'temperature' => 0.2
    ]
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $apiKey
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 60
]);

$body = curl_exec($ch);
if ($body === false) {
    fwrite(STDERR, 'cURL error: ' . curl_error($ch) . "\n");
    curl_close($ch);
    exit(1);
}

$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
$data = json_decode($body, true);

if ($status < 200 || $status >= 300 || !is_array($data)) {
    fwrite(STDERR, "Bedrock HTTP error {$status}: {$body}\n");
    exit(1);
}

$text = '';
foreach (($data['output']['message']['content'] ?? []) as $part) {
    $text .= $part['text'] ?? '';
}
echo $text . PHP_EOL;
Policies

Data & usage

Data training

AWS documentation states that Amazon Bedrock customer inputs and outputs are not used to train or improve the underlying foundation models by AWS or model providers as a general service behavior. Data handling can vary by model, endpoint, feature and applicable provider terms. Customers should review the selected model's documentation and terms before production use.

Data retention

Retention depends on endpoint, model, feature and account settings. Bedrock documents default abuse-detection retention behavior and special model cases where prompts and completions may be retained within the AWS boundary for up to 30 days for required review. Cross-Region inference may store retained data in destination Regions. Account-level retention controls and model-specific documentation should be reviewed for sensitive workloads.

Rate limits

Quotas are model-, Region-, account- and endpoint-dependent. bedrock-runtime generally uses per-model token quotas and, for some models, requests-per-minute quotas. bedrock-mantle uses separate input-token and output-token quotas for applicable models. Ex

Developer guide

Amazon Bedrock API: Developer Platform, Features, Pricing and Examples

Amazon Bedrock is Amazon Web Services’ managed developer platform for accessing foundation models from Amazon and third-party providers through unified APIs, SDKs, console playgrounds, agents, knowledge bases, web search, guardrails, fine-tuning, structured outputs and related generative-AI services.
Amazon Bedrock provides a managed API layer for building generative-AI applications without deploying foundation models yourself. The platform exposes model inference through the Bedrock Runtime APIs, with Converse and ConverseStream recommended for supported conversational workloads, while additional control-plane, agent, knowledge-base, customization and evaluation APIs support production application development.
Amazon Bedrock is AWS’s managed developer platform for accessing Amazon and third-party foundation models through common inference APIs, SDKs and console tools. It supports conversational inference, streaming, tool use, structured JSON outputs, multimodal content, web search, agents, knowledge bases, fine-tuning and model evaluation.

What is the Amazon Bedrock API?

Amazon Bedrock is an AWS platform for building applications with foundation models, which are general-purpose AI models that can generate or analyze content. Instead of deploying and operating models yourself, you call AWS APIs and select a supported model from Amazon or a third-party provider.

Bedrock includes several related API surfaces. The bedrock-runtime endpoint handles inference, meaning requests that ask a model to generate or analyze content. The bedrock endpoint provides control-plane operations such as model discovery, customization, evaluations and guardrails. Agent and knowledge-base features use bedrock-agent and bedrock-agent-runtime. AWS also documents a bedrock-mantle endpoint for selected compatible APIs.

For a new conversational application, AWS generally recommends the Converse API when the selected model supports it. Converse provides a common message format for supported models, including system instructions, multi-turn conversations, tool use and some multimodal content. Use ConverseStream when the application should display output as it is generated.

Who is Bedrock for?

Bedrock is intended for developers and organizations already using AWS or needing AWS identity, regional controls, usage monitoring and access to multiple model providers. It can support chat applications, document analysis, retrieval-augmented generation, agents, content generation, embeddings, image generation, automated workflows and model customization.

It is less suitable when you want a very simple standalone API with no AWS account, IAM configuration or regional infrastructure decisions. It is also not a single model: capabilities, prices, input formats and availability vary by the model and Region you choose.

Getting access and obtaining credentials

You need an AWS account, access to the relevant Bedrock models and permission to call the selected API operations. In production, use an IAM role, temporary AWS credentials or another AWS-managed identity pattern. IAM controls which users, roles and applications may discover models, invoke them, use tools or access other Bedrock resources.

AWS also provides Bedrock API keys for exploration and development. The documented long-term API keys expire after 30 days, so they should not normally be used as the permanent credential mechanism for a production service. API-key requests use an Authorization: Bearer header, while AWS SDK requests normally use standard AWS credential signing.

The main runtime endpoint follows this pattern:

https://bedrock-runtime.{region}.amazonaws.com

Before sending a request, verify that the chosen model is available in your Region and that your identity has permission to invoke it. Model access, supported Regions and capabilities are not identical across the Bedrock catalog.

Choosing an API and model

Start by deciding what the application needs rather than assuming every model supports every Bedrock feature.

RequirementTypical choiceImportant check
Multi-turn text or multimodal conversationConverseConfirm that the selected model supports Converse and the required content types.
Incremental response displayConverseStreamConfirm streaming support and handle response events incrementally.
Model-specific request featuresInvokeModelUse the model’s documented request and response format.
Bulk or asynchronous processingBatch or related Bedrock workflowsCheck model, Region and service-tier availability.
Knowledge-grounded answersKnowledge bases or application-managed retrievalCheck ingestion, retrieval, permissions and supported models.
Agentic workflowsBedrock agent capabilities or AgentCoreBedrock Agents Classic is not open to new customers; AWS directs new workloads toward AgentCore.

Keep the model ID configurable in application settings. This makes it easier to test another model, move between Regions or change models without rewriting application logic.

Making a first request with Converse

The following Python example uses the current Boto3 Bedrock Runtime client. Install the SDK with python -m pip install boto3, configure AWS credentials through the normal AWS credential chain, set AWS_REGION if needed and provide a model ID available in that Region.

import os
import boto3
from botocore.exceptions import BotoCoreError, ClientError

region = os.getenv("AWS_REGION", "us-east-1")
model_id = os.environ["BEDROCK_MODEL_ID"]
client = boto3.client("bedrock-runtime", region_name=region)

try:
    response = client.converse(
        modelId=model_id,
        system=[{"text": "Answer clearly and briefly for a developer."}],
        messages=[
            {
                "role": "user",
                "content": [{"text": "What is Amazon Bedrock?"}],
            }
        ],
        inferenceConfig={
            "maxTokens": 300,
            "temperature": 0.2,
        },
    )

    text = "".join(
        part.get("text", "")
        for part in response["output"]["message"].get("content", [])
    )
    print(text)
except (ClientError, BotoCoreError) as exc:
    print(f"Bedrock request failed: {exc}")
    raise

The equivalent request can be sent over HTTP using the Converse resource path /model/{modelId}/converse. With an API key, send the request to the regional runtime endpoint and include the bearer token:

curl --fail-with-body --silent --show-error 
  --request POST "https://bedrock-runtime.${AWS_REGION}/model/${BEDROCK_MODEL_ID}/converse" 
  --header 'Content-Type: application/json' 
  --header "Authorization: Bearer ${AWS_BEARER_TOKEN_BEDROCK}" 
  --data '{
    "messages": [
      {"role": "user", "content": [{"text": "Explain Bedrock in one sentence."}]}
    ],
    "inferenceConfig": {"maxTokens": 200, "temperature": 0.2}
  }'

For production, prefer IAM roles or temporary credentials over storing a long-lived API key in an application server.

Understanding the response

Converse returns an output message containing one or more content parts. A simple text application can join the text values, as shown in the Python example. The response also includes metadata such as usage and latency information, which can be useful for cost tracking and operational monitoring.

Do not assume that every response contains only text. Depending on the model and request, content can include other supported parts, tool-use requests or multimodal results. Application code should inspect the response structure and handle stop reasons, tool requests and errors explicitly.

How Amazon Bedrock pricing works

Bedrock does not have one universal API subscription price. Cost is primarily determined by the selected model, Region, modality, input and output usage and service tier. AWS publishes model-specific prices for tokens, images, embeddings and other units where applicable.

Capacity and responsiveness options can include the following, depending on the model and Region:

  • Standard: General on-demand inference.
  • Flex: Intended for workloads that can trade responsiveness for lower cost.
  • Priority: Intended for latency-sensitive requests and typically priced at a premium.
  • Batch: Intended for asynchronous bulk workloads and may provide discounted pricing.
  • Provisioned or Reserved Throughput: Intended for predictable production capacity where supported.

Estimate cost using the pricing page for the exact model and Region rather than applying a platform-wide rate. Monitor input tokens, output tokens, image or embedding usage and the service tier in your own telemetry.

Capabilities available through the platform

Conversation, streaming and multimodal input

Converse supports system instructions, multi-turn messages and common inference settings across compatible models. ConverseStream returns incremental events so a user interface can show text while generation is in progress instead of waiting for the complete response.

Compatible models may accept images and documents alongside text. The exact formats, size limits, Regions and model support must be checked for the selected model. Multimodal input does not mean every Bedrock model can process every file type.

Tool and function calling

Tool use lets a model request an application-defined function, such as looking up an order or querying an internal service. Your application executes the function, validates its arguments and sends the result back to the model. Bedrock does not automatically grant a model access to your systems; the application remains responsible for execution, authorization and safety checks.

A typical tool workflow is:

  1. Send the user message together with a tool specification.
  2. Inspect the response for a tool-use request.
  3. Validate the requested name and arguments.
  4. Run only the permitted application function.
  5. Send the tool result back in the expected message format.
  6. Continue the conversation and present the final response.

Structured outputs

For workflows that pass model results to software, free-form prose can be difficult to parse reliably. Supported Bedrock models and APIs can constrain results with JSON Schema output configuration or strict tool definitions. Verify support for the selected model and endpoint, then validate the returned data in your application even when a schema was supplied.

Supported models can accept documents or other content as input, subject to model-specific restrictions. Bedrock also provides knowledge-base capabilities for retrieval-augmented applications, where relevant information is retrieved before the model generates an answer.

Bedrock Web Search can provide current web information to supported models through the Responses API and IAM-controlled search and fetch operations. It is a separate capability with its own model, permission and availability requirements; do not assume it is available through every Converse request.

Depending on model availability, Bedrock supports fine-tuning, continued pre-training, reinforcement fine-tuning or other customization workflows. It also provides guardrails, evaluations, prompt management, flows and agent-related services.

The AWS Management Console includes text, chat and image playgrounds. These are useful for comparing prompts and testing supported models before writing integration code. The playground does not remove the need to configure IAM permissions, select a Region and confirm production quotas.

Streaming output with Boto3

Use converse_stream when the application should print or display text as events arrive:

stream = client.converse_stream(
    modelId=model_id,
    messages=[
        {
            "role": "user",
            "content": [{"text": "Give three practical Bedrock API tips."}],
        }
    ],
    inferenceConfig={"maxTokens": 300, "temperature": 0.2},
)

for event in stream["stream"]:
    delta = event.get("contentBlockDelta", {}).get("delta", {})
    if "text" in delta:
        print(delta["text"], end="", flush=True)
print()

Streaming responses are event sequences rather than one completed JSON message. Your code should handle text deltas and account for completion, tool-use and error events according to the selected model’s response contract.

Limits and production considerations

Quotas and throttling

Quotas depend on the model, Region, account, endpoint and request type. The runtime endpoint commonly uses per-model token quotas and, for some models, requests-per-minute quotas. Bedrock Mantle has separate input-token and output-token quotas for applicable models. Check AWS Service Quotas for exact values and possible increases.

Use bounded concurrency, exponential backoff with jitter and careful retry classification. A 429 response can indicate throttling or model-readiness conditions, while 503 and model-specific overload responses can indicate temporary capacity constraints. Do not blindly retry authorization failures, validation errors or malformed requests.

Regions, routing and latency

Latency depends on the model, Region, request size, traffic, service tier and routing. Cross-Region inference profiles can route requests across supported Regions. Standard, Priority, Flex, provisioned capacity and Reserved options provide different cost, latency and capacity trade-offs where supported.

Regional routing also matters for data residency. If cross-Region inference is used, review where retained inputs and outputs may be stored and whether that matches organizational requirements.

Privacy and data handling

AWS documentation states that Bedrock customer inputs and outputs are not used to train or improve the underlying foundation models by AWS or model providers as a general service behavior. Handling can vary by model, endpoint, feature and applicable provider terms.

Bedrock documents default abuse-detection and retention behavior, including special cases where content may be retained within the AWS boundary for review or abuse prevention. Review the exact model documentation, account-level retention controls, encryption settings, CloudTrail coverage and IAM permissions before processing regulated or sensitive information.

Model compatibility

Support for streaming, images, documents, structured outputs, tool use and fine-tuning is not uniform across the catalog. Treat model capability as a configuration that must be tested, not as a guarantee of the Bedrock platform as a whole. Keep capability checks and model-specific settings separate from business logic.

SDKs and documentation

AWS provides official SDKs for Python through Boto3, JavaScript and TypeScript through AWS SDK for JavaScript v3, Java, Go, .NET, C++, Ruby, PHP and other AWS SDK languages. The SDKs handle AWS request signing and expose service-specific request and response structures.

For a first implementation, use the Bedrock Runtime SDK client and the Converse operation. Then consult the model availability table, Runtime API Reference, quotas, pricing and model-specific capability documentation. This is important because a syntactically valid request can still fail if the model is unavailable in the chosen Region or does not support the requested feature.

When Amazon Bedrock is a good or poor choice

Bedrock is a good choice when you need managed access to several foundation-model providers, AWS IAM and regional controls, model selection, usage-based billing, console experimentation or integration with AWS infrastructure. It is also useful when an application needs a combination of inference, tools, knowledge bases, agents, structured outputs and model customization under one AWS service family.

It may be a poor choice when the team does not want AWS account and IAM administration, needs one extremely simple provider-specific API, requires a capability unavailable in the target Region or expects identical behavior across every model. The main evaluation task is to test the exact model, API surface, Region, quota, data-handling terms and pricing that the production application will use.

Sources 18