What is the Amazon Bedrock API?
Amazon Bedrock is an AWS platform for building applications with foundation models, which are general-purpose AI models that can generate or analyze content. Instead of deploying and operating models yourself, you call AWS APIs and select a supported model from Amazon or a third-party provider.
Bedrock includes several related API surfaces. The bedrock-runtime endpoint handles inference, meaning requests that ask a model to generate or analyze content. The bedrock endpoint provides control-plane operations such as model discovery, customization, evaluations and guardrails. Agent and knowledge-base features use bedrock-agent and bedrock-agent-runtime. AWS also documents a bedrock-mantle endpoint for selected compatible APIs.
For a new conversational application, AWS generally recommends the Converse API when the selected model supports it. Converse provides a common message format for supported models, including system instructions, multi-turn conversations, tool use and some multimodal content. Use ConverseStream when the application should display output as it is generated.
Who is Bedrock for?
Bedrock is intended for developers and organizations already using AWS or needing AWS identity, regional controls, usage monitoring and access to multiple model providers. It can support chat applications, document analysis, retrieval-augmented generation, agents, content generation, embeddings, image generation, automated workflows and model customization.
It is less suitable when you want a very simple standalone API with no AWS account, IAM configuration or regional infrastructure decisions. It is also not a single model: capabilities, prices, input formats and availability vary by the model and Region you choose.
Getting access and obtaining credentials
You need an AWS account, access to the relevant Bedrock models and permission to call the selected API operations. In production, use an IAM role, temporary AWS credentials or another AWS-managed identity pattern. IAM controls which users, roles and applications may discover models, invoke them, use tools or access other Bedrock resources.
AWS also provides Bedrock API keys for exploration and development. The documented long-term API keys expire after 30 days, so they should not normally be used as the permanent credential mechanism for a production service. API-key requests use an Authorization: Bearer header, while AWS SDK requests normally use standard AWS credential signing.
The main runtime endpoint follows this pattern:
https://bedrock-runtime.{region}.amazonaws.comBefore sending a request, verify that the chosen model is available in your Region and that your identity has permission to invoke it. Model access, supported Regions and capabilities are not identical across the Bedrock catalog.
Choosing an API and model
Start by deciding what the application needs rather than assuming every model supports every Bedrock feature.
| Requirement | Typical choice | Important check |
|---|---|---|
| Multi-turn text or multimodal conversation | Converse | Confirm that the selected model supports Converse and the required content types. |
| Incremental response display | ConverseStream | Confirm streaming support and handle response events incrementally. |
| Model-specific request features | InvokeModel | Use the model’s documented request and response format. |
| Bulk or asynchronous processing | Batch or related Bedrock workflows | Check model, Region and service-tier availability. |
| Knowledge-grounded answers | Knowledge bases or application-managed retrieval | Check ingestion, retrieval, permissions and supported models. |
| Agentic workflows | Bedrock agent capabilities or AgentCore | Bedrock Agents Classic is not open to new customers; AWS directs new workloads toward AgentCore. |
Keep the model ID configurable in application settings. This makes it easier to test another model, move between Regions or change models without rewriting application logic.
Making a first request with Converse
The following Python example uses the current Boto3 Bedrock Runtime client. Install the SDK with python -m pip install boto3, configure AWS credentials through the normal AWS credential chain, set AWS_REGION if needed and provide a model ID available in that Region.
import os
import boto3
from botocore.exceptions import BotoCoreError, ClientError
region = os.getenv("AWS_REGION", "us-east-1")
model_id = os.environ["BEDROCK_MODEL_ID"]
client = boto3.client("bedrock-runtime", region_name=region)
try:
response = client.converse(
modelId=model_id,
system=[{"text": "Answer clearly and briefly for a developer."}],
messages=[
{
"role": "user",
"content": [{"text": "What is Amazon Bedrock?"}],
}
],
inferenceConfig={
"maxTokens": 300,
"temperature": 0.2,
},
)
text = "".join(
part.get("text", "")
for part in response["output"]["message"].get("content", [])
)
print(text)
except (ClientError, BotoCoreError) as exc:
print(f"Bedrock request failed: {exc}")
raiseThe equivalent request can be sent over HTTP using the Converse resource path /model/{modelId}/converse. With an API key, send the request to the regional runtime endpoint and include the bearer token:
curl --fail-with-body --silent --show-error
--request POST "https://bedrock-runtime.${AWS_REGION}/model/${BEDROCK_MODEL_ID}/converse"
--header 'Content-Type: application/json'
--header "Authorization: Bearer ${AWS_BEARER_TOKEN_BEDROCK}"
--data '{
"messages": [
{"role": "user", "content": [{"text": "Explain Bedrock in one sentence."}]}
],
"inferenceConfig": {"maxTokens": 200, "temperature": 0.2}
}'For production, prefer IAM roles or temporary credentials over storing a long-lived API key in an application server.
Understanding the response
Converse returns an output message containing one or more content parts. A simple text application can join the text values, as shown in the Python example. The response also includes metadata such as usage and latency information, which can be useful for cost tracking and operational monitoring.
Do not assume that every response contains only text. Depending on the model and request, content can include other supported parts, tool-use requests or multimodal results. Application code should inspect the response structure and handle stop reasons, tool requests and errors explicitly.
How Amazon Bedrock pricing works
Bedrock does not have one universal API subscription price. Cost is primarily determined by the selected model, Region, modality, input and output usage and service tier. AWS publishes model-specific prices for tokens, images, embeddings and other units where applicable.
Capacity and responsiveness options can include the following, depending on the model and Region:
- Standard: General on-demand inference.
- Flex: Intended for workloads that can trade responsiveness for lower cost.
- Priority: Intended for latency-sensitive requests and typically priced at a premium.
- Batch: Intended for asynchronous bulk workloads and may provide discounted pricing.
- Provisioned or Reserved Throughput: Intended for predictable production capacity where supported.
Estimate cost using the pricing page for the exact model and Region rather than applying a platform-wide rate. Monitor input tokens, output tokens, image or embedding usage and the service tier in your own telemetry.
Capabilities available through the platform
Conversation, streaming and multimodal input
Converse supports system instructions, multi-turn messages and common inference settings across compatible models. ConverseStream returns incremental events so a user interface can show text while generation is in progress instead of waiting for the complete response.
Compatible models may accept images and documents alongside text. The exact formats, size limits, Regions and model support must be checked for the selected model. Multimodal input does not mean every Bedrock model can process every file type.
Tool and function calling
Tool use lets a model request an application-defined function, such as looking up an order or querying an internal service. Your application executes the function, validates its arguments and sends the result back to the model. Bedrock does not automatically grant a model access to your systems; the application remains responsible for execution, authorization and safety checks.
A typical tool workflow is:
- Send the user message together with a tool specification.
- Inspect the response for a tool-use request.
- Validate the requested name and arguments.
- Run only the permitted application function.
- Send the tool result back in the expected message format.
- Continue the conversation and present the final response.
Structured outputs
For workflows that pass model results to software, free-form prose can be difficult to parse reliably. Supported Bedrock models and APIs can constrain results with JSON Schema output configuration or strict tool definitions. Verify support for the selected model and endpoint, then validate the returned data in your application even when a schema was supplied.
Files, knowledge bases and web search
Supported models can accept documents or other content as input, subject to model-specific restrictions. Bedrock also provides knowledge-base capabilities for retrieval-augmented applications, where relevant information is retrieved before the model generates an answer.
Bedrock Web Search can provide current web information to supported models through the Responses API and IAM-controlled search and fetch operations. It is a separate capability with its own model, permission and availability requirements; do not assume it is available through every Converse request.
Customization, agents and developer tools
Depending on model availability, Bedrock supports fine-tuning, continued pre-training, reinforcement fine-tuning or other customization workflows. It also provides guardrails, evaluations, prompt management, flows and agent-related services.
The AWS Management Console includes text, chat and image playgrounds. These are useful for comparing prompts and testing supported models before writing integration code. The playground does not remove the need to configure IAM permissions, select a Region and confirm production quotas.
Streaming output with Boto3
Use converse_stream when the application should print or display text as events arrive:
stream = client.converse_stream(
modelId=model_id,
messages=[
{
"role": "user",
"content": [{"text": "Give three practical Bedrock API tips."}],
}
],
inferenceConfig={"maxTokens": 300, "temperature": 0.2},
)
for event in stream["stream"]:
delta = event.get("contentBlockDelta", {}).get("delta", {})
if "text" in delta:
print(delta["text"], end="", flush=True)
print()Streaming responses are event sequences rather than one completed JSON message. Your code should handle text deltas and account for completion, tool-use and error events according to the selected model’s response contract.
Limits and production considerations
Quotas and throttling
Quotas depend on the model, Region, account, endpoint and request type. The runtime endpoint commonly uses per-model token quotas and, for some models, requests-per-minute quotas. Bedrock Mantle has separate input-token and output-token quotas for applicable models. Check AWS Service Quotas for exact values and possible increases.
Use bounded concurrency, exponential backoff with jitter and careful retry classification. A 429 response can indicate throttling or model-readiness conditions, while 503 and model-specific overload responses can indicate temporary capacity constraints. Do not blindly retry authorization failures, validation errors or malformed requests.
Regions, routing and latency
Latency depends on the model, Region, request size, traffic, service tier and routing. Cross-Region inference profiles can route requests across supported Regions. Standard, Priority, Flex, provisioned capacity and Reserved options provide different cost, latency and capacity trade-offs where supported.
Regional routing also matters for data residency. If cross-Region inference is used, review where retained inputs and outputs may be stored and whether that matches organizational requirements.
Privacy and data handling
AWS documentation states that Bedrock customer inputs and outputs are not used to train or improve the underlying foundation models by AWS or model providers as a general service behavior. Handling can vary by model, endpoint, feature and applicable provider terms.
Bedrock documents default abuse-detection and retention behavior, including special cases where content may be retained within the AWS boundary for review or abuse prevention. Review the exact model documentation, account-level retention controls, encryption settings, CloudTrail coverage and IAM permissions before processing regulated or sensitive information.
Model compatibility
Support for streaming, images, documents, structured outputs, tool use and fine-tuning is not uniform across the catalog. Treat model capability as a configuration that must be tested, not as a guarantee of the Bedrock platform as a whole. Keep capability checks and model-specific settings separate from business logic.
SDKs and documentation
AWS provides official SDKs for Python through Boto3, JavaScript and TypeScript through AWS SDK for JavaScript v3, Java, Go, .NET, C++, Ruby, PHP and other AWS SDK languages. The SDKs handle AWS request signing and expose service-specific request and response structures.
For a first implementation, use the Bedrock Runtime SDK client and the Converse operation. Then consult the model availability table, Runtime API Reference, quotas, pricing and model-specific capability documentation. This is important because a syntactically valid request can still fail if the model is unavailable in the chosen Region or does not support the requested feature.
When Amazon Bedrock is a good or poor choice
Bedrock is a good choice when you need managed access to several foundation-model providers, AWS IAM and regional controls, model selection, usage-based billing, console experimentation or integration with AWS infrastructure. It is also useful when an application needs a combination of inference, tools, knowledge bases, agents, structured outputs and model customization under one AWS service family.
It may be a poor choice when the team does not want AWS account and IAM administration, needs one extremely simple provider-specific API, requires a capability unavailable in the target Region or expects identical behavior across every model. The main evaluation task is to test the exact model, API surface, Region, quota, data-handling terms and pricing that the production application will use.
