Developer platform

AI21 Studio API

Developer overview for AI21 Labs, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Jamba Chat Completions API
SDK support Official Python and TypeScript SDKs. Raw HTTP is also supported. The SDKs include chat completions, streaming, Maestro, library files, and conversational RAG functionality.
Rate limits Rate limits and quotas are account-, model-, plan-, and deployment-dependent. The reference documentation includes rate-limit guidance, while the SDK documents retry configuration and handling of 429 responses.
Platform

API overview

Endpoints

API access

Base URL https://api.ai21.com/studio/v1
Primary API Jamba Chat Completions API
Pricing

API pricing

Pricing model Usage-based token pricing with enterprise and private-deployment agreements

Hosted access is usage-based and account/model dependent; current commercial terms should be verified in the AI21 account or pricing documentation. Enterprise private deployments use negotiated pricing.

Developer experience

SDKs & usability

SDKs Official Python and TypeScript SDKs. Raw HTTP is also supported. The SDKs include chat completions, streaming, Maestro, library files, and conversational RAG functionality.
Ease of use Moderate to easy. The REST API follows a familiar chat-completions pattern, and official Python and TypeScript SDKs simplify authentication, requests, streaming, Maestro, and file operations.
Documentation Good and actively maintained. Current documentation covers AI21 Studio, Jamba, Maestro, SDKs, authentication, rate limits, files, retrieval, function calling, fine-tuning, deployment, and API references.
Latency No universal latency SLA or fixed latency figure was verified. Streaming is available for Jamba chat completions, and latency depends on model, prompt length, output length, workload, region, and deployment type.
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Fine-tuning
✓ Web search
✓ Structured outputs
✓ Playground
Feature notes

The current AI21 developer surface combines AI21 Studio hosted Jamba APIs with higher-level services. Jamba chat completions support multi-turn messages, streaming, and tool/function calling. The official Python and TypeScript SDKs support chat completions, streaming, Maestro, library files, and conversational RAG. AI21 Maestro provides first-party agent and orchestration capabilities, including planning, tools, web search, file search, polling, requirements, and validated output. File upload is available through the library API. Fine-tuning is documented for Jamba models through full fine-tuning, LoRA, and QLoRA, but exact hosted availability depends on the selected deployment and account. AI21 Studio documentation lists web search, file search, MCP, HTTP tools, and related workflow features. Native image input was not verified in the current first-party API documentation consulted. Ordinary Jamba chat completion JSON-Schema enforcement was not verified; Maestro validated output is the clearest documented structured-output capability.

Examples

API examples

set -euo pipefail

: "${AI21_API_KEY:?Set AI21_API_KEY first}"

curl --fail-with-body --silent --show-error \
  --request POST \
  --url "https://api.ai21.com/studio/v1/chat/completions" \
  --header "Authorization: Bearer ${AI21_API_KEY}" \
  --header "Content-Type: application/json" \
  --data '{
    "model": "jamba-mini",
    "messages": [
      {
        "role": "system",
        "content": "You are a concise technical assistant."
      },
      {
        "role": "user",
        "content": "Explain what an API key is in one sentence."
      }
    ],
    "temperature": 0.2,
    "max_tokens": 100
  }' | python3 -c 'import json,sys; data=json.load(sys.stdin); print(data["choices"][0]["message"]["content"])'

# Streaming example
curl --fail-with-body --silent --show-error \
  --request POST \
  --url "https://api.ai21.com/studio/v1/chat/completions" \
  --header "Authorization: Bearer ${AI21_API_KEY}" \
  --header "Content-Type: application/json" \
  --data '{
    "model": "jamba-mini",
    "messages": [{"role": "user", "content": "Write a short explanation of streaming responses."}],
    "stream": true,
    "max_tokens": 120
  }'
# Install: pip install ai21
import json
import os
from ai21 import AI21Client
from ai21.models.chat import ChatMessage

client = AI21Client(api_key=os.environ["AI21_API_KEY"])

messages = [
    ChatMessage(role="system", content="You are a concise technical assistant."),
    ChatMessage(role="user", content="What is an API key?"),
]

try:
    response = client.chat.completions.create(
        model="jamba-mini",
        messages=messages,
        temperature=0.2,
        max_tokens=100,
    )
    print(response.choices[0].message.content)
except Exception as exc:
    print(f"AI21 request failed: {exc}")
    raise

# Streaming
stream = client.chat.completions.create(
    model="jamba-mini",
    messages=[ChatMessage(role="user", content="Write a short explanation of streaming.")],
    stream=True,
)
for chunk in stream:
    content = getattr(chunk.choices[0].delta, "content", None)
    if content:
        print(content, end="", flush=True)
print()

# Portable JSON-output pattern; validate the parsed result in production.
json_response = client.chat.completions.create(
    model="jamba-mini",
    messages=[
        ChatMessage(role="system", content="Return valid JSON only with keys: answer and confidence."),
        ChatMessage(role="user", content="State whether 2+2 equals 4."),
    ],
    temperature=0,
    max_tokens=100,
)
print(json.loads(json_response.choices[0].message.content))
// Install: npm install ai21
import { AI21 } from "ai21";

const client = new AI21({
  apiKey: process.env.AI21_API_KEY,
});

try {
  const response = await client.chat.completions.create({
    model: "jamba-mini",
    messages: [
      { role: "system", content: "You are a concise technical assistant." },
      { role: "user", content: "What is an API key?" },
    ],
    temperature: 0.2,
    max_tokens: 100,
  });

  console.log(response.choices[0]?.message?.content ?? "");
} catch (error) {
  console.error("AI21 request failed:", error);
  process.exitCode = 1;
}

const stream = await client.chat.completions.create({
  model: "jamba-mini",
  messages: [{ role: "user", content: "Write a short explanation of streaming." }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
process.stdout.write("\n");

const jsonResponse = await client.chat.completions.create({
  model: "jamba-mini",
  messages: [
    { role: "system", content: "Return valid JSON only with keys: answer and confidence." },
    { role: "user", content: "State whether 2+2 equals 4." },
  ],
  temperature: 0,
  max_tokens: 100,
});

console.log(JSON.parse(jsonResponse.choices[0].message.content));
<?php
declare(strict_types=1);

$apiKey = getenv('AI21_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('Set AI21_API_KEY before running this script.');
}

$url = 'https://api.ai21.com/studio/v1/chat/completions';
$payload = [
    'model' => 'jamba-mini',
    'messages' => [
        ['role' => 'system', 'content' => 'You are a concise technical assistant.'],
        ['role' => 'user', 'content' => 'What is an API key?'],
    ],
    'temperature' => 0.2,
    'max_tokens' => 100,
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
        'Accept: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_CONNECTTIMEOUT => 10,
    CURLOPT_TIMEOUT => 120,
]);

$body = curl_exec($ch);
if ($body === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('cURL error: ' . $error);
}

$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? $body;
    throw new RuntimeException("AI21 HTTP {$status}: {$message}");
}

$content = $data['choices'][0]['message']['content'] ?? null;
if (!is_string($content)) {
    throw new RuntimeException('The response did not contain assistant content.');
}

echo $content . PHP_EOL;

// Portable JSON-output pattern; validate the result in production.
$payload['messages'] = [
    ['role' => 'system', 'content' => 'Return valid JSON only with keys: answer and confidence.'],
    ['role' => 'user', 'content' => 'State whether 2+2 equals 4.'],
];
$payload['temperature'] = 0;
$payload['max_tokens'] = 100;

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
        'Accept: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_TIMEOUT => 120,
]);
$jsonBody = curl_exec($ch);
if ($jsonBody === false) {
    throw new RuntimeException('JSON request failed: ' . curl_error($ch));
}
$jsonStatus = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
if ($jsonStatus < 200 || $jsonStatus >= 300) {
    throw new RuntimeException("AI21 JSON request failed with HTTP {$jsonStatus}");
}
$jsonData = json_decode($jsonBody, true, 512, JSON_THROW_ON_ERROR);
$structured = json_decode($jsonData['choices'][0]['message']['content'], true, 512, JSON_THROW_ON_ERROR);
print_r($structured);
Policies

Data & usage

Data training

AI21’s current terms govern customer data, outputs, service use, and account-specific entitlements. The public documentation and terms reviewed for this record do not provide a single universal statement that applies identically to every AI21 Studio plan and private deployment regarding whether API inputs are used for model training. Customers should confirm the applicable account, product, and contractual data-processing terms before sending sensitive data.

Data retention

Retention is deployment- and agreement-dependent. AI21 describes AI21 Studio, partner-cloud, AI21-managed private, and self-managed private deployment options, including VPC and on-premises arrangements for qualifying enterprise use cases. The public materials reviewed do not establish one universal retention period for every hosted API account. Confirm retention, deletion, residency, and logging controls in the applicable account or enterprise agreement.

Rate limits

Rate limits and quotas are account-, model-, plan-, and deployment-dependent. The reference documentation includes rate-limit guidance, while the SDK documents retry configuration and handling of 429 responses.

Developer guide

AI21 Studio API Guide: Jamba, Maestro, SDKs and Developer Features

AI21 Studio is AI21 Labs' current developer platform for hosted Jamba language-model APIs and higher-level services such as file libraries, conversational RAG, tool calling, web search, and Maestro orchestration. This guide explains access, authentication, chat completions, streaming, files, structured output, SDKs, pricing, deployment choices, and production considerations.
AI21 Studio gives developers access to AI21 Labs' hosted Jamba models through a REST API and official Python and TypeScript SDKs. The platform is designed for applications that need text generation, multi-turn conversations, streaming, tool use, retrieval over uploaded files, or agent-style workflows through Maestro. It is primarily a text and document-oriented developer platform, with commercial terms, quotas, and some deployment options varying by account and agreement.
AI21 Studio is AI21 Labs' developer platform for hosted Jamba chat completions and higher-level services including streaming, tool calling, file libraries, conversational RAG, web search, fine-tuning workflows, and Maestro orchestration. The guide covers authentication, SDKs, first requests, response handling, structured-output caveats, pricing, quotas, privacy, deployment choices, and situations where the API is or is not a good fit.

What is the AI21 Studio API?

AI21 Studio is AI21 Labs' current hosted developer platform. Its main API provides access to the Jamba family of language models through chat completions. A chat completion accepts an ordered list of messages, such as system, user, and assistant, and returns generated text.

The platform also includes services beyond ordinary text generation. Developers can upload documents to a managed library, use file search and conversational retrieval-augmented generation (RAG), define tools that an application can execute, and use AI21 Maestro for planning and orchestrating more complex workflows.

The primary hosted API base URL is https://api.ai21.com/studio/v1. AI21 also provides a developer console at https://studio.ai21.com/.

Who is it for?

AI21 Studio is a fit for developers building text-generation services, writing and summarization tools, enterprise knowledge assistants, retrieval-grounded applications, and agent workflows. It is especially relevant when an organization wants options such as managed private deployment, partner-cloud deployment, VPC deployment, or self-managed deployment.

It is less suitable if the main requirement is a broad consumer assistant, native image or audio generation, general-purpose web browsing without configuring the relevant API service, or a large ecosystem of consumer plugins and applications. Native image input was not verified in the current API documentation reviewed for this guide.

Getting access and creating an API key

  1. Create or use an AI21 account with access to AI21 Studio.
  2. Create an API key through the applicable AI21 developer or account interface.
  3. Store the key in a server-side environment variable such as AI21_API_KEY.
  4. Do not place the key in browser JavaScript, mobile-app code, public repositories, or client-side configuration that users can inspect.

API requests use bearer-style authentication. In production, keep the key in a secrets manager or protected deployment environment, rotate it when necessary, and avoid writing it to application logs.

Choosing the API and model

For a normal text-generation or conversational application, start with the Jamba Chat Completions API. A request supplies a current Jamba model identifier and a messages array. The research for this guide uses jamba-mini in examples; check the current AI21 model documentation and the models available to your account before deploying.

Use the chat-completions service when your application mainly needs generated text, multi-turn context, streaming, or tool calls. Use the library and retrieval services when answers must be grounded in uploaded documents. Use Maestro when the workflow needs planning, multiple tools, web search, file search, requirements, polling, or validated output.

These are related capabilities, but Maestro is a higher-level orchestration service rather than simply another chat-completion option.

Making your first request

The following raw HTTP example sends a chat-completion request. It uses the current hosted API pattern and keeps the API key on the server.

set -euo pipefail

: "${AI21_API_KEY:?Set AI21_API_KEY first}"

curl --fail-with-body --silent --show-error 
  --request POST 
  --url "https://api.ai21.com/studio/v1/chat/completions" 
  --header "Authorization: Bearer ${AI21_API_KEY}" 
  --header "Content-Type: application/json" 
  --data '{
    "model": "jamba-mini",
    "messages": [
      {
        "role": "system",
        "content": "You are a concise technical assistant."
      },
      {
        "role": "user",
        "content": "Explain what an API key is in one sentence."
      }
    ],
    "temperature": 0.2,
    "max_tokens": 100
  }'

The request contains three important parts: the model, the conversation messages, and generation settings. temperature influences variation, while max_tokens places an upper bound on the generated response. Exact model availability and account permissions should be checked in the current AI21 documentation.

Python SDK example

AI21's official Python package is installed with pip install ai21. The current SDK pattern uses AI21Client and chat completions.

import os
from ai21 import AI21Client
from ai21.models.chat import ChatMessage

client = AI21Client(api_key=os.environ["AI21_API_KEY"])

response = client.chat.completions.create(
    model="jamba-mini",
    messages=[
        ChatMessage(
            role="system",
            content="You are a concise technical assistant."
        ),
        ChatMessage(
            role="user",
            content="What is an API key?"
        ),
    ],
    temperature=0.2,
    max_tokens=100,
)

print(response.choices[0].message.content)

Understanding the response

A successful chat-completion response contains a choices collection. The generated assistant text is available in the first choice's message content, represented in the Python SDK as response.choices[0].message.content.

Do not assume that every successful HTTP response means the model produced the exact format your application needs. Validate required fields, handle empty or unexpected content, and apply application-level validation when the result is consumed by another system.

For structured data, a prompt can request valid JSON, but ordinary Jamba chat-completion JSON-Schema enforcement was not verified in the supplied research. Maestro's validated-output and requirements features are the clearest documented first-party structured-output capability. If you request JSON from a normal chat completion, parse and validate it in your application rather than trusting the model output blindly.

How pricing generally works

Hosted AI21 model access uses usage-based pricing, generally based on token consumption. The actual commercial terms can depend on the model, account, plan, deployment, and contract. Enterprise and private-deployment arrangements use negotiated or account-specific terms.

No universal price, quota, or rate limit should be assumed from the API documentation alone. Check the current AI21 account and pricing documentation before estimating costs. Track input and output usage, select an appropriate model, limit unnecessary context, and set application budgets where available.

Core capabilities available to developers

Chat completions and conversations

Jamba chat completions support system instructions, user and assistant messages, and multi-turn conversations. Your application is responsible for deciding how much conversation history to send and how to manage context for a particular use case.

Streaming responses

Streaming returns generated output incrementally instead of waiting for the complete response. This can make a user interface feel more responsive and lets a server begin forwarding text sooner. It does not remove the need for timeouts, error handling, or final-response validation.

stream = client.chat.completions.create(
    model="jamba-mini",
    messages=[
        ChatMessage(
            role="user",
            content="Write a short explanation of streaming responses."
        )
    ],
    stream=True,
)

for chunk in stream:
    content = getattr(chunk.choices[0].delta, "content", None)
    if content:
        print(content, end="", flush=True)
print()

Tool and function calling

Tool calling lets the model request an action defined by your application. A tool definition normally includes a function name, a description, and a JSON parameter schema. The model does not perform the external action itself: your application receives the tool call, validates its arguments, executes the function, and sends the result back in a subsequent conversation step.

This pattern is useful for controlled access to business systems, calculators, search services, databases, and other application functions. Treat tool arguments as untrusted input and enforce authorization and validation in the application.

Files, file search, and conversational RAG

The AI21 library API supports uploading, listing, retrieving, updating, downloading, and deleting workspace files. Uploaded content can be used with file search or conversational RAG so that responses are grounded in a document collection rather than relying only on the model's general language knowledge.

Maestro also documents file-search tools and file identifiers for agent workflows. File handling is primarily a text and document capability; it should not be confused with verified native image understanding in the hosted Jamba API.

Maestro orchestration

AI21 Maestro is intended for higher-level workflows involving planning, tools, web search, file search, requirements, polling, and validated output. It can be considered when a single chat-completion request has become difficult to manage because the application needs multiple coordinated steps.

Fine-tuning

AI21 documents fine-tuning approaches for Jamba models, including full fine-tuning, LoRA, and QLoRA. The exact availability, hosting path, and commercial terms depend on the selected product, deployment, and account, so do not assume that every method is available in every hosted account.

SDKs and developer tools

AI21 provides official Python and TypeScript SDKs. They cover core chat completions and also provide support for capabilities such as streaming, Maestro, library files, and conversational RAG. Raw HTTP remains a practical option for languages without an official SDK or for teams that want direct control over requests.

The AI21 Studio console provides a developer entry point and playground-style environment for trying the platform. The official documentation includes API references, authentication guidance, SDK documentation, function calling, file libraries, file search, web search, fine-tuning, usage and cost information, and rate-limit guidance.

A safer production request pattern

A production integration should add bounded timeouts, controlled retries, status-code handling, request logging without secrets, and output validation. The Python SDK documents configurable timeouts and automatic retries. Applications should still use bounded exponential backoff and treat HTTP 429 responses as throttling signals.

import os
from ai21 import AI21Client
from ai21.models.chat import ChatMessage

client = AI21Client(
    api_key=os.environ["AI21_API_KEY"],
    timeout=60,
    max_retries=2,
)

response = client.chat.completions.create(
    model="jamba-mini",
    messages=[
        ChatMessage(
            role="system",
            content="Return valid JSON only with keys: answer and confidence."
        ),
        ChatMessage(
            role="user",
            content="State whether 2+2 equals 4."
        ),
    ],
    temperature=0,
    max_tokens=100,
)

raw_content = response.choices[0].message.content
print(raw_content)
# Parse and validate raw_content with an application-level JSON validator.

The example requests JSON but deliberately does not treat the model response as trusted structured data. In a real service, parse the content, check its schema and values, and return a controlled error if validation fails.

Limits and production considerations

  • Quotas and rate limits: Limits depend on the account, model, plan, and deployment. The current materials do not establish one universal limit for all users.
  • Latency: No universal latency figure or SLA was verified. Latency varies with model, prompt size, output size, workload, region, and deployment type. Streaming can improve perceived responsiveness.
  • Availability: Model names, features, and entitlements can change. Confirm that the selected model and service are enabled for your account.
  • Data governance: Retention, residency, deletion, logging, and training-use terms can vary by product and deployment. Review the applicable account and enterprise agreement before sending sensitive data.
  • Privacy: AI21 offers deployment choices including AI21 Studio, partner-cloud, AI21-managed private, and self-managed private arrangements. Enterprise buyers should verify the exact controls rather than assuming that all hosted and private options have identical policies.
  • Tool safety: Tool calls can trigger real actions. Validate arguments, enforce user authorization, and require confirmation for irreversible operations.
  • Context management: Long conversations and retrieved documents increase token usage and may affect latency and cost. Select only the context needed for the task.

Advantages and limitations

Advantages

  • A familiar chat-completions interface for Jamba language models.
  • Official Python and TypeScript SDKs alongside direct REST access.
  • Streaming and tool/function calling for interactive applications.
  • Managed file libraries, file search, and conversational RAG for document-grounded workflows.
  • Maestro for planning, orchestration, tools, web search, file search, and validated-output workflows.
  • Deployment choices that can extend beyond a standard shared hosted API.

Limitations

  • Pricing, quotas, and some capabilities are account-, model-, plan-, or contract-dependent.
  • Ordinary chat-completion JSON-Schema enforcement was not verified; application-side validation remains necessary.
  • Native image input was not verified in the current first-party API documentation reviewed.
  • The platform is primarily text and document-oriented rather than a general multimodal media-generation service.
  • Private deployments and some enterprise capabilities may require sales-led or negotiated access.

When AI21 Studio is a good or poor choice

AI21 Studio is a good choice when you need hosted Jamba text generation, a conventional chat API, streaming, tool calls, document retrieval, or an enterprise-oriented path to private deployment. It is also worth evaluating when Maestro's orchestration and validated-output features match the workflow you are building.

It may be a poor choice if your application depends primarily on image, video, or audio generation; verified native image input; a universal fixed-price plan; or a broad consumer-assistant ecosystem. Before committing, test the models and services available to your account, confirm current pricing and limits, and obtain the data-processing terms required for your workload.

Sources 21