Developer platform

Cerebras Inference API

Developer overview for Cerebras, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Cerebras Inference API
SDK support Official Python package: cerebras_cloud_sdk. Official TypeScript/Node.js package: @cerebras/cerebras_cloud_sdk. OpenAI-compatible Python and Node.js clients can also be configured with the Cerebras base URL. Other languages can use HTTP.
Rate limits Limits are model-, organization-, project-, and plan-dependent and are measured by requests and tokens across minute, hour, and day windows. Current documentation gives free-tier examples such as 30 RPM and approximately 60K-64K TPM for several models, wh
Platform

API overview

Endpoints

API access

Base URL https://api.cerebras.ai/v1
Primary API Cerebras Inference API
Pricing

API pricing

Pricing model Free tier, pay-per-token developer access, and custom enterprise pricing

Current public model metadata lists model-specific input/output token prices; GPT OSS 120B is listed at approximately $0.35 per million input tokens and $0.75 per million output tokens. Free access has lower limits; enterprise pricing is custom.

Developer experience

SDKs & usability

SDKs Official Python package: cerebras_cloud_sdk. Official TypeScript/Node.js package: @cerebras/cerebras_cloud_sdk. OpenAI-compatible Python and Node.js clients can also be configured with the Cerebras base URL. Other languages can use HTTP.
Ease of use High for developers familiar with OpenAI-compatible APIs. The platform provides a conventional REST interface, bearer authentication, official Python and TypeScript SDKs, and OpenAI client compatibility.
Documentation Good and technically focused. Documentation covers quickstart, API reference, models, streaming, tool use, structured outputs, rate limits, projects, service tiers, files, and batch processing. Some newer features are preview-only.
Latency Cerebras positions the service for very low-latency, high-throughput inference. Current model documentation lists approximately 2,200 tokens per second for llama3.1-8b and approximately 3,000 tokens per second for gpt-oss-120b. Actual latency depends on m
Features

API capabilities

✓ Streaming
✓ Function calling
✓ File uploads
✓ Fine-tuning
✓ Structured outputs
✓ Playground
Feature notes

The recommended public architecture is the OpenAI-compatible Chat Completions API at /v1/chat/completions, accessed through the official Cerebras Python or TypeScript SDK, direct HTTP, or a compatible OpenAI client configured with the Cerebras base URL. Current platform capabilities include text chat completions, legacy-style text completions, streaming, tool/function calling, strict structured outputs, JSON mode, model listing, usage monitoring, project management, rate-limit headers, service tiers, and batch processing. The Files API is currently documented for Batch API input and is in private preview; it is not a general-purpose retrieval or multimodal file-ingestion API. Fine-tuned models are an enterprise capability rather than a general self-serve feature. Image input is not generally supported by the public shared model offering; current public model metadata reports vision false for the cited production model. No first-party persistent Assistants API or hosted agent runtime is documented in the current public API reference. Agentic applications can be built by combining multi-turn chat, tool calling, external state, and application-side orchestration. Model IDs and capability availability can change, so applications should consult the live model catalog before deployment. Implement bounded retries for transient 429 and 5xx errors, respect rate-limit headers, and set max_completion_tokens appropriately to avoid unnecessary quota estimation.

Examples

API examples

# Basic request
curl --fail-with-body --silent --show-error \
  https://api.cerebras.ai/v1/chat/completions \
  -H "Authorization: Bearer ${CEREBRAS_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-oss-120b",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain server-sent events in one paragraph."}
    ],
    "max_completion_tokens": 200
  }' | jq -r '.choices[0].message.content'

# Streaming request
curl --fail-with-body --silent --show-error \
  https://api.cerebras.ai/v1/chat/completions \
  -H "Authorization: Bearer ${CEREBRAS_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-oss-120b",
    "messages": [{"role": "user", "content": "Give three uses for fast inference."}],
    "stream": true,
    "max_completion_tokens": 200
  }'
# Install: pip install --upgrade cerebras_cloud_sdk
import json
import os
import sys
from cerebras.cloud.sdk import Cerebras

api_key = os.environ.get("CEREBRAS_API_KEY")
if not api_key:
    raise RuntimeError("CEREBRAS_API_KEY is not set")

client = Cerebras(api_key=api_key)

try:
    completion = client.chat.completions.create(
        model="gpt-oss-120b",
        messages=[
            {"role": "system", "content": "You are a helpful technical assistant."},
            {"role": "user", "content": "What is a token bucket rate limiter?"}
        ],
        max_completion_tokens=300,
    )
    print(completion.choices[0].message.content)

    stream = client.chat.completions.create(
        model="gpt-oss-120b",
        messages=[
            {"role": "user", "content": "Explain streaming inference briefly."}
        ],
        stream=True,
        max_completion_tokens=200,
    )
    for chunk in stream:
        content = chunk.choices[0].delta.content or ""
        print(content, end="", flush=True)
    print()

    schema = {
        "type": "object",
        "properties": {
            "summary": {"type": "string"},
            "priority": {"type": "string", "enum": ["low", "medium", "high"]}
        },
        "required": ["summary", "priority"],
        "additionalProperties": False,
    }
    structured = client.chat.completions.create(
        model="gpt-oss-120b",
        messages=[
            {"role": "system", "content": "Return a concise classification."},
            {"role": "user", "content": "Classify this issue: API requests are timing out."}
        ],
        response_format={
            "type": "json_schema",
            "json_schema": {"name": "classification", "strict": True, "schema": schema},
        },
        max_completion_tokens=200,
    )
    print(json.loads(structured.choices[0].message.content))
except Exception as exc:
    print(f"Cerebras API error: {exc}", file=sys.stderr)
    raise
// Install: npm install @cerebras/cerebras_cloud_sdk
import Cerebras from "@cerebras/cerebras_cloud_sdk";

const apiKey = process.env.CEREBRAS_API_KEY;
if (!apiKey) throw new Error("CEREBRAS_API_KEY is not set");

const client = new Cerebras({ apiKey });

async function main() {
  try {
    const completion = await client.chat.completions.create({
      model: "gpt-oss-120b",
      messages: [
        { role: "system", content: "You are a helpful technical assistant." },
        { role: "user", content: "Explain why low latency matters for agents." }
      ],
      max_completion_tokens: 250
    });
    console.log(completion.choices[0]?.message?.content ?? "");

    const stream = await client.chat.completions.create({
      model: "gpt-oss-120b",
      messages: [{ role: "user", content: "List three API reliability practices." }],
      stream: true,
      max_completion_tokens: 200
    });
    for await (const chunk of stream) {
      process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
    }
    process.stdout.write("\n");

    const toolCompletion = await client.chat.completions.create({
      model: "gpt-oss-120b",
      messages: [{ role: "user", content: "What is 27 multiplied by 14?" }],
      tools: [{
        type: "function",
        function: {
          name: "calculate",
          description: "Calculate a basic arithmetic expression.",
          strict: true,
          parameters: {
            type: "object",
            properties: { expression: { type: "string" } },
            required: ["expression"],
            additionalProperties: false
          }
        }
      }],
      max_completion_tokens: 200
    });
    console.log(JSON.stringify(toolCompletion.choices[0]?.message ?? {}));
  } catch (error) {
    console.error("Cerebras API error:", error);
    process.exitCode = 1;
  }
}

main();
<?php
$apiKey = getenv('CEREBRAS_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('CEREBRAS_API_KEY is not set');
}

$url = 'https://api.cerebras.ai/v1/chat/completions';
$payload = [
    'model' => 'gpt-oss-120b',
    'messages' => [
        ['role' => 'system', 'content' => 'You are a concise technical assistant.'],
        ['role' => 'user', 'content' => 'Explain how API retries should be bounded.']
    ],
    'max_completion_tokens' => 250
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
        'Accept: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_CONNECTTIMEOUT => 10,
    CURLOPT_TIMEOUT => 60
]);

$body = curl_exec($ch);
if ($body === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('Network error: ' . $error);
}
$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? $body;
    throw new RuntimeException("Cerebras API HTTP $status: $message");
}

$content = $data['choices'][0]['message']['content'] ?? null;
if ($content === null) {
    throw new RuntimeException('No assistant content in the response');
}
echo $content . PHP_EOL;
?>
Policies

Data & usage

Data training

Cerebras' current terms state that service content is not used for training or fine-tuning models. Usage and diagnostic data may be collected and processed to provide, maintain, monitor, and improve services, subject to the applicable terms and privacy policy. Enterprise customers should confirm any additional contractual controls with Cerebras.

Data retention

Cerebras' privacy policy states that inputs and outputs associated with its training, inference, and chatbot services are not retained, and that associated logs are deleted when no longer necessary to provide the services. Batch-uploaded files are a separate feature and are documented as retained for seven days by default, with configurable expiration between one hour and thirty days. Review current enterprise agreements for workload-specific retention and compliance requirements.

Rate limits

Limits are model-, organization-, project-, and plan-dependent and are measured by requests and tokens across minute, hour, and day windows. Current documentation gives free-tier examples such as 30 RPM and approximately 60K-64K TPM for several models, wh

Developer guide

Cerebras Inference API: Beginner Developer Guide

Cerebras Inference is an OpenAI-compatible developer API for fast hosted language-model inference. This guide explains access, API keys, model selection, chat completions, streaming, tool calling, structured outputs, files and batch processing, SDKs, pricing, rate limits, and production considerations.
Cerebras Inference gives developers a familiar API for building applications that generate text with models hosted on Cerebras infrastructure. It uses API-key authentication, supports official Python and TypeScript SDKs, and provides chat completions, streaming, tool calling, structured outputs, model discovery, project controls, and usage monitoring. Free access, pay-per-token billing, and enterprise arrangements are available, while some file and batch features remain in private preview.
Cerebras Inference is an OpenAI-compatible API for fast hosted language-model generation. Developers can use Python or TypeScript SDKs, chat completions, streaming, tool calling, structured outputs, model discovery, projects, and usage controls. Access includes free testing, pay-per-token billing, and enterprise options. Files and batch processing are preview features, while the public shared API is primarily text-focused and does not document persistent assistants or general image input.

What is the Cerebras Inference API?

Cerebras Inference is a hosted developer API for sending prompts to supported language models and receiving generated text. Its main interface is an OpenAI-compatible HTTP API at https://api.cerebras.ai/v1. The primary endpoint is /chat/completions, where an application sends a list of messages and receives an assistant response.

The service is intended for developers building applications, coding tools, agents, real-time assistants, and other systems that need low-latency, high-throughput text generation. It is not a consumer chatbot platform and does not provide a general-purpose hosted assistant with persistent conversations, built-in web search, or native image generation.

Model identifiers, pricing, capabilities, and deprecation status can change. Before deploying an application, check the current Cerebras model catalog rather than relying on a model name copied from an older example.

Who should use Cerebras Inference?

Cerebras Inference is a good fit for developers who already understand basic HTTP or SDK usage and want a fast hosted language-model endpoint. It is particularly relevant for interactive applications where response latency matters, coding agents, high-throughput workloads, and systems that need OpenAI-compatible request patterns.

It is a less suitable choice if your application depends on native image understanding, image generation, persistent server-side assistants, web browsing, or a broad consumer application ecosystem. Those functions are not documented as general capabilities of the public shared inference service.

Getting access and creating an API key

You need a Cerebras account and an API key. Developer access is available through the Cerebras Cloud environment, with free access subject to lower limits and self-serve pay-per-token access available for higher usage. Enterprise customers can arrange custom capacity and support.

Store the key in an environment variable or secret manager rather than placing it directly in source code. The standard variable used in Cerebras examples is CEREBRAS_API_KEY.

export CEREBRAS_API_KEY="your-api-key"

Requests authenticate with a bearer token. Treat the key as a server-side secret and avoid exposing it in browser code, public repositories, or client applications distributed to users.

Choose a supported model

Cerebras exposes a model catalog that lists supported model IDs and can provide capability, pricing, and status information. Select a model that is currently available for your account and endpoint. The examples below use gpt-oss-120b, but model availability can change.

When selecting a model, check whether it supports the features your application needs. Compatibility can vary for streaming, tool calling, structured outputs, context limits, and other options. Do not assume that every model supports every parameter.

Make your first request

The simplest request is a chat completion. It sends a system instruction that establishes behavior and a user message containing the task. The API returns a JSON response containing one or more choices.

curl --fail-with-body --silent --show-error 
  https://api.cerebras.ai/v1/chat/completions 
  -H "Authorization: Bearer ${CEREBRAS_API_KEY}" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "gpt-oss-120b",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain server-sent events in one paragraph."}
    ],
    "max_completion_tokens": 200
  }'

The messages array represents the conversation. For a multi-turn exchange, your application normally sends the relevant previous messages again; ordinary chat completions do not require a separate server-side conversation object.

Use the official SDKs

Cerebras provides an official Python package named cerebras_cloud_sdk and an official TypeScript/Node.js package named @cerebras/cerebras_cloud_sdk. Direct HTTP requests are also practical because the API uses conventional JSON and bearer-token authentication. Existing OpenAI-compatible clients can generally be configured with the Cerebras base URL, but individual parameters and feature support should still be checked.

pip install --upgrade cerebras_cloud_sdk
import os
from cerebras.cloud.sdk import Cerebras

client = Cerebras(api_key=os.environ["CEREBRAS_API_KEY"])

completion = client.chat.completions.create(
    model="gpt-oss-120b",
    messages=[
        {"role": "system", "content": "You are a helpful technical assistant."},
        {"role": "user", "content": "What is a token bucket rate limiter?"}
    ],
    max_completion_tokens=300,
)

print(completion.choices[0].message.content)

Understand the response

A successful chat-completion response contains a choices array. The generated assistant text is normally found at choices[0].message.content. Depending on the request and model, the response can also include tool calls, finish information, model identifiers, and usage data.

Do not assume that every response contains ordinary text. A tool-calling response may ask your application to invoke a function instead, and streamed responses deliver partial deltas rather than one completed message.

How Cerebras Inference pricing works

Cerebras offers free access with lower limits, self-serve pay-per-token access, and custom enterprise arrangements. Self-serve billing is based on model-specific input and output token rates. Public model metadata has listed an example of approximately $0.35 per million input tokens and $0.75 per million output tokens for GPT OSS 120B, but rates should be verified in the live model catalog and pricing documentation before being used in a budget or contract.

Enterprise agreements can include custom capacity, dedicated priority, custom model weights or fine-tuned models, service-level commitments, and dedicated support. Your actual limits and available models can depend on your organization, project, plan, and model.

Stream responses for interactive applications

Streaming sends partial output as it is generated through server-sent events. Set stream to true and append the text from each delta to the interface. Streaming is useful for chat interfaces because users can see the beginning of an answer without waiting for the complete response.

curl --fail-with-body --silent --show-error 
  https://api.cerebras.ai/v1/chat/completions 
  -H "Authorization: Bearer ${CEREBRAS_API_KEY}" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "gpt-oss-120b",
    "messages": [{"role": "user", "content": "Give three uses for fast inference."}],
    "stream": true,
    "max_completion_tokens": 200
  }'

Usage and timing information may be supplied only in the final stream chunk. Code consuming a stream should therefore tolerate chunks that contain only partial content or metadata.

Build tools and function calling

Tool calling lets a model request an operation that your application performs. For example, an application can expose a calculator, database lookup, search service, or business action through a JSON Schema. The model chooses whether to call the tool, your code executes it, and the result is sent back in a later model request.

This is client-orchestrated tool use, not a persistent hosted agent runtime. Your application remains responsible for authorization, validation, execution, error handling, and any external state.

const toolCompletion = await client.chat.completions.create({
  model: "gpt-oss-120b",
  messages: [{ role: "user", content: "What is 27 multiplied by 14?" }],
  tools: [{
    type: "function",
    function: {
      name: "calculate",
      description: "Calculate a basic arithmetic expression.",
      strict: true,
      parameters: {
        type: "object",
        properties: { expression: { type: "string" } },
        required: ["expression"],
        additionalProperties: false
      }
    }
  }],
  max_completion_tokens: 200
});

After receiving a tool call, validate its arguments before running the operation. Never allow a model-generated request to bypass the permissions and safety checks that your application would apply to a normal user request.

Use structured outputs when code needs predictable JSON

Structured outputs allow you to provide a JSON Schema and request a response that follows it. With strict mode enabled, the model is constrained by the supplied schema when the selected model supports the feature. This is more reliable for application data than asking for JSON in ordinary prose.

schema = {
    "type": "object",
    "properties": {
        "summary": {"type": "string"},
        "priority": {"type": "string", "enum": ["low", "medium", "high"]}
    },
    "required": ["summary", "priority"],
    "additionalProperties": False,
}

result = client.chat.completions.create(
    model="gpt-oss-120b",
    messages=[
        {"role": "user", "content": "Classify: API requests are timing out."}
    ],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "classification",
            "strict": True,
            "schema": schema,
        },
    },
    max_completion_tokens=200,
)

JSON mode is also documented, but structured outputs are preferable when the application needs a defined schema. Confirm support in the current model catalog before depending on either feature.

Files and batch processing

The Files API and Batch API are documented as private-preview features for asynchronous processing. Files are uploaded as JSONL input for batch jobs; they are not a general-purpose retrieval system or a documented multimodal attachment workflow.

The documented file limit is 200 MB. Files are retained for seven days by default, with configurable expiration from one hour to thirty days. Batch SDK support was not available in the documented private preview, so direct HTTP requests may be required. Confirm preview access and current behavior before designing a production workflow around these features.

Playground and developer tools

The Cerebras Cloud console provides a playground and project-management features. Projects can group API keys, members, usage analytics, and project-level rate limits, which is useful for separating development, testing, and production workloads.

The API also provides model discovery, usage monitoring, rate-limit information, and service tiers. Where enabled for an organization, flex traffic is lower priority while priority traffic receives higher priority. Availability depends on the organization and current service configuration.

Limits and production considerations

Limits are model-, organization-, project-, and plan-dependent. They can apply across request and token quotas measured over minute, hour, and day windows. Documentation gives free-tier examples such as 30 requests per minute and approximately 60,000 to 64,000 tokens per minute for several models, but the exact values should be read from the Cerebras console for the relevant account and model.

Response headers expose remaining quota and reset information. A 429 response indicates rate limiting. Cerebras recommends setting a realistic max_completion_tokens value because token quotas may be evaluated using estimated maximum consumption.

  • Store API keys in a secret manager or environment variable.
  • Handle 429, 5xx, timeout, and network errors with bounded exponential backoff.
  • Respect rate-limit headers instead of retrying continuously.
  • Track request IDs, usage, latency, and estimated cost.
  • Use separate projects for development, testing, and production.
  • Check the model catalog for deprecations before hard-coding model IDs.
  • Validate tool arguments and structured responses before using them in application logic.

Privacy and data handling

Cerebras states that it does not retain inputs and outputs associated with its training, inference, and chatbot services, deleting related logs when they are no longer necessary to provide the services. Its terms state that service content is not used by Cerebras for training or fine-tuning models.

Usage and diagnostic data may still be collected to provide, maintain, monitor, improve, analyze, and develop services. Batch-uploaded files have separate retention behavior. Review the current privacy policy, terms, and any enterprise agreement before sending regulated or sensitive data.

Advantages and limitations

AreaWhat to expect
SpeedCerebras positions the service for very low-latency, high-throughput inference.
IntegrationOpenAI-compatible HTTP patterns and official Python and TypeScript SDKs reduce migration effort.
Developer featuresChat completions, streaming, tool calling, structured outputs, model discovery, projects, and usage controls are available.
FilesFiles and batch processing are preview features intended for asynchronous JSONL processing, not general multimodal uploads.
Multimodal supportThe public shared offering is primarily text-oriented; general image input is not documented.
Hosted agentsNo first-party persistent Assistants API or hosted agent runtime is documented in the public API reference.

When Cerebras Inference is a good or poor choice

Choose Cerebras Inference when low latency, high throughput, OpenAI-compatible integration, tool calling, or fast coding and agent workflows are central requirements. It is also worth evaluating when you want a developer API with free access for testing and a path to pay-per-token or enterprise capacity.

Look elsewhere, or add your own orchestration layer, when you require native image understanding, image generation, web search, persistent assistant threads, consumer mobile applications, or a mature hosted agent runtime. Cerebras can still serve as the text-generation component in those systems, but the missing functions would need to come from your application or other services.

Sources 20