What is the Cerebras Inference API?
Cerebras Inference is a hosted developer API for sending prompts to supported language models and receiving generated text. Its main interface is an OpenAI-compatible HTTP API at https://api.cerebras.ai/v1. The primary endpoint is /chat/completions, where an application sends a list of messages and receives an assistant response.
The service is intended for developers building applications, coding tools, agents, real-time assistants, and other systems that need low-latency, high-throughput text generation. It is not a consumer chatbot platform and does not provide a general-purpose hosted assistant with persistent conversations, built-in web search, or native image generation.
Model identifiers, pricing, capabilities, and deprecation status can change. Before deploying an application, check the current Cerebras model catalog rather than relying on a model name copied from an older example.
Who should use Cerebras Inference?
Cerebras Inference is a good fit for developers who already understand basic HTTP or SDK usage and want a fast hosted language-model endpoint. It is particularly relevant for interactive applications where response latency matters, coding agents, high-throughput workloads, and systems that need OpenAI-compatible request patterns.
It is a less suitable choice if your application depends on native image understanding, image generation, persistent server-side assistants, web browsing, or a broad consumer application ecosystem. Those functions are not documented as general capabilities of the public shared inference service.
Getting access and creating an API key
You need a Cerebras account and an API key. Developer access is available through the Cerebras Cloud environment, with free access subject to lower limits and self-serve pay-per-token access available for higher usage. Enterprise customers can arrange custom capacity and support.
Store the key in an environment variable or secret manager rather than placing it directly in source code. The standard variable used in Cerebras examples is CEREBRAS_API_KEY.
export CEREBRAS_API_KEY="your-api-key"Requests authenticate with a bearer token. Treat the key as a server-side secret and avoid exposing it in browser code, public repositories, or client applications distributed to users.
Choose a supported model
Cerebras exposes a model catalog that lists supported model IDs and can provide capability, pricing, and status information. Select a model that is currently available for your account and endpoint. The examples below use gpt-oss-120b, but model availability can change.
When selecting a model, check whether it supports the features your application needs. Compatibility can vary for streaming, tool calling, structured outputs, context limits, and other options. Do not assume that every model supports every parameter.
Make your first request
The simplest request is a chat completion. It sends a system instruction that establishes behavior and a user message containing the task. The API returns a JSON response containing one or more choices.
curl --fail-with-body --silent --show-error
https://api.cerebras.ai/v1/chat/completions
-H "Authorization: Bearer ${CEREBRAS_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "gpt-oss-120b",
"messages": [
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain server-sent events in one paragraph."}
],
"max_completion_tokens": 200
}'The messages array represents the conversation. For a multi-turn exchange, your application normally sends the relevant previous messages again; ordinary chat completions do not require a separate server-side conversation object.
Use the official SDKs
Cerebras provides an official Python package named cerebras_cloud_sdk and an official TypeScript/Node.js package named @cerebras/cerebras_cloud_sdk. Direct HTTP requests are also practical because the API uses conventional JSON and bearer-token authentication. Existing OpenAI-compatible clients can generally be configured with the Cerebras base URL, but individual parameters and feature support should still be checked.
pip install --upgrade cerebras_cloud_sdkimport os
from cerebras.cloud.sdk import Cerebras
client = Cerebras(api_key=os.environ["CEREBRAS_API_KEY"])
completion = client.chat.completions.create(
model="gpt-oss-120b",
messages=[
{"role": "system", "content": "You are a helpful technical assistant."},
{"role": "user", "content": "What is a token bucket rate limiter?"}
],
max_completion_tokens=300,
)
print(completion.choices[0].message.content)Understand the response
A successful chat-completion response contains a choices array. The generated assistant text is normally found at choices[0].message.content. Depending on the request and model, the response can also include tool calls, finish information, model identifiers, and usage data.
Do not assume that every response contains ordinary text. A tool-calling response may ask your application to invoke a function instead, and streamed responses deliver partial deltas rather than one completed message.
How Cerebras Inference pricing works
Cerebras offers free access with lower limits, self-serve pay-per-token access, and custom enterprise arrangements. Self-serve billing is based on model-specific input and output token rates. Public model metadata has listed an example of approximately $0.35 per million input tokens and $0.75 per million output tokens for GPT OSS 120B, but rates should be verified in the live model catalog and pricing documentation before being used in a budget or contract.
Enterprise agreements can include custom capacity, dedicated priority, custom model weights or fine-tuned models, service-level commitments, and dedicated support. Your actual limits and available models can depend on your organization, project, plan, and model.
Stream responses for interactive applications
Streaming sends partial output as it is generated through server-sent events. Set stream to true and append the text from each delta to the interface. Streaming is useful for chat interfaces because users can see the beginning of an answer without waiting for the complete response.
curl --fail-with-body --silent --show-error
https://api.cerebras.ai/v1/chat/completions
-H "Authorization: Bearer ${CEREBRAS_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "gpt-oss-120b",
"messages": [{"role": "user", "content": "Give three uses for fast inference."}],
"stream": true,
"max_completion_tokens": 200
}'Usage and timing information may be supplied only in the final stream chunk. Code consuming a stream should therefore tolerate chunks that contain only partial content or metadata.
Build tools and function calling
Tool calling lets a model request an operation that your application performs. For example, an application can expose a calculator, database lookup, search service, or business action through a JSON Schema. The model chooses whether to call the tool, your code executes it, and the result is sent back in a later model request.
This is client-orchestrated tool use, not a persistent hosted agent runtime. Your application remains responsible for authorization, validation, execution, error handling, and any external state.
const toolCompletion = await client.chat.completions.create({
model: "gpt-oss-120b",
messages: [{ role: "user", content: "What is 27 multiplied by 14?" }],
tools: [{
type: "function",
function: {
name: "calculate",
description: "Calculate a basic arithmetic expression.",
strict: true,
parameters: {
type: "object",
properties: { expression: { type: "string" } },
required: ["expression"],
additionalProperties: false
}
}
}],
max_completion_tokens: 200
});After receiving a tool call, validate its arguments before running the operation. Never allow a model-generated request to bypass the permissions and safety checks that your application would apply to a normal user request.
Use structured outputs when code needs predictable JSON
Structured outputs allow you to provide a JSON Schema and request a response that follows it. With strict mode enabled, the model is constrained by the supplied schema when the selected model supports the feature. This is more reliable for application data than asking for JSON in ordinary prose.
schema = {
"type": "object",
"properties": {
"summary": {"type": "string"},
"priority": {"type": "string", "enum": ["low", "medium", "high"]}
},
"required": ["summary", "priority"],
"additionalProperties": False,
}
result = client.chat.completions.create(
model="gpt-oss-120b",
messages=[
{"role": "user", "content": "Classify: API requests are timing out."}
],
response_format={
"type": "json_schema",
"json_schema": {
"name": "classification",
"strict": True,
"schema": schema,
},
},
max_completion_tokens=200,
)JSON mode is also documented, but structured outputs are preferable when the application needs a defined schema. Confirm support in the current model catalog before depending on either feature.
Files and batch processing
The Files API and Batch API are documented as private-preview features for asynchronous processing. Files are uploaded as JSONL input for batch jobs; they are not a general-purpose retrieval system or a documented multimodal attachment workflow.
The documented file limit is 200 MB. Files are retained for seven days by default, with configurable expiration from one hour to thirty days. Batch SDK support was not available in the documented private preview, so direct HTTP requests may be required. Confirm preview access and current behavior before designing a production workflow around these features.
Playground and developer tools
The Cerebras Cloud console provides a playground and project-management features. Projects can group API keys, members, usage analytics, and project-level rate limits, which is useful for separating development, testing, and production workloads.
The API also provides model discovery, usage monitoring, rate-limit information, and service tiers. Where enabled for an organization, flex traffic is lower priority while priority traffic receives higher priority. Availability depends on the organization and current service configuration.
Limits and production considerations
Limits are model-, organization-, project-, and plan-dependent. They can apply across request and token quotas measured over minute, hour, and day windows. Documentation gives free-tier examples such as 30 requests per minute and approximately 60,000 to 64,000 tokens per minute for several models, but the exact values should be read from the Cerebras console for the relevant account and model.
Response headers expose remaining quota and reset information. A 429 response indicates rate limiting. Cerebras recommends setting a realistic max_completion_tokens value because token quotas may be evaluated using estimated maximum consumption.
- Store API keys in a secret manager or environment variable.
- Handle 429, 5xx, timeout, and network errors with bounded exponential backoff.
- Respect rate-limit headers instead of retrying continuously.
- Track request IDs, usage, latency, and estimated cost.
- Use separate projects for development, testing, and production.
- Check the model catalog for deprecations before hard-coding model IDs.
- Validate tool arguments and structured responses before using them in application logic.
Privacy and data handling
Cerebras states that it does not retain inputs and outputs associated with its training, inference, and chatbot services, deleting related logs when they are no longer necessary to provide the services. Its terms state that service content is not used by Cerebras for training or fine-tuning models.
Usage and diagnostic data may still be collected to provide, maintain, monitor, improve, analyze, and develop services. Batch-uploaded files have separate retention behavior. Review the current privacy policy, terms, and any enterprise agreement before sending regulated or sensitive data.
Advantages and limitations
| Area | What to expect |
|---|---|
| Speed | Cerebras positions the service for very low-latency, high-throughput inference. |
| Integration | OpenAI-compatible HTTP patterns and official Python and TypeScript SDKs reduce migration effort. |
| Developer features | Chat completions, streaming, tool calling, structured outputs, model discovery, projects, and usage controls are available. |
| Files | Files and batch processing are preview features intended for asynchronous JSONL processing, not general multimodal uploads. |
| Multimodal support | The public shared offering is primarily text-oriented; general image input is not documented. |
| Hosted agents | No first-party persistent Assistants API or hosted agent runtime is documented in the public API reference. |
When Cerebras Inference is a good or poor choice
Choose Cerebras Inference when low latency, high throughput, OpenAI-compatible integration, tool calling, or fast coding and agent workflows are central requirements. It is also worth evaluating when you want a developer API with free access for testing and a path to pay-per-token or enterprise capacity.
Look elsewhere, or add your own orchestration layer, when you require native image understanding, image generation, web search, persistent assistant threads, consumer mobile applications, or a mature hosted agent runtime. Cerebras can still serve as the text-generation component in those systems, but the missing functions would need to come from your application or other services.
