Developer platform

xAI API

Developer overview for xAI, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Responses API
SDK support Official Python SDK xai-sdk with native gRPC support; OpenAI Python and JavaScript SDK compatibility through the xAI base URL; JavaScript support through @ai-sdk/xai and Vercel AI SDK; gRPC and raw HTTP are also supported.
Rate limits Per-model limits use requests per second and tokens per minute for language and embedding models. Limits scale with spend-based tiers: Tier 0 at $0, Tier 1 at $50, Tier 2 at $250, Tier 3 at $1,000, Tier 4 at $5,000, with Enterprise available on request. H
Platform

API overview

Endpoints

API access

Base URL https://api.x.ai
Primary API Responses API
Pricing

API pricing

Pricing model Usage-based pricing billed by tokens, tool invocations, images, video duration, audio duration or characters depending on the service; prepaid credits and enterprise invoicing are available.

The documented Grok 4.7 price is $2 per 1M short-context input tokens and $6 per 1M output tokens. Long-context, cached-input, reasoning, media, voice and server-side tool prices vary by model and operation.

Developer experience

SDKs & usability

SDKs Official Python SDK xai-sdk with native gRPC support; OpenAI Python and JavaScript SDK compatibility through the xAI base URL; JavaScript support through @ai-sdk/xai and Vercel AI SDK; gRPC and raw HTTP are also supported.
Ease of use High for developers familiar with OpenAI-compatible APIs. API keys are created in the Console, the REST API follows familiar schemas, and official or compatible SDK options are available.
Documentation Strong and actively maintained, with quickstarts, model catalog, REST and gRPC references, capability guides, pricing, rate limits, security FAQs and code examples. Some APIs and integrations are marked early access or may change.
Latency Variable by model, context length, reasoning configuration, output size and tool usage. Streaming is recommended for interactive and agentic requests. Responses API WebSocket mode is available for repeated turns and can reduce connection overhead.
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

The current platform includes Responses API, Chat Completions, image, video, voice, files, batches, models, management and gRPC interfaces. Responses API is the preferred architecture for new agentic applications and supports multi-turn state, streaming, server-side tools and custom function calling. Chat Completions remains supported for conventional conversational workloads. Structured Outputs supports JSON Schema and JSON object modes. Web Search and X Search provide current information when enabled; models do not automatically have real-time knowledge without search tools. Image input is available on vision-capable models, with documented JPEG and PNG support and a maximum image size of 20 MiB. Files can be uploaded and attached to document-understanding workflows, with a documented maximum file size of 512 MB. The Console provides API keys, billing, usage monitoring, model access and a Playground. Enterprise options include SSO, audit logging, monthly invoicing, data residency, custom rate limits, dedicated support and optional Zero Data Retention. Fine-tuning was not confirmed in the current first-party developer documentation reviewed and is therefore left unknown.

Examples

API examples

curl --fail-with-body https://api.x.ai/v1/responses \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-4.7",
    "input": [
      {
        "role": "system",
        "content": "You are a concise technical assistant."
      },
      {
        "role": "user",
        "content": "Return a JSON object describing three benefits of streaming APIs."
      }
    ],
    "text": {
      "format": {
        "type": "json_schema",
        "name": "benefits",
        "schema": {
          "type": "object",
          "properties": {
            "benefits": {
              "type": "array",
              "items": {"type": "string"}
            }
          },
          "required": ["benefits"],
          "additionalProperties": false
        },
        "strict": true
      }
    }
  }' | jq -r '.output[]?.content[]?.text // empty'

# Streaming example
curl -N https://api.x.ai/v1/responses \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-4.7",
    "input": "Explain server-sent events in two sentences.",
    "stream": true
  }'
# Install: pip install xai-sdk pydantic
import json
import os
from pydantic import BaseModel
from xai_sdk import Client
from xai_sdk.chat import system, user

class Answer(BaseModel):
    summary: str
    points: list[str]

api_key = os.environ["XAI_API_KEY"]
client = Client(api_key=api_key)

chat = client.chat.create(
    model="grok-4.7",
    response_format=Answer,
)
chat.append(system("You are a concise technical assistant."))
chat.append(user("Explain why streaming is useful for an API client."))

response = chat.sample()
result = Answer.model_validate_json(response.content)
print(result.summary)
for point in result.points:
    print(f"- {point}")

# Multi-turn streaming conversation
chat.append(user("Now explain the main tradeoff."))
for accumulated, chunk in chat.stream():
    if chunk.content:
        print(chunk.content, end="", flush=True)
print()

# Client-side function calling
from xai_sdk.tools import function

weather_tool = function(
    name="get_weather",
    description="Get the current weather for a city.",
    parameters={
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
        "additionalProperties": False,
    },
)

tool_chat = client.chat.create(model="grok-4.7", tools=[weather_tool])
tool_chat.append(user("What is the weather in Seattle?"))
tool_response = tool_chat.sample()
for call in tool_response.tool_calls:
    print(call.function.name, call.function.arguments)
// Install: npm install openai
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.XAI_API_KEY,
  baseURL: "https://api.x.ai/v1",
});

async function main() {
  try {
    const response = await client.responses.create({
      model: "grok-4.7",
      input: [
        {
          role: "system",
          content: "You are a concise technical assistant.",
        },
        {
          role: "user",
          content: "Explain the value of streaming APIs in two sentences.",
        },
      ],
    });

    console.log(response.output_text);

    const stream = await client.responses.create({
      model: "grok-4.7",
      input: "Give me three practical API retry tips.",
      stream: true,
    });

    for await (const event of stream) {
      if (event.type === "response.output_text.delta") {
        process.stdout.write(event.delta);
      }
    }
    process.stdout.write("\n");

    const toolResponse = await client.responses.create({
      model: "grok-4.7",
      input: "Find the weather for Seattle using the available function.",
      tools: [
        {
          type: "function",
          name: "get_weather",
          description: "Get the current weather for a city.",
          parameters: {
            type: "object",
            properties: { city: { type: "string" } },
            required: ["city"],
            additionalProperties: false,
          },
        },
      ],
    });

    for (const item of toolResponse.output ?? []) {
      if (item.type === "function_call") {
        console.log(item.name, item.arguments);
      }
    }
  } catch (error) {
    console.error(error);
    process.exitCode = 1;
  }
}

main();
<?php
$apiKey = getenv('XAI_API_KEY');
if (!$apiKey) {
    fwrite(STDERR, "XAI_API_KEY is not set\n");
    exit(1);
}

$url = 'https://api.x.ai/v1/responses';
$payload = [
    'model' => 'grok-4.7',
    'input' => [
        [
            'role' => 'system',
            'content' => 'You are a concise technical assistant.'
        ],
        [
            'role' => 'user',
            'content' => 'Return a short explanation of why API streaming is useful.'
        ]
    ]
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
        'Accept: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_CONNECTTIMEOUT => 15,
    CURLOPT_TIMEOUT => 120
]);

$body = curl_exec($ch);
if ($body === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('cURL error: ' . $error);
}

$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? $body;
    throw new RuntimeException('xAI API error (' . $status . '): ' . $message);
}

$text = $data['output_text'] ?? '';
if ($text === '' && isset($data['output'])) {
    foreach ($data['output'] as $item) {
        foreach ($item['content'] ?? [] as $content) {
            if (($content['type'] ?? '') === 'output_text') {
                $text .= $content['text'] ?? '';
            }
        }
    }
}

echo $text . PHP_EOL;
Policies

Data & usage

Data training

xAI states that it does not train on customer API inputs or outputs without explicit permission. Business and enterprise data is not used to train models under the enterprise FAQ, although xAI may offer specific arrangements involving permission to train on business data. Consumer Grok policies are separate and should not be applied to API customers.

Data retention

By default, API requests and responses are stored on xAI servers for 30 days for abuse auditing and are encrypted at rest, after which the content is automatically deleted. Eligible teams can enable team-level Zero Data Retention, which prevents prompts and outputs from being persisted. ZDR disables storage-dependent capabilities including stateful Responses API, Files, Collections and Batch API. Enterprise contracts or data-processing terms may define additional controls.

Rate limits

Per-model limits use requests per second and tokens per minute for language and embedding models. Limits scale with spend-based tiers: Tier 0 at $0, Tier 1 at $50, Tier 2 at $250, Tier 3 at $1,000, Tier 4 at $5,000, with Enterprise available on request. H

Developer guide

xAI API: Beginner's Guide to Grok, Pricing, Tools and SDKs

The xAI API gives developers programmatic access to Grok text and reasoning models, agentic Responses API workflows, search tools, file analysis, structured outputs, image and video services, voice capabilities, and custom function calling through an OpenAI-compatible platform.
The xAI API is designed for developers building applications with Grok rather than using Grok only through a chat interface. It provides a recommended Responses API for new agentic applications, a supported Chat Completions interface for conventional conversational workloads, usage-based pricing, streaming, tool calling, file uploads, structured JSON output, and official or compatible SDKs. This guide explains how to obtain credentials, make a first request, choose an API pattern, estimate costs, and account for operational and privacy limitations.
The xAI API provides usage-priced access to Grok text and reasoning models, agentic Responses API workflows, search, files, structured JSON output, custom functions, image and video services, voice, and other developer tools. It supports OpenAI-compatible REST requests, the official Python SDK, JavaScript integrations, streaming, a Playground, spend-based rate-limit tiers, and optional Zero Data Retention.

What is the xAI API and when should you use it?

The xAI API is a developer platform for integrating Grok models and related services into applications. Instead of sending prompts manually through grok.com, your application sends authenticated HTTP requests to xAI and receives generated text, structured data, tool calls, or media results.

It is suitable for chat applications, document analysis, research assistants, coding tools, automated workflows, customer-support systems, and agentic applications that need search or external functions. The platform also exposes image generation, video generation, voice, file, batch, model-management, and related endpoints.

For new applications that need tools, multi-turn state, or agent-like behavior, xAI recommends the Responses API. Chat Completions remains useful for straightforward text generation, manually managed conversation histories, and existing applications built around the OpenAI-compatible chat schema.

How to get xAI API access

Create an account in the xAI Console, add prepaid credits or arrange enterprise billing, and create an API key. Treat the key like a password: store it in a server-side secret manager or environment variable, not in browser code, mobile application bundles, source-control repositories, or publicly shared examples.

export XAI_API_KEY="your_api_key"

Requests use bearer-token authentication and the primary API host is https://api.x.ai. Versioned inference endpoints normally use the /v1 path, including /v1/responses and /v1/chat/completions.

Choosing an API and model

Use the Responses API for new agentic work. It can represent generated text, tool calls, search activity, and other response items, and supports streaming, custom functions, server-side tools, and multi-turn continuation through previous_response_id. Stateful features are unavailable when Zero Data Retention is enabled.

Choose Chat Completions when your application already has a conventional sequence of roles and messages or when you want to preserve an OpenAI-compatible integration. The exact model catalog and aliases can change, so check xAI's current Models documentation before hard-coding a dated identifier. The supplied current documentation lists grok-4.7 as a flagship text model with a 500,000-token context window.

Make a first request with the Responses API

The following request sends a simple text prompt. The input can be a string for a basic request or a structured collection of messages for more control.

curl --fail-with-body https://api.x.ai/v1/responses 
  -H "Authorization: Bearer $XAI_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "grok-4.7",
    "input": "Explain server-sent events in two sentences."
  }'

The response contains generated output and usage information. Applications should read the returned output rather than assume that a fixed array position or a single text field will always be present. For many SDK integrations, a convenience property such as output_text provides the combined text.

Use an SDK instead of raw HTTP

xAI provides the xai-sdk Python package with native gRPC support. The REST API is also compatible with the OpenAI Python and JavaScript SDKs when the base URL is changed to xAI's endpoint. JavaScript developers can additionally use the @ai-sdk/xai provider with the Vercel AI SDK.

This JavaScript example uses the current OpenAI-compatible client style and the Responses API:

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.XAI_API_KEY,
  baseURL: "https://api.x.ai/v1",
});

const response = await client.responses.create({
  model: "grok-4.7",
  input: [
    { role: "system", content: "You are a concise technical assistant." },
    { role: "user", content: "List three benefits of API streaming." }
  ]
});

console.log(response.output_text);

Use the official Python SDK when you want its higher-level chat interface, typed response formats, or native gRPC behavior. Use raw HTTP when you need direct control over requests or are integrating from a language without a suitable SDK.

Stream responses for interactive applications

Streaming sends partial results as they become available instead of waiting for the complete response. xAI supports server-sent events (SSE) when stream is set to true. This is useful for chat interfaces and long-running reasoning or tool workflows because users can see progress sooner.

curl -N https://api.x.ai/v1/responses 
  -H "Authorization: Bearer $XAI_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "grok-4.7",
    "input": "Give me three practical API retry tips.",
    "stream": true
  }'

JavaScript SDK streams expose events such as response.output_text.delta. Your client should handle completion events, errors, and interruptions rather than treating every event as text. For repeated sequential Responses API turns, xAI also provides WebSocket mode, which can reduce connection overhead.

How xAI API pricing works

xAI uses usage-based pricing rather than a single unlimited API subscription. Text usage can be billed by input tokens, cached-input tokens, reasoning tokens, and output tokens. Prices vary by model and can differ for short and long contexts.

The supplied documentation lists grok-4.7 at $2 per 1 million short-context input tokens and $6 per 1 million output tokens. Long-context requests, cached input, reasoning, and other model families have separate rates. Image, video, voice, speech-to-text, text-to-speech, and server-side tools can use different pricing units.

Responses can include actual request-cost information in usage metadata through cost_in_usd_ticks. Include that data in application monitoring so you can compare cost by user, feature, model, and workflow. Tool calls may increase both latency and cost because a single user request can trigger several model or service operations.

Capabilities available through the platform

  • Text and reasoning: Generate answers, summaries, code, classifications, and other text outputs with Grok models.
  • Search: Use server-side Web Search and X Search when an application needs current information. Models do not automatically have real-time knowledge without an enabled search tool.
  • Vision and files: Vision-capable models can accept image inputs, including supported JPEG and PNG images. The documentation lists a 20 MiB maximum image size. The Files API supports document uploads, with a documented maximum file size of 512 MB.
  • Media and voice: Separate services provide image generation, video generation, voice, speech-to-text, and text-to-speech capabilities.
  • Structured output: Request JSON Schema-constrained output or JSON-object output when downstream code needs predictable fields.
  • Agentic workflows: Combine model responses with search, code execution, file search, collections, MCP integrations, and custom functions.

Use function calling to connect application logic

Function calling lets a model request an operation that your application implements. For example, the model can produce a structured request for get_weather; your server validates the arguments, calls a weather service, and sends the result back for the next model turn. The model does not execute your private function automatically.

const response = await client.responses.create({
  model: "grok-4.7",
  input: "Find the weather for Seattle using the available function.",
  tools: [{
    type: "function",
    name: "get_weather",
    description: "Get the current weather for a city.",
    parameters: {
      type: "object",
      properties: { city: { type: "string" } },
      required: ["city"],
      additionalProperties: false
    }
  }]
});

for (const item of response.output ?? []) {
  if (item.type === "function_call") {
    console.log(item.name, item.arguments);
  }
}

Define narrow tools with explicit JSON Schema, validate every argument on your server, apply authorization checks, and do not execute arbitrary model-generated commands. xAI also offers server-side tools such as Web Search, X Search, code execution, collections search, attachment search, image generation, and MCP integrations.

Structured outputs and file workflows

Structured Outputs can constrain a response to a supplied JSON Schema. The REST API supports formats including json_schema and json_object. This is useful when your application needs to store results in a database or pass them to another program, but your code should still handle refusal, incomplete output, validation errors, and unexpected content.

Files can be uploaded and attached to document-understanding workflows or used with collections for semantic search. Supported document types include common text, code, CSV, JSON, PDF, and other formats. File handling and collection features depend on storage on xAI's side and therefore are not available when Zero Data Retention disables storage-dependent capabilities.

Playground and developer tools

The xAI Console provides API-key management, billing, usage monitoring, team administration, and a Playground for testing prompts and models before integrating them into an application. The platform also provides REST and gRPC interfaces, model documentation, capability guides, and API references.

Use the Playground for quick experiments, but test production behavior through the same SDK or HTTP path your application will use. Record the selected model, prompt or system instructions, tool definitions, input size, output size, latency, and cost so that changes can be evaluated consistently.

Rate limits, reliability, and privacy

Language and embedding rate limits use requests per second and tokens per minute, assigned per model and team. The documented spend-based tiers begin at Tier 0 with $0 cumulative spend, followed by Tier 1 at $50, Tier 2 at $250, Tier 3 at $1,000, and Tier 4 at $5,000; Enterprise limits are available by request. Media and voice services use operation-specific request, concurrency, or session limits. Exceeding a limit returns HTTP 429.

Use exponential backoff for transient rate-limit responses, add request timeouts, and avoid concentrating a full minute of traffic into one second. Latency varies with model, context length, reasoning configuration, output size, and tool usage. Streaming is generally preferable for interactive or agentic operations.

xAI states that it does not train on customer API inputs or outputs without explicit permission. By default, API requests and responses may be retained for 30 days for abuse auditing and are encrypted at rest. Eligible teams can enable Zero Data Retention, which prevents prompt and output persistence but disables stateful Responses, Files, Collections, and Batch operations. Enterprise customers may obtain additional controls such as SSO, audit logging, invoicing, data residency, custom limits, and dedicated support.

When xAI API is a good or poor choice

The xAI API is a good fit when an application needs Grok models, current information through Web Search or X Search, tool-enabled workflows, multimodal services, file analysis, or an OpenAI-compatible integration path. Its Responses API is particularly relevant for applications that combine model output with functions, search, files, or other agentic steps.

It may be a poor fit when your workload requires fixed, easily predictable costs despite variable tool and media usage, strict availability of every feature under Zero Data Retention, or limits and model behavior that never change. Search results and generated content still require application-level verification. Before production deployment, confirm the current model catalog, prices, rate limits, regional availability, retention terms, and the capabilities of the specific model you plan to use.

Sources 18