What is the xAI API and when should you use it?
The xAI API is a developer platform for integrating Grok models and related services into applications. Instead of sending prompts manually through grok.com, your application sends authenticated HTTP requests to xAI and receives generated text, structured data, tool calls, or media results.
It is suitable for chat applications, document analysis, research assistants, coding tools, automated workflows, customer-support systems, and agentic applications that need search or external functions. The platform also exposes image generation, video generation, voice, file, batch, model-management, and related endpoints.
For new applications that need tools, multi-turn state, or agent-like behavior, xAI recommends the Responses API. Chat Completions remains useful for straightforward text generation, manually managed conversation histories, and existing applications built around the OpenAI-compatible chat schema.
How to get xAI API access
Create an account in the xAI Console, add prepaid credits or arrange enterprise billing, and create an API key. Treat the key like a password: store it in a server-side secret manager or environment variable, not in browser code, mobile application bundles, source-control repositories, or publicly shared examples.
export XAI_API_KEY="your_api_key"Requests use bearer-token authentication and the primary API host is https://api.x.ai. Versioned inference endpoints normally use the /v1 path, including /v1/responses and /v1/chat/completions.
Choosing an API and model
Use the Responses API for new agentic work. It can represent generated text, tool calls, search activity, and other response items, and supports streaming, custom functions, server-side tools, and multi-turn continuation through previous_response_id. Stateful features are unavailable when Zero Data Retention is enabled.
Choose Chat Completions when your application already has a conventional sequence of roles and messages or when you want to preserve an OpenAI-compatible integration. The exact model catalog and aliases can change, so check xAI's current Models documentation before hard-coding a dated identifier. The supplied current documentation lists grok-4.7 as a flagship text model with a 500,000-token context window.
Make a first request with the Responses API
The following request sends a simple text prompt. The input can be a string for a basic request or a structured collection of messages for more control.
curl --fail-with-body https://api.x.ai/v1/responses
-H "Authorization: Bearer $XAI_API_KEY"
-H "Content-Type: application/json"
-d '{
"model": "grok-4.7",
"input": "Explain server-sent events in two sentences."
}'The response contains generated output and usage information. Applications should read the returned output rather than assume that a fixed array position or a single text field will always be present. For many SDK integrations, a convenience property such as output_text provides the combined text.
Use an SDK instead of raw HTTP
xAI provides the xai-sdk Python package with native gRPC support. The REST API is also compatible with the OpenAI Python and JavaScript SDKs when the base URL is changed to xAI's endpoint. JavaScript developers can additionally use the @ai-sdk/xai provider with the Vercel AI SDK.
This JavaScript example uses the current OpenAI-compatible client style and the Responses API:
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.XAI_API_KEY,
baseURL: "https://api.x.ai/v1",
});
const response = await client.responses.create({
model: "grok-4.7",
input: [
{ role: "system", content: "You are a concise technical assistant." },
{ role: "user", content: "List three benefits of API streaming." }
]
});
console.log(response.output_text);Use the official Python SDK when you want its higher-level chat interface, typed response formats, or native gRPC behavior. Use raw HTTP when you need direct control over requests or are integrating from a language without a suitable SDK.
Stream responses for interactive applications
Streaming sends partial results as they become available instead of waiting for the complete response. xAI supports server-sent events (SSE) when stream is set to true. This is useful for chat interfaces and long-running reasoning or tool workflows because users can see progress sooner.
curl -N https://api.x.ai/v1/responses
-H "Authorization: Bearer $XAI_API_KEY"
-H "Content-Type: application/json"
-d '{
"model": "grok-4.7",
"input": "Give me three practical API retry tips.",
"stream": true
}'JavaScript SDK streams expose events such as response.output_text.delta. Your client should handle completion events, errors, and interruptions rather than treating every event as text. For repeated sequential Responses API turns, xAI also provides WebSocket mode, which can reduce connection overhead.
How xAI API pricing works
xAI uses usage-based pricing rather than a single unlimited API subscription. Text usage can be billed by input tokens, cached-input tokens, reasoning tokens, and output tokens. Prices vary by model and can differ for short and long contexts.
The supplied documentation lists grok-4.7 at $2 per 1 million short-context input tokens and $6 per 1 million output tokens. Long-context requests, cached input, reasoning, and other model families have separate rates. Image, video, voice, speech-to-text, text-to-speech, and server-side tools can use different pricing units.
Responses can include actual request-cost information in usage metadata through cost_in_usd_ticks. Include that data in application monitoring so you can compare cost by user, feature, model, and workflow. Tool calls may increase both latency and cost because a single user request can trigger several model or service operations.
Capabilities available through the platform
- Text and reasoning: Generate answers, summaries, code, classifications, and other text outputs with Grok models.
- Search: Use server-side Web Search and X Search when an application needs current information. Models do not automatically have real-time knowledge without an enabled search tool.
- Vision and files: Vision-capable models can accept image inputs, including supported JPEG and PNG images. The documentation lists a 20 MiB maximum image size. The Files API supports document uploads, with a documented maximum file size of 512 MB.
- Media and voice: Separate services provide image generation, video generation, voice, speech-to-text, and text-to-speech capabilities.
- Structured output: Request JSON Schema-constrained output or JSON-object output when downstream code needs predictable fields.
- Agentic workflows: Combine model responses with search, code execution, file search, collections, MCP integrations, and custom functions.
Use function calling to connect application logic
Function calling lets a model request an operation that your application implements. For example, the model can produce a structured request for get_weather; your server validates the arguments, calls a weather service, and sends the result back for the next model turn. The model does not execute your private function automatically.
const response = await client.responses.create({
model: "grok-4.7",
input: "Find the weather for Seattle using the available function.",
tools: [{
type: "function",
name: "get_weather",
description: "Get the current weather for a city.",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
additionalProperties: false
}
}]
});
for (const item of response.output ?? []) {
if (item.type === "function_call") {
console.log(item.name, item.arguments);
}
}Define narrow tools with explicit JSON Schema, validate every argument on your server, apply authorization checks, and do not execute arbitrary model-generated commands. xAI also offers server-side tools such as Web Search, X Search, code execution, collections search, attachment search, image generation, and MCP integrations.
Structured outputs and file workflows
Structured Outputs can constrain a response to a supplied JSON Schema. The REST API supports formats including json_schema and json_object. This is useful when your application needs to store results in a database or pass them to another program, but your code should still handle refusal, incomplete output, validation errors, and unexpected content.
Files can be uploaded and attached to document-understanding workflows or used with collections for semantic search. Supported document types include common text, code, CSV, JSON, PDF, and other formats. File handling and collection features depend on storage on xAI's side and therefore are not available when Zero Data Retention disables storage-dependent capabilities.
Playground and developer tools
The xAI Console provides API-key management, billing, usage monitoring, team administration, and a Playground for testing prompts and models before integrating them into an application. The platform also provides REST and gRPC interfaces, model documentation, capability guides, and API references.
Use the Playground for quick experiments, but test production behavior through the same SDK or HTTP path your application will use. Record the selected model, prompt or system instructions, tool definitions, input size, output size, latency, and cost so that changes can be evaluated consistently.
Rate limits, reliability, and privacy
Language and embedding rate limits use requests per second and tokens per minute, assigned per model and team. The documented spend-based tiers begin at Tier 0 with $0 cumulative spend, followed by Tier 1 at $50, Tier 2 at $250, Tier 3 at $1,000, and Tier 4 at $5,000; Enterprise limits are available by request. Media and voice services use operation-specific request, concurrency, or session limits. Exceeding a limit returns HTTP 429.
Use exponential backoff for transient rate-limit responses, add request timeouts, and avoid concentrating a full minute of traffic into one second. Latency varies with model, context length, reasoning configuration, output size, and tool usage. Streaming is generally preferable for interactive or agentic operations.
xAI states that it does not train on customer API inputs or outputs without explicit permission. By default, API requests and responses may be retained for 30 days for abuse auditing and are encrypted at rest. Eligible teams can enable Zero Data Retention, which prevents prompt and output persistence but disables stateful Responses, Files, Collections, and Batch operations. Enterprise customers may obtain additional controls such as SSO, audit logging, invoicing, data residency, custom limits, and dedicated support.
When xAI API is a good or poor choice
The xAI API is a good fit when an application needs Grok models, current information through Web Search or X Search, tool-enabled workflows, multimodal services, file analysis, or an OpenAI-compatible integration path. Its Responses API is particularly relevant for applications that combine model output with functions, search, files, or other agentic steps.
It may be a poor fit when your workload requires fixed, easily predictable costs despite variable tool and media usage, strict availability of every feature under Zero Data Retention, or limits and model behavior that never change. Search results and generated content still require application-level verification. Before production deployment, confirm the current model catalog, prices, rate limits, regional availability, retention terms, and the capabilities of the specific model you plan to use.
