What is the Z.ai API?
The Z.ai API is a developer platform for calling GLM models and related services from software. Its primary interface is an OpenAI-compatible chat-completions API, with support for ordinary responses, server-sent-event streaming, function calling, structured output, and multimodal message content.
The platform also documents specialized APIs for image generation, asynchronous video generation, audio transcription, OCR and layout parsing, web search, web reading, tokenization, translation, and agent workflows. These capabilities are not necessarily available through every model or endpoint, so an application should verify the documentation for the exact operation it plans to use.
Z.ai is a good fit when an application needs GLM models, long-context or reasoning-oriented workloads, coding assistance, multimodal input, external tools, or a provider with OpenAI-compatible request patterns. It may be a poor fit when an organization requires universally consistent regional availability, a single clearly documented data-retention policy for all accounts, or a mature consumer-focused personalization and mobile ecosystem.
Getting access and creating an API key
Start at the Z.ai developer or model API platform and create an account or sign in. API keys are created or claimed through the Z.ai API portal. Keep the key on the server side and load it from an environment variable rather than placing it in browser code, source control, logs, or user-visible configuration.
The general API uses the base URL https://api.z.ai/api/paas/v4. Coding-plan customers use the separate base URL https://api.z.ai/api/coding/paas/v4. These endpoints should not be mixed: the coding endpoint is intended for coding-plan access, while the general endpoint is used for normal API resource packages and prepaid balances.
- General API base URL:
https://api.z.ai/api/paas/v4 - Coding-plan base URL:
https://api.z.ai/api/coding/paas/v4 - Core operation:
POST /chat/completions - Authentication:
Authorization: Bearer YOUR_API_KEY - Request format: JSON over HTTPS
Choosing an API and model
For a first text-generation integration, use the general chat-completions endpoint and select a GLM model enabled for your account. Current documentation highlights GLM-5.3 and related GLM models, but the available catalog can vary by account, region, plan, and service changes. Confirm the exact model identifier in the current Z.ai documentation or account portal instead of assuming that every published model is available to every user.
Choose the operation based on the task. Chat completions are the general-purpose starting point. Multimodal chat is appropriate when the selected model accepts image or other supported media content. Web search, web reading, OCR, image, video, audio, translation, and agent operations have their own documentation and may use different request shapes or asynchronous workflows.
For coding-plan access, use the coding endpoint and the quota associated with that plan. General prepaid API balances and coding-plan quotas are separate. This distinction matters when a request succeeds against one endpoint but fails because the account has no corresponding resource or entitlement on the other.
Make your first request
The following request sends a short conversation to the general chat-completions endpoint. Set ZAI_API_KEY in the shell before running it. The model name is an example based on the current documentation; replace it with a model enabled for your account if necessary.
set -euo pipefail
: "${ZAI_API_KEY:?Set ZAI_API_KEY first}"
curl --fail-with-body --silent --show-error
--request POST
--url https://api.z.ai/api/paas/v4/chat/completions
--header "Authorization: Bearer ${ZAI_API_KEY}"
--header "Content-Type: application/json"
--data '{
"model": "glm-5.3",
"messages": [
{
"role": "system",
"content": "You are a concise technical assistant."
},
{
"role": "user",
"content": "Explain API rate limiting in three bullet points."
}
],
"temperature": 0.2
}'The request contains a model identifier and an ordered messages array. A system message can establish general behavior, while the user message contains the current task. For a multi-turn conversation, your application stores the relevant history and sends it again with the next request.
Understand the response
A successful non-streaming response follows the familiar chat-completions pattern. The generated text is normally read from the first choice's message content:
response.choices[0].message.contentApplications should still check for missing choices, empty content, tool calls, and error responses. Do not assume every successful response contains ordinary text: a model may return a request to call a tool, or a specialized operation may return a different documented structure.
For production code, parse the JSON response, inspect the HTTP status, record a request identifier when provided, and handle provider errors without exposing the API key or sensitive prompt content in logs.
How Z.ai API pricing works
General API usage is usage-based rather than a single unlimited subscription. Z.ai promotes token usage bundles and prepaid API resources, with pricing depending on the selected model and account arrangement. Enterprise, volume, or business pricing may be handled through sales. Public prices and model availability can change, so check the current Z.ai pricing pages before estimating a project budget.
Coding plans are separate from ordinary API billing. They use their own subscription quotas and coding endpoint. A coding-plan subscription should not be treated as a general-purpose balance for every API operation.
For budgeting, measure both input and output usage for representative prompts, include retries and tool calls, and account for multimodal or specialized operations separately where the documentation uses different pricing. Monitor account consumption and quota rather than relying only on a static estimate.
Streaming, conversations, and multimodal input
Set stream to true to receive incremental server-sent events. Streaming lets an interface display text as it arrives instead of waiting for the complete response. A client should process delta content, recognize the terminating event, and handle possible error payloads.
For multi-turn applications, keep conversation state in your own application. Preserve assistant tool-call messages, tool-call identifiers, and the matching tool results accurately. This is particularly important when a conversation alternates between model output and external actions.
The chat-completions API supports multimodal message content, including text and image inputs, with additional documentation for audio, video, and file content. Media support is model-dependent. Select a model that explicitly accepts the required media type and follow the documented content-part format rather than assuming that every GLM model accepts every kind of input.
Function calling, tools, and MCP
Function calling allows the model to request an operation that your application implements. You describe a function with a name, explanation, and JSON Schema-like parameter definition. The model does not safely execute arbitrary code by itself: your application must validate the arguments, enforce permissions, run the function, and send the result back in a follow-up request.
A typical flow is:
- Send the user message and available tool definitions.
- Inspect the response for a tool call.
- Validate the requested function and its arguments.
- Execute only an approved operation.
- Append the assistant tool-call message and the tool result.
- Send another completion request so the model can produce the final answer.
Z.ai also documents MCP server calling. MCP can connect a model to external servers using documented transports such as streamable HTTP or SSE. This can support web search, visual analysis, document processing, and other external operations, but the same security rules apply: restrict available tools, validate inputs, set timeouts, and treat external tool results as untrusted data.
Files and specialized developer APIs
File upload is documented in agent workflows, including translation workflows that use reference materials. General multimodal chat can accept file-like content where supported by the selected model and endpoint. File support is therefore operation-specific, not a universal promise that any file can be attached to any chat request.
Depending on the use case, developers can also access APIs for:
- Image generation
- Asynchronous video generation
- Audio transcription
- OCR and layout parsing
- Web search and web reading
- Tokenization and translation
- Specialized agent conversations, including slide and translation workflows
Before implementing a file or media workflow, check whether the operation is synchronous or asynchronous, which model or endpoint accepts the content, how results are retrieved, and what account permissions or quotas apply.
Request structured JSON output
Compatible models and request modes support structured output controls. A JSON-object response can be requested with response_format, but the returned text should still be parsed and validated against the application's own schema. Structured output does not remove the need for handling unsupported models, malformed responses, missing fields, or account-specific limitations.
curl --fail-with-body --silent --show-error
--request POST
--url https://api.z.ai/api/paas/v4/chat/completions
--header "Authorization: Bearer ${ZAI_API_KEY}"
--header "Content-Type: application/json"
--data '{
"model": "glm-5.3",
"messages": [
{
"role": "user",
"content": "Return a JSON object with keys name and priority for a timeout bug."
}
],
"response_format": {"type": "json_object"}
}'After receiving the response, parse the message content as JSON and validate types and required fields. If the application requires a strict schema, reject or safely repair invalid data instead of passing it directly to business logic.
Python SDK and integration options
Z.ai documents an official Python SDK distributed as zai-sdk, as well as an official Java SDK. The following Python example uses the documented ZaiClient pattern and includes ordinary, streaming, and structured-output requests.
import json
import os
import sys
from zai import ZaiClient
api_key = os.environ.get("ZAI_API_KEY")
if not api_key:
raise SystemExit("Set ZAI_API_KEY before running this example")
client = ZaiClient(api_key=api_key)
try:
response = client.chat.completions.create(
model="glm-5.3",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Give me two practical API retry recommendations."},
],
temperature=0.2,
)
print(response.choices[0].message.content)
except Exception as exc:
print(f"Z.ai request failed: {exc}", file=sys.stderr)
raise
stream = client.chat.completions.create(
model="glm-5.3",
messages=[
{"role": "user", "content": "Explain server-sent events in four short steps."}
],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()
structured = client.chat.completions.create(
model="glm-5.3",
messages=[
{"role": "user", "content": "Return JSON with keys title and severity for a
database timeout bug."}
],
response_format={"type": "json_object"},
)
print(json.loads(structured.choices[0].message.content))JavaScript and PHP applications can call the HTTP API directly. The current documentation reviewed for this guide verifies Python and Java SDKs; it does not verify a first-party JavaScript or PHP SDK in the documentation index.
Playground and developer tools
Z.ai provides a model API page and an API portal for experimenting with access, models, and usage. The documentation includes quick starts, HTTP examples, SDK guides, API references, streaming, function calling, MCP, structured output, web search, agents, and specialized media APIs.
Use the portal or playground for small manual tests, but reproduce important tests in code before production. A playground request may use a model, setting, account entitlement, or quota that differs from the environment used by your application.
Limits and production considerations
Z.ai publishes a dedicated rate-limit reference, but limits can vary by model, account, plan, and service access. There is no single universal limit that should be hard-coded into every client. Implement bounded retries with exponential backoff for transient failures, respect HTTP error responses, and monitor quota and usage.
There is also no universal public latency guarantee verified for every model or region. Latency depends on model choice, prompt and output size, reasoning settings, service load, network path, and whether streaming is enabled. Set application timeouts, avoid unbounded retries, and design user interfaces that can tolerate delayed or partial results.
Keep secrets outside prompts and source code. Minimize sensitive data, apply retention controls in your own systems, and review the current API terms and data-processing addendum before sending confidential or regulated information. The public materials reviewed do not establish one universal API retention period or a simple training-use rule that applies to every API customer and region.
Finally, model and feature availability can change. Verify supported models at deployment time, handle unsupported-feature errors, test long-context and tool workflows under realistic load, and keep provider-specific behavior behind an application interface if portability matters.
Advantages and limitations
Where the API is a strong choice
- OpenAI-compatible chat-completions patterns can shorten the path from an existing integration.
- The platform covers text, reasoning, coding, multimodal understanding, tools, search, agents, and specialized media operations.
- Official Python and Java SDK documentation is available alongside direct HTTP integration.
- MCP and function calling support can connect models to application services and external tools.
- General API access uses prepaid resources and usage bundles, which can suit teams that prefer usage-based purchasing.
What to evaluate carefully
- Model, feature, quota, and regional availability can vary.
- General API resources and coding-plan quotas use different endpoints and billing arrangements.
- File, audio, video, structured-output, and tool features depend on the specific model or operation.
- Rate limits and latency are account- and workload-dependent rather than covered by one universal public guarantee.
- API data retention, deletion, logging, and training-use terms should be confirmed for the specific account or contract.
When should you use the Z.ai API?
Choose Z.ai when you need GLM-based text or coding features, multimodal analysis, external tools, web search, agent workflows, or specialized image, video, audio, OCR, and translation operations. Its OpenAI-compatible core interface also makes it practical for developers who want a conventional chat-completions integration.
Evaluate alternatives carefully when your application depends on uniform availability across regions, a highly mature mobile or consumer-assistant ecosystem, extensive personalization tooling, or fully transparent universal data-retention and training controls. In all cases, run representative tests with the exact model, endpoint, account, region, prompt sizes, and tool workflows that production will use.
