What is the Baidu Qianfan API?
Baidu Qianfan is Baidu Intelligent Cloud's platform for using hosted AI models and building AI applications. It combines a model-inference API with higher-level services for Agents, workflows, file processing, knowledge bases, model management, and fine-tuning.
The current v2 inference API is designed around familiar chat-completion concepts. A request specifies a model and a messages array containing roles such as system, user, and assistant. Supported models can return ordinary or streamed responses, and compatible models can use function calling and multimodal inputs.
Qianfan is intended for developers building applications rather than for casual chatbot use. It is a suitable starting point for Chinese-language applications, Baidu ecosystem integrations, document and file workflows, Agent applications, and services that need access to several hosted models through one cloud platform.
Two related developer surfaces
Qianfan has two closely related parts:
- v2 inference API: Direct model calls through endpoints such as
POST /v2/chat/completionsandGET /v2/models. - Qianfan and AppBuilder platform: Higher-level Agents, workflows, custom tools, knowledge bases, file inputs, persistent conversations, and application components.
These surfaces share the Qianfan developer ecosystem, but a feature available in an Agent workflow is not automatically available for every basic chat-completions request.
How to get access and obtain credentials
Normal inference and file APIs require a Qianfan API key. Keys are created and managed through Baidu Intelligent Cloud identity and access-management tools. Store the key on your server or in an environment variable; do not place it in browser JavaScript, mobile applications, source control, or client-side HTML.
Some platform-management and fine-tuning operations use Baidu Cloud access-key and secret-key request signing instead of the simpler Bearer API-key method. Check the documentation for the specific endpoint before implementing administrative or training workflows.
The main access sequence is:
- Create or use a Baidu Intelligent Cloud account and enable the relevant Qianfan services.
- Create an API key with permissions appropriate for the models or services your application will call.
- Check the current model catalog and pricing console.
- Save the key as a server-side environment variable such as
QIANFAN_API_KEY. - Send authenticated requests to the current v2 base URL:
https://qianfan.baidubce.com/v2.
Account, region, service-tier, quota, and model availability requirements can affect access. Confirm the current requirements in the Baidu Intelligent Cloud console rather than assuming that every listed model is enabled for every account.
How to choose a model or API
Use the model-discovery endpoint instead of permanently hard-coding an old model name:
GET https://qianfan.baidubce.com/v2/modelsThe model catalog can expose identifiers, supported modalities, context limits, token limits, and model-specific pricing information. Use this information to select a model that matches the task and to detect changes before deployment.
For a conventional text assistant, begin with a compatible chat-completions model. For image input or other media, select a model whose catalog entry explicitly supports the required modality. For tools, Agents, file processing, or knowledge-base retrieval, evaluate the Qianfan or AppBuilder workflow surface rather than assuming that a basic completion request provides those features.
Model capabilities are not uniform. Streaming, function calling, JSON-oriented responses, multimodal input, web search, and other features can depend on the selected model, endpoint, account, or platform component.
Make a first request
The simplest request is an HTTPS POST to the v2 chat-completions endpoint. It uses a Bearer API key, a model identifier, and a messages array.
curl --fail-with-body --silent --show-error --request POST "https://qianfan.baidubce.com/v2/chat/completions"
--header "Authorization: Bearer ${QIANFAN_API_KEY}"
--header "Content-Type: application/json"
--data '{
"model": "ernie-4.5-turbo-128k",
"messages": [
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain what an API key is in one paragraph."}
],
"stream": false
}'The model name in this example is an example of the current request pattern, not a guarantee that the model is enabled for every account. In production, verify the identifier and capabilities through the current model catalog.
The same request with the OpenAI Python client
Qianfan's v2 documentation demonstrates the current OpenAI-compatible client pattern. Install or upgrade the current openai package and point its base URL at Qianfan:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["QIANFAN_API_KEY"],
base_url="https://qianfan.baidubce.com/v2",
)
response = client.chat.completions.create(
model="ernie-4.5-turbo-128k",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "What is an API key?"},
],
stream=False,
)
print(response.choices[0].message.content)This compatibility layer is convenient for developers who already know the OpenAI client interface. Direct HTTPS remains useful for languages without a suitable provider-specific SDK or when you need precise control over requests.
Understand the response
A successful chat-completions response normally contains a choices array. The generated assistant text is available at choices[0].message.content in a non-streaming response. Applications should still validate that the expected fields exist and handle provider errors, empty choices, interrupted connections, and model-specific response differences.
Conversation memory is normally implemented by your application. Retain the relevant previous messages and send them again with the next request. The basic chat endpoint should not be treated as a permanent conversation store unless you are using a separate Agent or application service that provides conversation management.
How Qianfan pricing works
Qianfan generally uses usage-based pricing. Text inference may be charged according to input and output tokens, while image, video, web-search, batch, fine-tuning, and other services can use different units or billing rules. The platform also offers token plans and quota-based service options.
There is no single price that applies to every Qianfan model. Prices can change by model and service, so production systems should read the current model catalog and pricing console rather than embedding assumptions in documentation or code. A model's price, context limit, and quota should be treated as configuration.
Budgeting should account for both sides of a conversation: retained history increases input tokens, while longer generated answers increase output usage. Tool calls, file processing, Agents, media generation, and fine-tuning may introduce separate charges or quota requirements.
Core capabilities available to developers
Chat and streaming
Chat requests use role-based messages and can support multi-turn conversations when your application resends the relevant history. Compatible endpoints can stream incremental output by setting stream to true. Streaming lets an interface display text as it arrives, improving perceived responsiveness, but it does not necessarily reduce total processing time.
const stream = await client.chat.completions.create({
model: "ernie-4.5-turbo-128k",
messages: [
{ role: "system", content: "Answer briefly." },
{ role: "user", content: "Name two HTTP methods." },
],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Streaming clients must handle partial output, network interruptions, provider errors, and final usage metadata where supplied. Do not assume that every chunk contains text.
Function calling and custom tools
Function calling lets a compatible model request that your application run a named function. You provide a function description and JSON Schema-like parameters. The model can return a tool call; your application validates the arguments, executes the local or remote operation, and sends the result back in a later conversation request.
The model does not automatically gain permission to perform arbitrary actions. Your application remains responsible for authentication, authorization, argument validation, rate limits, side effects, and user confirmation for sensitive operations.
Qianfan's Agent and workflow services add a broader tool layer, including custom tools, workflow components, file inputs, conversation identifiers, and structured content types. Tool compatibility should be checked for the selected model and application component.
Files and multimodal input
Supported Qianfan Agent workflows can accept uploaded text, spreadsheet, image, and audio files. Uploaded files may be associated with conversations or used by supported agents such as code interpreters, browser-use agents, deep-research agents, and AI assistants.
Image input is available through supported multimodal models, Agent file uploads, or component workflows. File support is therefore broader than simply attaching any file to any chat-completions request. Confirm accepted file types, size rules, processing behavior, and model compatibility in the endpoint documentation.
Agents, workflows, and knowledge bases
AppBuilder and related Qianfan components support visual and code-based development of RAG applications, Agents, workflows, user-interface-oriented applications, custom components, and knowledge bases. These higher-level tools are useful when an application needs retrieval, multi-step execution, persistent workflow state, or managed file processing.
They are separate from the basic v2 chat-completions call. Start with direct inference when your application only needs prompt-and-response generation; use an Agent or workflow when the platform-managed orchestration is part of the requirement.
Structured outputs
Function calling provides a relatively consistent structured mechanism by placing arguments in a tool-call schema. Some model and endpoint combinations also support JSON-oriented response formats, but strict structured-output behavior is not uniform across the model catalog.
If your application requires a specific schema, confirm support for the chosen model, request format, and endpoint. Validate the returned content locally and handle malformed or incomplete output instead of assuming that a text response is valid JSON.
Fine-tuning and data management
Qianfan includes fine-tuning and related data-management APIs. Fine-tuning tasks can use uploaded files or managed datasets, while task-management endpoints can expose status, model, training mode, progress, metrics, and checkpoint information.
Fine-tuning and administrative operations may use Baidu Cloud access-key and secret-key signing and require additional identity-management permissions. They should be planned separately from the simpler Bearer-key inference integration.
SDKs, documentation, and developer tools
The official Qianfan Python SDK is available, and Baidu Cloud provides SDK resources for multiple languages. The current v2 documentation also demonstrates the OpenAI Python client for compatible inference calls. JavaScript and PHP applications can use raw HTTPS or an appropriate HTTP client when a provider-specific integration is not required.
The Qianfan console provides a playground and platform-management tools. The official documentation includes the model list API, inference reference, Agent file-upload guidance, custom tool documentation, fine-tuning references, and pricing information. Use the model catalog and console together when testing because the available models, quotas, and prices depend on the account and service configuration.
Limits and production considerations
- Model compatibility: Features such as streaming, function calling, multimodal input, JSON-oriented output, and web search are model- or endpoint-dependent.
- Quotas: Request-per-minute and token-per-minute limits vary by model, account, service tier, region, and purchased quota. Check the account console for effective limits.
- Latency: Response time varies with model size, prompt and output length, reasoning mode, concurrency, traffic, service tier, and Agent or tool execution.
- Regional and account access: Availability, payment methods, model entitlements, and feature access can vary by account, device, region, and platform configuration.
- API generations: Keep the v2 base URL, package versions, authentication method, and method syntax consistent. Do not combine examples from older Qianfan interfaces with the current OpenAI-compatible v2 pattern.
- Security: Keep API keys server-side, use environment variables or a secret manager, limit permissions, and avoid logging sensitive prompts, file contents, or credentials.
- Reliability: Use timeouts, bounded retries with exponential backoff, response validation, provider-specific error handling, and safe handling of partial streams.
- Data handling: The supplied research does not verify a universal Qianfan policy stating whether all API inputs are retained or used for training. Review current Baidu Cloud terms and endpoint-specific data policies before sending confidential information.
When Qianfan is a good or poor choice
Qianfan is a good choice when you need Baidu-hosted ERNIE or selected third-party models, strong Chinese-language support, access to Baidu's cloud ecosystem, OpenAI-compatible chat requests, Agent workflows, knowledge bases, file processing, or model fine-tuning. It is also useful when a project benefits from a Chinese cloud provider's integrated services rather than assembling separate inference and orchestration systems.
It may be a poor choice when your application requires globally uniform availability, fully standardized English documentation, one consistent capability set across all models, or pricing that does not vary by model and service. The distinction between direct inference, Agents, workflows, and fine-tuning also creates more platform complexity than a minimal single-endpoint API.
For a first implementation, begin with the v2 model catalog, make one authenticated non-streaming chat request, add local response validation, and then evaluate streaming, tools, files, or Agents only when the application needs them.
