Developer platform

Baidu Qianfan API

Developer overview for Baidu, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Qianfan v2 OpenAI-compatible inference API
SDK support Official Qianfan Python SDK is available. Baidu Cloud SDK resources cover multiple languages. The current v2 documentation also supports the OpenAI Python client for compatible inference requests. Direct HTTPS integration is available for JavaScript and P
Rate limits Limits vary by model, account, service tier, and purchased quota. Qianfan documentation describes default and purchased TPM/RPM quota models; exact production limits should be checked in the account console and model documentation.
Platform

API overview

Endpoints

API access

Base URL https://qianfan.baidubce.com/v2
Primary API Qianfan v2 OpenAI-compatible inference API
Pricing

API pricing

Pricing model Usage-based pricing by model and service; token-based inference, token plans, quota plans, and separate charges for selected tools, media, batch, or training services.

Prices vary by model and may use input/output token rates, image or video units, web-search calls, token plans, or TPM/RPM quota plans. The current model catalog exposes model-specific pricing.

Developer experience

SDKs & usability

SDKs Official Qianfan Python SDK is available. Baidu Cloud SDK resources cover multiple languages. The current v2 documentation also supports the OpenAI Python client for compatible inference requests. Direct HTTPS integration is available for JavaScript and P
Ease of use Moderate. The v2 API is straightforward for OpenAI-compatible chat requests, but Baidu-specific authentication, model availability, regional/account requirements, Agents, and fine-tuning add platform complexity.
Documentation Good and extensive, with current Chinese-language API reference pages, model catalog documentation, SDK material, Agent documentation, and endpoint examples. English coverage is more limited.
Latency No single platform-wide latency guarantee was verified. Latency varies by model, prompt and output length, reasoning mode, concurrency, traffic period, service tier, and Agent or tool execution. Streaming improves time-to-first-visible-output but not nece
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Fine-tuning
✓ Web search
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

The current platform has two closely related developer surfaces: the v2 model inference API and the broader Qianfan/AppBuilder Agent platform. The v2 inference API supports chat completions, model discovery, streaming for compatible models, and function calling for compatible models. Agent and workflow APIs add persistent conversations, custom components, tools, file inputs, knowledge bases, and structured content types. File upload is available through /v2/agent/file/upload for supported agents and workflows. Fine-tuning and data-management APIs exist but may use Baidu Cloud AK/SK signing and additional IAM permissions. Web search is available as a platform or model-dependent capability rather than a universal property of every chat model. Image input is available through supported multimodal models, Agent file uploads, or component workflows. Structured JSON output is model- and endpoint-dependent; function calling provides a more consistently documented structured mechanism.

Examples

API examples

curl --fail-with-body --silent --show-error --request POST "https://qianfan.baidubce.com/v2/chat/completions" \
  --header "Authorization: Bearer ${QIANFAN_API_KEY}" \
  --header "Content-Type: application/json" \
  --data '{
    "model": "ernie-4.5-turbo-128k",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain what an API key is in one paragraph."}
    ],
    "stream": false
  }' | jq -r '.choices[0].message.content'

curl --fail-with-body --silent --show-error --request POST "https://qianfan.baidubce.com/v2/chat/completions" \
  --header "Authorization: Bearer ${QIANFAN_API_KEY}" \
  --header "Content-Type: application/json" \
  --data '{
    "model": "ernie-4.5-turbo-128k",
    "messages": [
      {"role": "system", "content": "Return a concise answer."},
      {"role": "user", "content": "List two benefits of API versioning."}
    ],
    "stream": true
  }'
# Install or upgrade the current OpenAI client:
# python -m pip install --upgrade openai

import json
import os
from openai import OpenAI

api_key = os.environ["QIANFAN_API_KEY"]
client = OpenAI(
    api_key=api_key,
    base_url="https://qianfan.baidubce.com/v2",
)

messages = [
    {"role": "system", "content": "You are a concise technical assistant."},
    {"role": "user", "content": "What is an API key?"},
]

try:
    response = client.chat.completions.create(
        model="ernie-4.5-turbo-128k",
        messages=messages,
        stream=False,
    )
    print(response.choices[0].message.content)

    messages.append({
        "role": "assistant",
        "content": response.choices[0].message.content,
    })
    messages.append({
        "role": "user",
        "content": "Give one security best practice.",
    })

    stream = client.chat.completions.create(
        model="ernie-4.5-turbo-128k",
        messages=messages,
        stream=True,
    )
    for chunk in stream:
        if not chunk.choices:
            continue
        delta = chunk.choices[0].delta
        if delta.content:
            print(delta.content, end="", flush=True)
    print()
except Exception as exc:
    print(f"Qianfan request failed: {exc}")
// Install the current OpenAI-compatible client:
// npm install openai

import OpenAI from "openai";

const apiKey = process.env.QIANFAN_API_KEY;
if (!apiKey) {
  throw new Error("QIANFAN_API_KEY is not set");
}

const client = new OpenAI({
  apiKey,
  baseURL: "https://qianfan.baidubce.com/v2",
});

try {
  const response = await client.chat.completions.create({
    model: "ernie-4.5-turbo-128k",
    messages: [
      { role: "system", content: "You are a concise technical assistant." },
      { role: "user", content: "Explain API versioning in two sentences." },
    ],
    stream: false,
  });

  console.log(response.choices[0]?.message?.content ?? "");

  const stream = await client.chat.completions.create({
    model: "ernie-4.5-turbo-128k",
    messages: [
      { role: "system", content: "Answer briefly." },
      { role: "user", content: "Name two HTTP methods." },
    ],
    stream: true,
  });

  for await (const chunk of stream) {
    process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
  }
  process.stdout.write("\n");
} catch (error) {
  console.error("Qianfan request failed:", error);
  process.exitCode = 1;
}
<?php

$apiKey = getenv('QIANFAN_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('QIANFAN_API_KEY is not set');
}

$url = 'https://qianfan.baidubce.com/v2/chat/completions';
$payload = [
    'model' => 'ernie-4.5-turbo-128k',
    'messages' => [
        [
            'role' => 'system',
            'content' => 'You are a concise technical assistant.'
        ],
        [
            'role' => 'user',
            'content' => 'Explain what an API key is in one paragraph.'
        ]
    ],
    'stream' => false
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
        'Accept: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_UNESCAPED_UNICODE | JSON_THROW_ON_ERROR),
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_CONNECTTIMEOUT => 10,
    CURLOPT_TIMEOUT => 120,
]);

$body = curl_exec($ch);
if ($body === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('Network error: ' . $error);
}

$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($body, true);
if ($status < 200 || $status >= 300) {
    $message = $data['message'] ?? $body;
    throw new RuntimeException('Qianfan API error (' . $status . '): ' . $message);
}

$content = $data['choices'][0]['message']['content'] ?? null;
if ($content === null) {
    throw new RuntimeException('The response did not contain assistant content');
}

echo $content . PHP_EOL;
Policies

Data & usage

Rate limits

Limits vary by model, account, service tier, and purchased quota. Qianfan documentation describes default and purchased TPM/RPM quota models; exact production limits should be checked in the account console and model documentation.

Developer guide

Baidu Qianfan API: Beginner's Guide to the Current Developer Platform

Baidu Qianfan is Baidu Intelligent Cloud's developer platform for hosted model inference and AI application development. Its current v2 API provides OpenAI-compatible chat completions, model discovery, streaming, function calling, multimodal input, file workflows, Agents, knowledge bases, and fine-tuning. This guide explains access, authentication, model selection, requests, pricing, SDK options, capabilities, limitations, and production considerations.
Baidu Qianfan is the current developer-facing API platform for Baidu's generative AI services. Developers can call Baidu ERNIE models and selected third-party models through the v2 inference API, or use the broader Qianfan and AppBuilder platform to create Agents, workflows, knowledge-based applications, and file-enabled experiences. The most accessible starting point is the OpenAI-compatible chat-completions interface at qianfan.baidubce.com/v2.
Baidu Qianfan is Baidu Intelligent Cloud's developer platform for model inference and AI application development. Its v2 API supports OpenAI-compatible chat completions, model discovery, streaming, function calling, multimodal files, Agents, workflows, knowledge bases, and fine-tuning. Pricing and capabilities vary by model, account, endpoint, and service tier.

What is the Baidu Qianfan API?

Baidu Qianfan is Baidu Intelligent Cloud's platform for using hosted AI models and building AI applications. It combines a model-inference API with higher-level services for Agents, workflows, file processing, knowledge bases, model management, and fine-tuning.

The current v2 inference API is designed around familiar chat-completion concepts. A request specifies a model and a messages array containing roles such as system, user, and assistant. Supported models can return ordinary or streamed responses, and compatible models can use function calling and multimodal inputs.

Qianfan is intended for developers building applications rather than for casual chatbot use. It is a suitable starting point for Chinese-language applications, Baidu ecosystem integrations, document and file workflows, Agent applications, and services that need access to several hosted models through one cloud platform.

Two related developer surfaces

Qianfan has two closely related parts:

  • v2 inference API: Direct model calls through endpoints such as POST /v2/chat/completions and GET /v2/models.
  • Qianfan and AppBuilder platform: Higher-level Agents, workflows, custom tools, knowledge bases, file inputs, persistent conversations, and application components.

These surfaces share the Qianfan developer ecosystem, but a feature available in an Agent workflow is not automatically available for every basic chat-completions request.

How to get access and obtain credentials

Normal inference and file APIs require a Qianfan API key. Keys are created and managed through Baidu Intelligent Cloud identity and access-management tools. Store the key on your server or in an environment variable; do not place it in browser JavaScript, mobile applications, source control, or client-side HTML.

Some platform-management and fine-tuning operations use Baidu Cloud access-key and secret-key request signing instead of the simpler Bearer API-key method. Check the documentation for the specific endpoint before implementing administrative or training workflows.

The main access sequence is:

  1. Create or use a Baidu Intelligent Cloud account and enable the relevant Qianfan services.
  2. Create an API key with permissions appropriate for the models or services your application will call.
  3. Check the current model catalog and pricing console.
  4. Save the key as a server-side environment variable such as QIANFAN_API_KEY.
  5. Send authenticated requests to the current v2 base URL: https://qianfan.baidubce.com/v2.

Account, region, service-tier, quota, and model availability requirements can affect access. Confirm the current requirements in the Baidu Intelligent Cloud console rather than assuming that every listed model is enabled for every account.

How to choose a model or API

Use the model-discovery endpoint instead of permanently hard-coding an old model name:

GET https://qianfan.baidubce.com/v2/models

The model catalog can expose identifiers, supported modalities, context limits, token limits, and model-specific pricing information. Use this information to select a model that matches the task and to detect changes before deployment.

For a conventional text assistant, begin with a compatible chat-completions model. For image input or other media, select a model whose catalog entry explicitly supports the required modality. For tools, Agents, file processing, or knowledge-base retrieval, evaluate the Qianfan or AppBuilder workflow surface rather than assuming that a basic completion request provides those features.

Model capabilities are not uniform. Streaming, function calling, JSON-oriented responses, multimodal input, web search, and other features can depend on the selected model, endpoint, account, or platform component.

Make a first request

The simplest request is an HTTPS POST to the v2 chat-completions endpoint. It uses a Bearer API key, a model identifier, and a messages array.

curl --fail-with-body --silent --show-error --request POST "https://qianfan.baidubce.com/v2/chat/completions" 
  --header "Authorization: Bearer ${QIANFAN_API_KEY}" 
  --header "Content-Type: application/json" 
  --data '{
    "model": "ernie-4.5-turbo-128k",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain what an API key is in one paragraph."}
    ],
    "stream": false
  }'

The model name in this example is an example of the current request pattern, not a guarantee that the model is enabled for every account. In production, verify the identifier and capabilities through the current model catalog.

The same request with the OpenAI Python client

Qianfan's v2 documentation demonstrates the current OpenAI-compatible client pattern. Install or upgrade the current openai package and point its base URL at Qianfan:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["QIANFAN_API_KEY"],
    base_url="https://qianfan.baidubce.com/v2",
)

response = client.chat.completions.create(
    model="ernie-4.5-turbo-128k",
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "What is an API key?"},
    ],
    stream=False,
)

print(response.choices[0].message.content)

This compatibility layer is convenient for developers who already know the OpenAI client interface. Direct HTTPS remains useful for languages without a suitable provider-specific SDK or when you need precise control over requests.

Understand the response

A successful chat-completions response normally contains a choices array. The generated assistant text is available at choices[0].message.content in a non-streaming response. Applications should still validate that the expected fields exist and handle provider errors, empty choices, interrupted connections, and model-specific response differences.

Conversation memory is normally implemented by your application. Retain the relevant previous messages and send them again with the next request. The basic chat endpoint should not be treated as a permanent conversation store unless you are using a separate Agent or application service that provides conversation management.

How Qianfan pricing works

Qianfan generally uses usage-based pricing. Text inference may be charged according to input and output tokens, while image, video, web-search, batch, fine-tuning, and other services can use different units or billing rules. The platform also offers token plans and quota-based service options.

There is no single price that applies to every Qianfan model. Prices can change by model and service, so production systems should read the current model catalog and pricing console rather than embedding assumptions in documentation or code. A model's price, context limit, and quota should be treated as configuration.

Budgeting should account for both sides of a conversation: retained history increases input tokens, while longer generated answers increase output usage. Tool calls, file processing, Agents, media generation, and fine-tuning may introduce separate charges or quota requirements.

Core capabilities available to developers

Chat and streaming

Chat requests use role-based messages and can support multi-turn conversations when your application resends the relevant history. Compatible endpoints can stream incremental output by setting stream to true. Streaming lets an interface display text as it arrives, improving perceived responsiveness, but it does not necessarily reduce total processing time.

const stream = await client.chat.completions.create({
  model: "ernie-4.5-turbo-128k",
  messages: [
    { role: "system", content: "Answer briefly." },
    { role: "user", content: "Name two HTTP methods." },
  ],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Streaming clients must handle partial output, network interruptions, provider errors, and final usage metadata where supplied. Do not assume that every chunk contains text.

Function calling and custom tools

Function calling lets a compatible model request that your application run a named function. You provide a function description and JSON Schema-like parameters. The model can return a tool call; your application validates the arguments, executes the local or remote operation, and sends the result back in a later conversation request.

The model does not automatically gain permission to perform arbitrary actions. Your application remains responsible for authentication, authorization, argument validation, rate limits, side effects, and user confirmation for sensitive operations.

Qianfan's Agent and workflow services add a broader tool layer, including custom tools, workflow components, file inputs, conversation identifiers, and structured content types. Tool compatibility should be checked for the selected model and application component.

Files and multimodal input

Supported Qianfan Agent workflows can accept uploaded text, spreadsheet, image, and audio files. Uploaded files may be associated with conversations or used by supported agents such as code interpreters, browser-use agents, deep-research agents, and AI assistants.

Image input is available through supported multimodal models, Agent file uploads, or component workflows. File support is therefore broader than simply attaching any file to any chat-completions request. Confirm accepted file types, size rules, processing behavior, and model compatibility in the endpoint documentation.

Agents, workflows, and knowledge bases

AppBuilder and related Qianfan components support visual and code-based development of RAG applications, Agents, workflows, user-interface-oriented applications, custom components, and knowledge bases. These higher-level tools are useful when an application needs retrieval, multi-step execution, persistent workflow state, or managed file processing.

They are separate from the basic v2 chat-completions call. Start with direct inference when your application only needs prompt-and-response generation; use an Agent or workflow when the platform-managed orchestration is part of the requirement.

Structured outputs

Function calling provides a relatively consistent structured mechanism by placing arguments in a tool-call schema. Some model and endpoint combinations also support JSON-oriented response formats, but strict structured-output behavior is not uniform across the model catalog.

If your application requires a specific schema, confirm support for the chosen model, request format, and endpoint. Validate the returned content locally and handle malformed or incomplete output instead of assuming that a text response is valid JSON.

Fine-tuning and data management

Qianfan includes fine-tuning and related data-management APIs. Fine-tuning tasks can use uploaded files or managed datasets, while task-management endpoints can expose status, model, training mode, progress, metrics, and checkpoint information.

Fine-tuning and administrative operations may use Baidu Cloud access-key and secret-key signing and require additional identity-management permissions. They should be planned separately from the simpler Bearer-key inference integration.

SDKs, documentation, and developer tools

The official Qianfan Python SDK is available, and Baidu Cloud provides SDK resources for multiple languages. The current v2 documentation also demonstrates the OpenAI Python client for compatible inference calls. JavaScript and PHP applications can use raw HTTPS or an appropriate HTTP client when a provider-specific integration is not required.

The Qianfan console provides a playground and platform-management tools. The official documentation includes the model list API, inference reference, Agent file-upload guidance, custom tool documentation, fine-tuning references, and pricing information. Use the model catalog and console together when testing because the available models, quotas, and prices depend on the account and service configuration.

Limits and production considerations

  • Model compatibility: Features such as streaming, function calling, multimodal input, JSON-oriented output, and web search are model- or endpoint-dependent.
  • Quotas: Request-per-minute and token-per-minute limits vary by model, account, service tier, region, and purchased quota. Check the account console for effective limits.
  • Latency: Response time varies with model size, prompt and output length, reasoning mode, concurrency, traffic, service tier, and Agent or tool execution.
  • Regional and account access: Availability, payment methods, model entitlements, and feature access can vary by account, device, region, and platform configuration.
  • API generations: Keep the v2 base URL, package versions, authentication method, and method syntax consistent. Do not combine examples from older Qianfan interfaces with the current OpenAI-compatible v2 pattern.
  • Security: Keep API keys server-side, use environment variables or a secret manager, limit permissions, and avoid logging sensitive prompts, file contents, or credentials.
  • Reliability: Use timeouts, bounded retries with exponential backoff, response validation, provider-specific error handling, and safe handling of partial streams.
  • Data handling: The supplied research does not verify a universal Qianfan policy stating whether all API inputs are retained or used for training. Review current Baidu Cloud terms and endpoint-specific data policies before sending confidential information.

When Qianfan is a good or poor choice

Qianfan is a good choice when you need Baidu-hosted ERNIE or selected third-party models, strong Chinese-language support, access to Baidu's cloud ecosystem, OpenAI-compatible chat requests, Agent workflows, knowledge bases, file processing, or model fine-tuning. It is also useful when a project benefits from a Chinese cloud provider's integrated services rather than assembling separate inference and orchestration systems.

It may be a poor choice when your application requires globally uniform availability, fully standardized English documentation, one consistent capability set across all models, or pricing that does not vary by model and service. The distinction between direct inference, Agents, workflows, and fine-tuning also creates more platform complexity than a minimal single-endpoint API.

For a first implementation, begin with the v2 model catalog, make one authenticated non-streaming chat request, add local response validation, and then evaluate streaming, tools, files, or Agents only when the application needs them.

Sources 10