Model catalog

Baidu Models

Browse the AI models associated with Baidu. Compare current and historical models by family, capabilities, context window, availability and intended use.

21 models tracked
21 Total models
17 Model families
10 Model types
19 Current / accessible
All models

Baidu model catalog

Baidu logo
Embedding-V1

Embedding-V1

Semantic search, vector retrieval, recommendation, semantic matching, knowledge bases, and retrieval-augmented generation

Type Embedding
Context 384
Speed 7/10
Status

Current and accessible through Baidu Qianfan and AI Studio embedding APIs

Input ¥0.0005 per 1,000 input tokens
Output ¥0.0005 per 1,000 completion tokens in current model-list metadata; the model practically returns vector embeddings rather than generated completion text
View model →

Multimodal understanding, long-context Chinese and English applications, complex reasoning, coding, tool-enabled agents, and enterprise workloads

Type Multimodal
Context 249K
Reasoning 8/10
Speed 7/10
Multimodal Image input Audio input
Status

Current and accessible through Baidu Qianfan as of September 25, 2026

Input ¥0.006 per 1K input tokens for requests up to 32K tokens; ¥0.010 per 1K input tokens above 32K
Output ¥0.024 per 1K output tokens for requests up to 32K tokens; ¥0.040 per 1K output tokens above 32K
View model →
Baidu logo
ERNIE-iRAG

ERNIE-iRAG-1.0

Realistic text-to-image generation, reference-grounded visual creation, commercial-style imagery, and applications needing reduced generative-artificiality.

Type Image Generation
Reasoning 2/10
Speed 8/10
Media output
Status

Available; listed in Baidu AI Cloud's inference-service model updates and supported for offline batch inference in Qianfan ModelBuilder documentation.

View model →

Object removal, masked image repainting, image variation, and batch image-editing workflows.

Type Image Editing
Reasoning 1/10
Speed 6/10
Multimodal Image input Media output
Status

Available with access dependent on Baidu Qianfan service configuration; listed for Qianfan batch inference. Related AI作画-iRAG版 API sales stopped on 2026-04-30, but that product notice does not explicitly state that the Qianfan ERNIE-iRAG-Edit endpoi

View model →

Lightweight local text generation, Chinese and English language experimentation, compact conversational systems, and domain adaptation

Type Lightweight
Context 131K
Reasoning 3/10
Speed 8/10
Fine-tuning Streaming
Status

Current open-weight model

View model →

Fast, low-cost Chinese text generation, long-context applications, enterprise agents, content creation, reasoning, and code assistance.

Type General Purpose
Context 138K
Reasoning 7/10
Speed 9/10
Tool use Web search
Status

Current and accessible through Baidu Qianfan; canonical API endpoint is ernie-4.5-turbo-128k.

Input ¥0.0008 per 1,000 input tokens; web-search augmentation is priced at ¥0.004 per 1,000 tokens where applicable.
Output ¥0.0032 per 1,000 output tokens.
View model →
Baidu logo
ERNIE 4.5 Turbo

ERNIE 4.5 Turbo VL

Cost-efficient multimodal understanding, image and video analysis, OCR, document comprehension, translation, visual question answering, and code-related tasks

Type Multimodal
Context 128K
Reasoning 7/10
Speed 8/10
Multimodal Image input Video input
Status

Current and available

Input ¥0.003 per 1,000 tokens
Output ¥0.009 per 1,000 tokens
View model →

Local or self-hosted multimodal applications, visual question answering, document and chart understanding, image analysis, video understanding, and efficient vision-language inference

Type Multimodal
Context 131K
Reasoning 7/10
Speed 8/10
Multimodal Image input Video input
Status

Current open-weight model; publicly released under the Apache 2.0 license

View model →
Baidu logo
ERNIE 5

ERNIE 5.1

Agentic workflows, web-search-assisted tasks, reasoning, Chinese-language knowledge work, creative writing, and general-purpose assistant applications.

Type General Purpose
Reasoning 8/10
Speed 8/10
Tool use Web search
Status

Current; officially released and accessible through Baidu's ERNIE website and AI Studio playground. Public Qianfan API availability for an exact ERNIE 5.1 endpoint was not verified.

View model →

Code completion, code generation, unit-test generation, code optimization, code explanation, and fine-tuning for software-development workflows.

Type Coding
Context 128K
Reasoning 4/10
Speed 6/10
Fine-tuning
Status

Retired; Qianfan ModelBuilder listed August 14, 2025 as the retirement date.

View model →

Chinese-language reasoning, long-form analysis, complex calculations, literary and document writing, agent workflows, and function calling

Type Reasoning
Context 33K
Reasoning 8/10
Speed 8/10
Tool use Web search Streaming
Status

Ready; currently accessible through Baidu Qianfan as ernie-x1-turbo-32k

View model →
Baidu logo
ERNIE X1

ERNIE X1.1

Deep reasoning, Chinese and English question answering, mathematics, coding, factual responses, tool calling, web-grounded applications, and agent workflows

Type Reasoning
Context 66K
Reasoning 8/10
Speed 7/10
Tool use Web search Streaming
Status

Current preview model

Input ¥0.001 per 1K input tokens
Output ¥0.004 per 1K output tokens
View model →

Low-cost image-to-video generation, short social clips, product animation, marketing assets, and turning still images into dynamic scenes

Type Video Generation
Reasoning 1/10
Speed 8/10
Multimodal Image input Media output
Status

Current and available through Baidu Qianfan API

Input CNY 1.00 per 5-second video
View model →
Baidu logo
MuseSteamer 2.0

Baidu MuseSteamer 2.0

Chinese image-to-video generation, audiovisual storytelling, marketing videos, multi-person dialogue, synchronized speech, sound effects, and cinematic short-form content

Type Multimodal
Speed 7/10
Multimodal Image input Media output
Status

Current model family; concrete Qianfan variants include Turbo, Lite, Pro, Turbo-I2V-Audio, and Turbo-I2V-Effect

View model →
Baidu logo
MuseSteamer Air

MuseSteamer-Air-Image

Low-cost text-to-image generation, marketing visuals, creative assets, illustrations, and rapid image prototyping

Type Image Generation
Reasoning 1/10
Speed 8/10
Media output
Status

Current and available

Input ¥0.05 per image at 1024x1024
Output ¥0.05 per generated image at 1024x1024
View model →
Baidu logo
PaddleOCR-VL

PaddleOCR-VL-1.5

Multilingual OCR and structured parsing of complex documents, including tables, formulas, charts, seals, scanned pages, warped documents, and screen photographs

Type Multimodal
Context 131K
Reasoning 3/10
Speed 8/10
Multimodal Image input
Status

Available; superseded by PaddleOCR-VL-1.6

View model →
Baidu logo
PP-Structure

PP-StructureV3

Document parsing, OCR, layout analysis, table extraction, and structured document understanding

Type Multimodal
Multimodal Image input
View model →

Fast enterprise question answering, summarization, workflow nodes, agent response generation, and text processing with a 32K context window

Type Lightweight
Context 33K
Reasoning 5/10
Speed 8/10
Tool use Streaming
Status

Available through Baidu’s documented Wenxin Workshop API; not found in the current Qianfan V2 model-list documentation, so new integrations should verify legacy endpoint compatibility.

View model →
Baidu logo
Qianfan-Agent-Lite

Qianfan-Agent-Lite-128K

Long-context Agent planning, task decomposition, component selection, and function-calling workflows.

Type Lightweight
Context 128K
Reasoning 6/10
Speed 8/10
Tool use
Status

Current platform-listed planning model; standalone lifecycle status is not separately published.

View model →

Enterprise agent workflows, intent recognition, instruction routing, and tool-calling tasks

Type Other
Context 33K
Reasoning 4/10
Speed 8/10
Tool use
Status

Legacy; current public availability not verified

View model →

Fast enterprise question answering, lightweight agent planning, component selection, and streaming text applications

Type Lightweight
Context 8K
Reasoning 4/10
Speed 8/10
Tool use Streaming
Status

Current and listed as available in Baidu Qianfan documentation; fast agent-oriented model

View model →