ERNIE

ERNIE 5.0

by Baidu · Current and accessible through Baidu Qianfan as of September 25, 2026

ERNIE 5.0 is Baidu’s omni-modal foundation model available through Qianfan. It accepts text, images, audio, and video, supports reasoning, coding, streaming, and tool use, and offers a documented 248,832-token total context. The endpoint currently returns text only, with tiered pricing based on request length. It is best suited to multimodal analysis, long documents, enterprise workflows, coding, and text-based agents rather than native media generation.

Text Reasoning Coding
ERNIE 5.0 is Baidu’s native omni-modal foundation model, designed to combine text, image, audio, and video understanding in one model. It is available as the ernie-5.0 endpoint through Baidu Qianfan and is aimed at complex reasoning, coding, long documents, multimodal analysis, and tool-enabled applications. The documented endpoint currently produces text rather than images, audio, or video, so its practical role is multimodal understanding and text-based agent work rather than media generation.
Outputs

What ERNIE 5.0 can produce

Text
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
7/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family ERNIE
Model type Multimodal
Context window 249K tokens
Maximum output 66K tokens
Release date November 2025
Status Current and accessible through Baidu Qianfan as of September 25, 2026
Knowledge cutoff notes

Baidu has not published a direct, authoritative knowledge-cutoff date for the exact ERNIE 5.0 model in the reviewed first-party documentation.

Model notes

ERNIE 5.0 is Baidu's native omni-modal foundation model and is independently identifiable as the Qianfan endpoint ernie-5.0. Baidu describes the model as jointly modeling text, images, audio, and video and reports approximately 2.4 trillion total parameters in its technical report. The current Qianfan API metadata lists text, image, voice, and video inputs but text output only, so multimodal_output and image/audio/video output are set to 0 for the documented endpoint. Current API metadata reports a 248,832-token total context length, 121,856 maximum input tokens, and 65,536 maximum output tokens, while Baidu's international model table rounds the context window to 128K and lists a 119K maximum input length. Pricing is tiered by input length and is expressed in Chinese yuan per 1,000 tokens. Baidu also lists a web-search charge of ¥0.004 per search request. ERNIE 5.0 has a later successor, ERNIE 5.1, but remains listed as an accessible model.

Cost

Model pricing

Input ¥0.006 per 1K input tokens for requests up to 32K tokens; ¥0.010 per 1K input tokens above 32K
Output ¥0.024 per 1K output tokens for requests up to 32K tokens; ¥0.040 per 1K output tokens above 32K
Model guide

ERNIE 5.0: Baidu’s Long-Context Omni-Modal Model for Reasoning and Agents

ERNIE 5.0 is Baidu’s omni-modal foundation model for understanding text, images, audio, and video while supporting reasoning, coding, tool use, and long-context applications. Its Qianfan endpoint accepts multimodal input but currently returns text, with a documented 248,832-token context limit and tiered per-token pricing.

What is ERNIE 5.0?

ERNIE 5.0 is a foundation model from Baidu and part of the ERNIE model family. Baidu describes it as a unified omni-modal model that jointly models text, images, audio, and video. The company’s technical material reports approximately 2.4 trillion total parameters, although parameter count should be treated as a provider-published architectural claim rather than a direct measure of performance.

For practical use, the independently identifiable Qianfan endpoint is ernie-5.0. The current Qianfan model metadata lists text, image, voice, and video as supported input types, while listing text as the output type. This distinction matters: ERNIE 5.0 can analyze multimedia content through the documented endpoint, but that endpoint is not documented as a native image, audio, or video generator.

Where ERNIE 5.0 fits in Baidu’s lineup

ERNIE 5.0 sits in Baidu’s foundation-model catalog and is accessible through the Qianfan platform. It is intended for developers and organizations building applications rather than only for casual chatbot use. Typical applications include multimodal question answering, document processing, research assistants, coding tools, agent workflows, and enterprise systems that need to combine several kinds of input.

Baidu has since published ERNIE 5.1, which is a later successor, but ERNIE 5.0 remains listed as an accessible model in the supplied Qianfan catalog information. That makes it relevant when an application has already been built around its endpoint, pricing, or behavior. The existence of a successor is also a reason to verify current model availability and migration guidance before starting a new production integration.

Supported modalities and output behavior

ERNIE 5.0’s main technical distinction is the range of inputs it can interpret. The current endpoint supports:

  • Text input: ordinary prompts, instructions, documents, and code.
  • Image input: visual question answering and image-based analysis.
  • Audio or voice input: analysis of spoken or other supported audio content.
  • Video input: analysis of video content.
  • Text output: written answers, explanations, extracted information, code, and tool-oriented responses.

The Qianfan metadata sets image, audio, and video output to zero for this endpoint. In other words, ERNIE 5.0 may help interpret a video or image and describe what it finds, but the documented model endpoint should not be selected when the application requires it to return a generated image, soundtrack, or video. Baidu’s wider consumer ecosystem advertises media-creation features, but those should not be conflated with the output capabilities of the ernie-5.0 Qianfan endpoint.

Context window and output limits

The current Qianfan metadata reports a total context length of 248,832 tokens. It lists a maximum input length of 121,856 tokens and a maximum output length of 65,536 tokens. A token is a unit of text used by the model; it may represent a whole short word, part of a longer word, punctuation, or another fragment. These limits are therefore much larger than equivalent character counts.

The large context is useful for applications that need to work across long reports, collections of documents, extended codebases, or lengthy conversation history. It does not mean that every request can combine the maximum input and maximum output simultaneously without regard to platform rules. Developers should rely on the limits returned by the current Qianfan API documentation and leave room for system instructions, tool messages, and any additional request overhead.

Baidu’s international model information rounds the context window to 128K and lists a maximum input length of approximately 119K. The more detailed current Qianfan metadata gives the 248,832-token total context figure and 121,856-token input figure. The difference appears to reflect catalog presentation and counting conventions, so applications should use the endpoint-specific limits exposed in their target region.

Reasoning, coding, and tool use

ERNIE 5.0 is positioned for complex reasoning, coding, planning, and agent-style workflows. Baidu’s official descriptions highlight reasoning, coding, multimodal understanding, agentic planning, and tool use. The Qianfan metadata also marks tool use and streaming as supported.

In practice, this makes the model suitable for tasks such as examining an image together with a written specification, extracting findings from a video and turning them into a report, debugging code while considering a large project context, or deciding which external function to call during a workflow. Tool support means the model can participate in an application that exposes functions or services; it does not mean that ERNIE 5.0 independently has unrestricted access to every external system.

The supplied evaluation records assign ERNIE 5.0 a reasoning score of 8 and a coding score of 8. These are editorial or database evaluations, not scores published by Baidu, and they should be used as directional comparisons rather than formal benchmark results. Baidu’s own public claims establish the model’s intended reasoning, coding, and tool-use capabilities, but the supplied research does not provide a standardized benchmark table for verifying those scores.

ERNIE 5.0 pricing

The Qianfan pricing supplied for ERNIE 5.0 is tiered by input length and is quoted in Chinese yuan per 1,000 tokens:

Usage tierInput priceOutput price
Requests up to 32K tokens¥0.006 per 1K input tokens¥0.024 per 1K output tokens
Requests above 32K tokens¥0.010 per 1K input tokens¥0.040 per 1K output tokens

Baidu also lists a separate web-search charge of ¥0.004 per search request. These prices are API usage prices, not a consumer subscription price. Actual billing can depend on the Qianfan account, region, selected services, and the platform’s current commercial terms, so developers should confirm the live pricing page before estimating production costs.

The price structure creates a clear trade-off. Shorter requests are relatively inexpensive, while long-context requests cost more for both input and output. An application that does not need multimodal input, very large context, or complex reasoning may obtain better economics from a smaller or faster model. Conversely, ERNIE 5.0 can reduce the need to split multimodal or long-document tasks across several specialized models.

Main strengths and limitations

Strengths

  • Broad input coverage: one endpoint can accept text, images, audio, and video.
  • Long context: the documented 248,832-token total context supports unusually large prompts and document collections.
  • Complex task support: the model is designed for reasoning, coding, planning, and tool-enabled workflows.
  • Text-based integration: text output is useful for reports, extracted data, code, classifications, decisions, and conversational interfaces.
  • Streaming and web search: the supplied Qianfan metadata marks streaming and tool use as supported and identifies a web-search option.

Limitations

  • No documented media output on the endpoint: image, audio, and video generation should not be assumed from its omni-modal input capability.
  • Long-context pricing: requests above 32K tokens have higher listed input and output rates.
  • Regional and platform considerations: Qianfan availability, pricing, and documentation may differ by region, and Baidu’s materials are often oriented toward Chinese-language users.
  • Model-catalog change: ERNIE 5.1 is a later successor, so endpoint status and recommended model choices may change.
  • Unverified knowledge cutoff: Baidu has not published a direct authoritative knowledge-cutoff date for the exact ERNIE 5.0 model in the reviewed sources.

When to choose ERNIE 5.0

Choose ERNIE 5.0 when the application must interpret more than text and the same workflow may receive images, audio, or video. It is particularly suitable for Chinese and English enterprise applications involving long documents, multimodal research, coding assistance, visual or video analysis, and agents that need to call tools. Its large context can also be valuable when preserving extensive source material is more important than minimizing per-request cost.

It may be a good fit for a system that receives a meeting recording, related slides, and written instructions, then produces a structured textual brief. Other examples include analyzing an image alongside a technical manual, reviewing a long code repository with supporting documentation, or building a tool-enabled research assistant that returns text-based findings.

Another model may be more appropriate when the task is simple text generation, extremely latency-sensitive, or highly cost-sensitive. A specialized media-generation model is the better choice when the required result is a new image, audio track, or video. If a project is beginning from scratch, Baidu’s later ERNIE 5.1 should also be evaluated because it is the named successor, although the supplied research does not provide enough comparable specifications to declare it universally better.

Overall assessment

ERNIE 5.0 is best understood as a text-output model with unusually broad multimodal input and a very large context window. Its value comes from bringing text, image, audio, and video understanding together with reasoning, coding, streaming, and tool-oriented application design. The main practical compromise is that the documented endpoint does not generate non-text media, while its higher long-context tier can be more expensive than a smaller or narrower model.

For developers using Baidu Qianfan, ERNIE 5.0 is a strong candidate for multimodal and long-context workloads where a single model can simplify the application architecture. For straightforward prompts or media creation, a more specialized or less expensive option may offer a better balance.


Answers to Frequently Asked Questions

How much does ERNIE 5.0 cost on Qianfan?
For requests up to 32K tokens, the listed price is ¥0.006 per 1K input tokens and ¥0.024 per 1K output tokens. For requests above 32K tokens, the rates are ¥0.010 per 1K input tokens and ¥0.040 per 1K output tokens. Baidu also lists web search at ¥0.004 per search request, but developers should confirm current regional pricing before deployment.
How large is ERNIE 5.0's context window?
Current Qianfan metadata reports a total context length of 248,832 tokens, with a maximum input length of 121,856 tokens and a maximum output length of 65,536 tokens. Developers should verify the limits for their target region and leave room for system instructions, tools, and other request overhead.
What is ERNIE 5.0?
ERNIE 5.0 is a foundation model from Baidu that jointly processes text, images, audio, and video. The Qianfan endpoint is identified as ernie-5.0 and is designed for multimodal understanding, long-context tasks, reasoning, coding, and agent workflows.
What input and output modalities does ERNIE 5.0 support?
The Qianfan endpoint supports text, image, voice or audio, and video inputs. Its documented output type is text, so it can analyze multimedia content and produce written responses, but it should not be assumed to generate images, audio, or video.


Sources 6
Provider

About Baidu