ERNIE 4.5

ERNIE 4.5 Turbo

by Baidu · Current and accessible through Baidu Qianfan; canonical API endpoint is ernie-4.5-turbo-128k.

ERNIE 4.5 Turbo is Baidu’s text-only model for fast, inexpensive Chinese-language generation and long-context workloads. Available on Qianfan as ernie-4.5-turbo-128k, it supports a 138,240-token context, up to 12,288 output tokens, function calling, and optional web-search augmentation. It is well suited to enterprise agents, document processing, content creation, reasoning, and code assistance, but it does not support image, audio, or video input or output.

Text Reasoning Coding
ERNIE 4.5 Turbo is Baidu’s speed- and cost-oriented member of the ERNIE 4.5 model family. It is designed for text generation rather than multimodal interaction: applications send text and receive text, with optional tool and web-search support through Baidu Qianfan. The model is particularly relevant to developers building Chinese-language services, long-document workflows, enterprise agents, content-generation systems, and coding assistants that need predictable costs and quick responses. Its 128K-class context window is substantial, but users should not confuse that capacity with multimodal input or frontier-level reasoning. The supplied provider documentation identifies the current Qianfan endpoint as ernie-4.5-turbo-128k.
Outputs

What ERNIE 4.5 Turbo can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Web search
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family ERNIE 4.5
Model type General Purpose
Context window 138K tokens
Maximum output 12K tokens
Release date 2025-04-25
Status Current and accessible through Baidu Qianfan; canonical API endpoint is ernie-4.5-turbo-128k.
Knowledge cutoff notes

Baidu's current public model documentation reviewed for this record does not state a definitive knowledge-cutoff date for ERNIE 4.5 Turbo.

Model notes

ERNIE 4.5 Turbo is Baidu's text-generation model released at the Create 2025 developer conference. The current Qianfan API identifier is ernie-4.5-turbo-128k. Baidu documentation describes improvements in hallucination reduction, logical reasoning, coding, speed, and cost relative to ERNIE 4.5. The API model metadata specifies text-to-text operation, a 138,240-token context length, 125,952 maximum input tokens, and 12,288 maximum completion tokens. The model is listed as supporting function calling in Baidu's model-configuration documentation. Web-search pricing in the model-list API indicates first-party web-search support, but web search does not change the model's underlying knowledge cutoff. Editorial scores are comparative estimates, not provider benchmarks.

Cost

Model pricing

Input ¥0.0008 per 1,000 input tokens; web-search augmentation is priced at ¥0.004 per 1,000 tokens where applicable.
Output ¥0.0032 per 1,000 output tokens.
Model guide

ERNIE 4.5 Turbo: Baidu’s Fast, Low-Cost 128K Text Model

ERNIE 4.5 Turbo is Baidu’s text-only general-purpose model for fast, inexpensive Chinese-language generation, long-context applications, enterprise agents, content creation, reasoning, and code assistance. Available through Baidu Qianfan as ernie-4.5-turbo-128k, it offers a 138,240-token context window, up to 12,288 output tokens, function calling, and optional web-search augmentation. Its main trade-off is that it does not accept or produce images, audio, or video.

What is ERNIE 4.5 Turbo?

ERNIE 4.5 Turbo is a general-purpose text-generation model provided by Baidu. Baidu released it at the Create 2025 developer conference, positioning it as a faster and more economical option in the ERNIE 4.5 family. The model is currently accessible through Baidu Qianfan, where its canonical API identifier is ernie-4.5-turbo-128k.

In practical terms, it is intended for applications that need to understand prompts and generate written responses at relatively low cost. Typical workloads include Chinese-language writing and rewriting, summarization, question answering, document processing, retrieval-augmented generation, enterprise assistants, agent workflows, reasoning tasks, and code assistance.

The “Turbo” designation is important for positioning, but it should not be treated as a universal performance claim. The available research describes improvements in speed, cost, logical reasoning, coding, and hallucination reduction relative to ERNIE 4.5. Those are provider-positioning claims; the comparative scores below are editorial estimates rather than published benchmark results.

Verified specifications and limits

SpecificationDetails
ProviderBaidu
Model familyERNIE 4.5
Current Qianfan endpointernie-4.5-turbo-128k
Release dateApril 25, 2025
Input and outputText input to text output
Context length138,240 tokens
Maximum input allocation125,952 tokens, according to the API model metadata
Maximum completion12,288 tokens
Function callingSupported in Baidu’s model-configuration documentation
Knowledge cutoffNot publicly verified in the reviewed Baidu documentation

A token is a unit of text used for processing and billing; it may represent part of a word, a complete word, punctuation, or characters in languages such as Chinese. The 138,240-token context limit represents the model’s total working window, while the documented maximum input and completion figures describe how that window can be allocated for a request. In a long-document application, the prompt, retrieved material, instructions, and generated answer all need to fit within the applicable limits.

Text-only operation with tool support

ERNIE 4.5 Turbo is documented as a text-to-text model. It does not natively accept images, audio, or video, and it does not generate images, audio, video, or other non-text media. This distinction matters because Baidu’s broader consumer AI ecosystem advertises multimodal features, but those product-level capabilities should not be attributed automatically to this specific Qianfan model.

The model does support function calling. Function calling allows an application to describe external operations—such as querying a database, looking up an order, or invoking an internal business service—and ask the model to produce a structured tool request. The external application still performs the operation and returns the result; the model does not independently execute arbitrary code merely because function calling is available.

Baidu’s model-list information also identifies web-search augmentation as available at an additional price where applicable. Web search can provide access to retrieved web information during a request, but it does not establish a definitive knowledge-cutoff date for the underlying model. Developers should also distinguish provider-supported search from a guarantee that every answer is current or correct.

Pricing and cost profile

The supplied Qianfan pricing information lists ERNIE 4.5 Turbo at ¥0.0008 per 1,000 input tokens and ¥0.0032 per 1,000 output tokens. Web-search augmentation is listed separately at ¥0.004 per 1,000 tokens where applicable.

UsageListed price
Input tokens¥0.0008 per 1,000 tokens
Output tokens¥0.0032 per 1,000 tokens
Web-search augmentation¥0.004 per 1,000 tokens where applicable

Input and output are priced separately, and generated tokens cost more per thousand than input tokens under these listed rates. Actual project cost depends on prompt length, response length, request volume, and whether search augmentation is used. The prices should therefore be treated as the documented Qianfan rates supplied for this record, not as a promise that every Baidu channel or future account configuration will use identical billing.

Reasoning, coding, speed, and reliability

ERNIE 4.5 Turbo is aimed at routine and moderately demanding reasoning rather than applications that require independently established frontier reasoning performance. Baidu describes improvements in logical reasoning and hallucination reduction compared with ERNIE 4.5. The model also supports coding assistance, including tasks such as generating snippets, explaining code, transforming text between formats, and helping reason through implementation problems.

For editorial comparison, the supplied evaluation assigns reasoning 7/10, coding 7/10, speed 9/10, and cost 9/10. These are comparative editorial estimates, not Baidu-published benchmark scores. They express the model’s intended trade-off: fast and inexpensive text generation with useful reasoning and coding ability, rather than maximum depth on the hardest problems.

Reliability still depends on the prompt, source material, application design, and verification process. The model’s undocumented knowledge-cutoff date is an additional reason to use retrieval or web search for time-sensitive information. Function calling can connect the model to authoritative systems, but the application should validate tool arguments and outputs before taking consequential actions.

Where ERNIE 4.5 Turbo fits best

  • Chinese-language applications: writing, rewriting, summarization, customer support, classification, and question answering where low latency and low token cost matter.
  • Long-context document workflows: analyzing large text collections, preparing summaries, extracting information, and comparing sections of lengthy documents within the supported context window.
  • Enterprise agents: assistants that need to select or call business tools, retrieve information, and produce a textual response.
  • Content operations: drafting product copy, reports, outlines, internal communications, and other text where high request volume makes pricing important.
  • Code assistance: code explanation, generation, transformation, and troubleshooting when the application does not require a specialized programming model.
  • Search-enhanced answers: workflows that benefit from Baidu web-search augmentation and can account for its separate usage charge.

Its combination of a large context window and low listed token prices can be useful for applications that process substantial text repeatedly. However, a large context limit does not guarantee that every detail will receive equal attention, so document segmentation, retrieval, and output validation may still improve results.

Limitations and when another option may be better

The clearest limitation is modality. Choose another model or service if the application must interpret photographs, screenshots, audio, or video, or if it must generate media. ERNIE 4.5 Turbo itself is documented as text-only even though Baidu’s wider consumer offerings include multimodal features.

It may also be a poor fit when independently verified structured-output guarantees are essential. The supplied record does not verify JSON mode or JSON Schema output support. Developers who require strict machine-readable responses should confirm the current Qianfan documentation and add application-level validation rather than assuming that function calling is equivalent to a general JSON-mode guarantee.

For extremely difficult mathematical, scientific, or multi-step reasoning tasks, a more reasoning-focused model may be preferable if its additional latency and cost are justified. Conversely, for simple classification or short completions, an even smaller model could be more economical. A specialized coding model may also be a better choice for large repositories or programming tasks requiring deeper codebase analysis, but no specific alternative is identified in the supplied research.

Users should also account for deployment context. ERNIE 4.5 Turbo is available through Baidu Qianfan, and the reviewed sources do not establish a globally standardized price or availability policy across all regions and accounts. Teams operating outside Baidu’s primary ecosystem should verify account access, regional availability, billing, data handling, and service-level requirements before committing to production use.

When to choose ERNIE 4.5 Turbo

Choose ERNIE 4.5 Turbo when the workload is primarily text-based, benefits from Chinese-language capability and long context, and values fast responses and low listed token costs. It is especially sensible for high-volume generation, document workflows, enterprise tool use, and applications where a 128K-class context window is more important than multimodal input.

Choose a different option when native media understanding or generation is required, when strict structured-output behavior has been verified as a hard requirement, or when the task demands the strongest independently measured reasoning available regardless of cost and latency. In all cases, test representative prompts in the target language and workload: the practical balance between speed, quality, cost, and reliability depends on the application’s documents, instructions, tools, and validation rules.


Answers to Frequently Asked Questions

When should I choose ERNIE 4.5 Turbo?
Choose ERNIE 4.5 Turbo for primarily text-based applications that benefit from Chinese-language capability, long context, fast responses, and low listed token costs. It fits high-volume generation, long-document workflows, enterprise agents, search-enhanced answers, and code assistance. Another option may be better for multimodal tasks, strict structured-output requirements, or the most demanding mathematical, scientific, and multi-step reasoning workloads.
Does ERNIE 4.5 Turbo support images, audio, video, or function calling?
ERNIE 4.5 Turbo is documented as a text-to-text model and does not natively accept or generate images, audio, or video. It does support function calling, allowing applications to connect the model to databases, business services, and other external tools. The application remains responsible for executing tools and validating their inputs and outputs.
How much does ERNIE 4.5 Turbo cost?
The listed Baidu Qianfan prices are ¥0.0008 per 1,000 input tokens and ¥0.0032 per 1,000 output tokens. Web-search augmentation is listed separately at ¥0.004 per 1,000 tokens where applicable. Actual costs depend on prompt length, response length, request volume, and search usage.
What is ERNIE 4.5 Turbo?
ERNIE 4.5 Turbo is Baidu’s general-purpose text-generation model in the ERNIE 4.5 family. It is available through Baidu Qianfan under the API identifier `ernie-4.5-turbo-128k` and is designed for fast, low-cost workloads such as Chinese-language writing, summarization, document processing, enterprise assistants, retrieval-augmented generation, reasoning, and code assistance.
What is the context window and maximum output length of ERNIE 4.5 Turbo?
ERNIE 4.5 Turbo has a total context length of 138,240 tokens. Baidu’s API metadata lists a maximum input allocation of 125,952 tokens and a maximum completion length of 12,288 tokens. The prompt, retrieved content, instructions, and generated response must fit within the applicable limits.


Sources 4
Provider

About Baidu