ERNIE 4.5

ERNIE-4.5-0.3B

by Baidu · Current open-weight model

Baidu’s ERNIE-4.5-0.3B is a compact dense text-generation model with approximately 0.36 billion parameters and a 131,072-token context window. It is available as an Apache 2.0 open-weight model with PaddlePaddle, FastDeploy, and Transformers deployment options, plus SFT, LoRA, and DPO fine-tuning workflows.

Text Reasoning Coding
ERNIE-4.5-0.3B is the smallest dense model in Baidu’s ERNIE 4.5 open-weight model family. With approximately 0.36 billion parameters and a documented 131,072-token context length, it targets developers who need a compact Chinese- and English-capable text-generation model that can be deployed locally and adapted with ERNIEKit.
Outputs

What ERNIE-4.5-0.3B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

3/10 Reasoning
3/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family ERNIE 4.5
Model type Lightweight
Context window 131K tokens
Release date 2025-06-30
Status Current open-weight model
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was identified in the reviewed first-party documentation.

Model notes

ERNIE-4.5-0.3B is the post-trained dense language model in the ERNIE 4.5 family. Baidu also publishes ERNIE-4.5-0.3B-Base as a separate pretrained checkpoint. Framework-specific identifiers include ERNIE-4.5-0.3B-Paddle for PaddlePaddle weights and ERNIE-4.5-0.3B-PT for Transformers-compatible weights; these are deployment formats rather than separate model families. The model is open-weight and has no official per-token hosted API price documented in the reviewed sources. Baidu documents SFT, SFT-LoRA, DPO, and DPO-LoRA workflows through ERNIEKit. Editorial scores are comparative estimates, not vendor benchmarks.

Model guide

ERNIE-4.5-0.3B: Baidu’s Compact Open-Weight Model for Local Text Generation

ERNIE-4.5-0.3B is Baidu’s compact dense language model in the ERNIE 4.5 family. It accepts text and produces text, supports local deployment under the Apache License 2.0, and is designed for resource-efficient inference, fine-tuning, experimentation, and lightweight language applications.

What is ERNIE-4.5-0.3B?

ERNIE-4.5-0.3B is a compact dense causal language model from Baidu’s ERNIE 4.5 family. It is designed to generate and complete text rather than process images, audio, or video. The model was announced as part of Baidu’s open-source ERNIE 4.5 release on June 30, 2025.

The “0.3B” label refers to the model’s approximate size. Baidu’s published configuration contains about 0.36 billion parameters, making this model substantially smaller than the medium and large language models commonly used for demanding reasoning, coding, and broad knowledge tasks. Its smaller size is useful when local deployment speed, lower hardware requirements, or inexpensive experimentation matters more than maximum capability.

ERNIE-4.5-0.3B is the post-trained version intended for conversational and instruction-oriented use. Baidu also publishes an ERNIE-4.5-0.3B-Base checkpoint for text completion and further adaptation. The Base checkpoint is a related model, not a different name for the post-trained model reviewed here.

Technical specifications and context length

The verified specifications supplied for ERNIE-4.5-0.3B describe a dense transformer language model with 18 layers, 16 query heads, and 2 key-value heads. Its documented maximum context length is 131,072 tokens. A context window is the amount of text the model can consider in one request, including the input and any generated continuation, although the practical usable amount can depend on the deployment configuration.

SpecificationERNIE-4.5-0.3B
ProviderBaidu
Model familyERNIE 4.5
ArchitectureDense causal language model
Approximate parameters0.36 billion
Transformer layers18
Query heads / key-value heads16 / 2
Maximum documented context131,072 tokens
Input and outputText input and text output
LicenseApache License 2.0
Maximum output tokensNot specified in the supplied documentation

The 131,072-token context claim comes from the published model configuration and should not be interpreted as a guarantee that every local deployment will handle that length efficiently. Memory use, available hardware, inference settings, and serving software can affect the practical limit.

Modalities and core capabilities

ERNIE-4.5-0.3B is text-only. It accepts text and returns text; it does not natively provide image, audio, video, speech, music, or embedding output according to the supplied model data. It also is not documented as having built-in web search, tool calling, function calling, or action execution.

This makes the model suitable for ordinary language-generation tasks such as drafting, rewriting, classification-style prompting, extraction, completion, and lightweight conversation. Developers can build surrounding application logic that calls external tools, but that capability would belong to the application rather than being a verified native feature of ERNIE-4.5-0.3B.

The research does not provide vendor benchmark results or a model-specific knowledge cutoff. It is therefore safer to treat the model as a general-purpose compact generator rather than assume a particular level of factual freshness or guaranteed performance on specialized tasks.

Reasoning, coding, and reliability trade-offs

ERNIE-4.5-0.3B can follow prompts and produce structured or explanatory text, but its small parameter count limits how much complex reasoning can be expected from it. The supplied editorial assessment rates its reasoning capability at 3 out of 10 and its coding capability at 3 out of 10. These are comparative editorial scores, not Baidu-published benchmarks.

In practical terms, the model may be appropriate for straightforward transformations, short explanations, basic code completion, and constrained domain workflows after fine-tuning. It is less appropriate for multi-step mathematical reasoning, difficult software engineering, long chains of dependent instructions, or applications where factual reliability must be consistently high. Larger models are generally a better choice when the task depends on extensive world knowledge, robust planning, or sophisticated code generation.

Its small size offers a different advantage: lower resource demands than larger models. The supplied editorial scores rate speed at 8 out of 10 and cost at 9 out of 10. These scores are also subjective comparisons rather than guaranteed measurements. Actual speed and cost depend on hardware, quantization, batch size, sequence length, and serving framework.

Local deployment and fine-tuning

Baidu makes ERNIE-4.5-0.3B available through official Hugging Face repositories and related deployment documentation. The published framework-specific identifiers include ERNIE-4.5-0.3B-Paddle for PaddlePaddle weights and ERNIE-4.5-0.3B-PT for a Transformers-compatible checkpoint. These names identify weight formats and deployment paths rather than separate model families.

PaddlePaddle-based FastDeploy documentation supports local inference and demonstrates OpenAI-compatible serving for the Paddle-format model. The Transformers-compatible checkpoint provides another route for developers already working in the broader Transformers ecosystem. The appropriate choice depends on the surrounding software stack, hardware, and operational requirements.

ERNIEKit supports supervised fine-tuning, LoRA-based parameter-efficient fine-tuning, and direct preference optimization workflows for the model. Supervised fine-tuning adapts the model using examples of desired inputs and outputs. LoRA changes a relatively small set of added parameters instead of retraining the entire model, which can reduce training requirements. Direct preference optimization is intended to align responses with preference data. These options make the model practical for domain-specific experimentation when a small checkpoint is preferable to a much larger one.

Pricing and access

ERNIE-4.5-0.3B is an open-weight model released under the Apache License 2.0. The supplied sources do not document an official per-token hosted API price for this model. Consequently, there is no verified input or output price to report. Users who run it locally will instead incur infrastructure, hardware, hosting, and operational costs based on their own deployment.

Open-weight availability does not mean that every use is cost-free. Running the model may require a suitable computer or rented compute environment, and commercial users remain responsible for complying with the Apache 2.0 license and any applicable third-party deployment terms. The model’s smaller footprint can nevertheless make local testing and high-volume lightweight inference more economical than using a hosted large model.

Main strengths and limitations

Strengths

  • Small footprint: Approximately 0.36 billion parameters makes it easier to experiment with than larger ERNIE 4.5 variants.
  • Long documented context: The published 131,072-token context length is substantial for a model of this size, subject to deployment constraints.
  • Open-weight deployment: The Apache 2.0 license and downloadable checkpoints support local experimentation and application development.
  • Multiple tooling paths: Developers can use PaddlePaddle, FastDeploy, or a Transformers-compatible checkpoint.
  • Adaptation support: ERNIEKit documents SFT, LoRA, and DPO workflows.
  • Text generation focus: Its narrow modality profile keeps the model suited to applications that do not need multimodal processing.

Limitations

  • Limited model capacity: It is not positioned for frontier-level reasoning, difficult coding, or consistently high-reliability knowledge work.
  • Text-only operation: It does not natively understand or generate images, audio, video, or speech.
  • No verified native tools: Web search, function calling, and action output are not documented in the supplied research.
  • Unknown output ceiling: A model-specific maximum output-token limit was not supplied.
  • No documented hosted price: Users need to arrange local or third-party infrastructure rather than rely on a verified official per-token rate.
  • No verified knowledge cutoff: The reviewed first-party documentation does not identify one.

Best use cases

ERNIE-4.5-0.3B is a strong fit when the application needs a compact, locally deployable text model rather than the highest available intelligence. Suitable examples include:

  • Local text generation, completion, rewriting, and summarization experiments
  • Chinese- and English-language prototyping
  • Lightweight conversational interfaces
  • Domain adaptation with supervised fine-tuning or LoRA
  • Educational projects involving model serving or fine-tuning
  • High-volume, lower-complexity text processing where a small model can reduce infrastructure costs
  • Applications that benefit from a long configured context but do not require multimodal input

For production use, developers should test the exact prompts, languages, context lengths, and fine-tuning data relevant to their application. The model’s compact size can make iteration convenient, but it does not remove the need for evaluation and safeguards.

When to choose ERNIE-4.5-0.3B

Choose ERNIE-4.5-0.3B when local control, low resource requirements, open-weight access, and fine-tuning flexibility matter more than maximum reasoning or coding performance. It is particularly attractive for developers who already use Baidu’s PaddlePaddle, FastDeploy, or ERNIEKit tooling and want a small checkpoint for experimentation or specialized text applications.

A larger language model is more appropriate when the task requires difficult reasoning, advanced programming assistance, broad factual coverage, or consistently reliable instruction following. A multimodal model is necessary for image, audio, or video understanding. A hosted API model may also be preferable when the team does not want to manage hardware, serving, updates, or inference capacity.

Within the ERNIE 4.5 family, the larger or multimodal variants may offer capabilities beyond this model, but the supplied research does not provide enough detailed specifications for a numerical comparison. The clear positioning of ERNIE-4.5-0.3B is as the compact dense text model: easier to deploy and adapt, but less capable on demanding tasks than larger alternatives.

Bottom line

ERNIE-4.5-0.3B is a practical open-weight option for developers who need a small text-generation model that can run locally and be adapted with Baidu’s tooling. Its approximately 0.36 billion parameters, Apache 2.0 license, documented 131,072-token context, and support for SFT, LoRA, and DPO make it useful for research, education, prototyping, and focused language applications. Its main trade-off is limited capacity: it should not be selected for frontier reasoning, demanding coding, multimodal work, or applications that require a documented hosted API and guaranteed high reliability.


Answers to Frequently Asked Questions

Does ERNIE-4.5-0.3B support fine-tuning?
Yes. ERNIEKit documents supervised fine-tuning, LoRA-based parameter-efficient fine-tuning, and direct preference optimization for the model. These methods can adapt it to specialized domains while reducing the resources needed compared with retraining a larger model from scratch.
What are the best use cases and limitations of ERNIE-4.5-0.3B?
It is well suited to local text generation, rewriting, summarization, lightweight chat, high-volume simple text processing, education, prototyping, and domain adaptation with SFT or LoRA. Its small size also reduces resource requirements. However, it is less suitable for difficult reasoning, advanced coding, consistently reliable knowledge work, multimodal tasks, and applications requiring native web search or tool calling.
Can ERNIE-4.5-0.3B run locally, and what license does it use?
Yes. ERNIE-4.5-0.3B is available as an open-weight model for local deployment under the Apache License 2.0. Developers can use PaddlePaddle and FastDeploy through ERNIE-4.5-0.3B-Paddle or use the Transformers-compatible ERNIE-4.5-0.3B-PT checkpoint.
What is ERNIE-4.5-0.3B?
ERNIE-4.5-0.3B is a compact dense causal language model from Baidu’s ERNIE 4.5 family. It is designed for text generation, completion, instruction following, and lightweight conversational applications, with approximately 0.36 billion parameters.
What is the context length of ERNIE-4.5-0.3B?
The model has a documented maximum context length of 131,072 tokens. The practical usable limit may be lower depending on the hardware, inference settings, memory availability, and serving framework used for deployment.


Sources 5
Provider

About Baidu