ERNIE 4.5 VL

ERNIE-4.5-VL-28B-A3B

by Baidu · Current open-weight model; publicly released under the Apache 2.0 license

Baidu's ERNIE-4.5-VL-28B-A3B is a 28B-total-parameter, 3B-active-parameter open-weight vision-language MoE model. It supports text, image, and video input, text output, a 131K context window, local deployment, and fine-tuning through Baidu's ERNIEKit ecosystem.

Text Reasoning Coding
Released on June 30, 2025, ERNIE-4.5-VL-28B-A3B is Baidu's lightweight ERNIE 4.5 vision-language model for visual question answering, document and chart understanding, image analysis, and multimodal reasoning. Its sparse MoE design reduces active computation while retaining a substantially larger total parameter capacity.
Outputs

What ERNIE-4.5-VL-28B-A3B can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
6/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family ERNIE 4.5 VL
Model type Multimodal
Context window 131K tokens
Release date 2025-06-30
Status Current open-weight model; publicly released under the Apache 2.0 license
Knowledge cutoff notes

No authoritative knowledge-cutoff date for the exact model was identified in the official model card, release announcement, or repository documentation reviewed.

Model notes

ERNIE-4.5-VL-28B-A3B is a post-trained multimodal Mixture-of-Experts chat model with 28B total parameters and approximately 3B activated parameters per token. Official model materials describe text and vision inputs, while the official processor examples also expose video processing. The model produces text rather than images, audio, or video. Baidu released the weights and inference code under Apache 2.0. Official deployment documentation supports Hugging Face Transformers, vLLM, PaddlePaddle, and FastDeploy; single-card deployment documentation states that at least 80GB of GPU memory is required. The model family supports thinking and non-thinking modes, but the separately named ERNIE-4.5-VL-28B-A3B-Thinking model should not be treated as the same exact model record. Editorial scores are comparative estimates, not vendor-provided ratings. No authoritative hosted API price, maximum output-token limit, knowledge cutoff, JSON-mode guarantee, prompt-caching feature, or batch API for this exact model was verified.

Model guide

ERNIE-4.5-VL-28B-A3B: Open-Weight Vision-Language Model for Efficient Visual Reasoning

Baidu's ERNIE-4.5-VL-28B-A3B is an open-weight multimodal Mixture-of-Experts model with 28 billion total parameters and approximately 3 billion activated parameters per token. It accepts text, images, and video and produces text, with a 131,072-token context window.

What is ERNIE-4.5-VL-28B-A3B?

ERNIE-4.5-VL-28B-A3B is Baidu's open-weight vision-language model for applications that need to interpret text together with visual information. It can process text, images, and video, then respond with text. Typical tasks include asking questions about an image, extracting meaning from documents, interpreting charts, analyzing visual scenes, and examining video content.

The model was released on June 30, 2025, as part of Baidu's ERNIE 4.5 family. It is not a general-purpose image, video, or audio generation system: its verified output is text. Baidu released the model weights and inference code under the Apache 2.0 license, making it suitable for local or self-hosted deployment subject to the license and the operator's infrastructure.

The name describes its architecture and scale. The model has 28 billion total parameters, but its Mixture-of-Experts (MoE) design activates approximately 3 billion parameters per token. In practical terms, the model retains a larger pool of learned capacity while using a smaller active portion for each part of an inference request. This is the main reason it can occupy a position between compact vision-language models and much larger dense models.

Where it fits in Baidu's ERNIE lineup

ERNIE-4.5-VL-28B-A3B is the lightweight 28B-total-parameter, approximately 3B-active-parameter vision-language member of the ERNIE 4.5 family. Its focus is multimodal understanding rather than consumer chat features or native media generation. The family also includes separately named variants, including ERNIE-4.5-VL-28B-A3B-Thinking. That Thinking model should not be treated as the same model record, even though the standard model family supports thinking and non-thinking modes.

This distinction matters when selecting weights or comparing deployment requirements. A reference to ERNIE-4.5-VL-28B-A3B does not automatically establish the specifications, behavior, or resource requirements of another ERNIE 4.5 variant.

Inputs, outputs, and context window

The verified input modalities are text, images, and video. The official processor examples expose video processing in addition to the model documentation's text-and-vision descriptions. The model produces text only; it does not directly return images, audio, or video.

CapabilityVerified status
Text inputSupported
Image inputSupported
Video inputSupported through the official processing examples
Text outputSupported
Image, audio, or video outputNot supported as native model output
Context window131,072 tokens
Maximum output tokensNot verified in the supplied official materials

A 131,072-token context window is useful for long documents, extended visual-grounding tasks, and prompts that combine instructions with substantial text and media references. It does not mean that every deployment can accept an arbitrarily large file or video: practical limits may also depend on the processor, image or video representation, available GPU memory, and the serving framework. No authoritative maximum output-token limit for this exact model was verified.

Architecture and computational trade-off

The model's sparse MoE architecture is its most important technical distinction. A conventional dense model uses the full parameter set for every token. In an MoE model, different expert components can be selected for different tokens. ERNIE-4.5-VL-28B-A3B has 28 billion parameters in total but activates approximately 3 billion per token, according to the supplied model information.

This design can reduce active computation compared with a dense model of the same total size, although it does not eliminate the need for substantial memory. Baidu's deployment documentation states that single-card deployment requires at least 80 GB of GPU memory. The actual hardware requirement can vary with quantization, framework, batch size, context length, and serving configuration; the 80 GB figure should therefore be read as the documented baseline for the referenced deployment approach rather than a universal requirement for every possible setup.

The model is supported in deployment documentation for Hugging Face Transformers, vLLM, PaddlePaddle, and FastDeploy. These options make it more practical for teams that want to run the model themselves instead of relying on a provider-managed endpoint. They also place responsibility for GPU capacity, scaling, monitoring, access control, and model updates on the operator.

Reasoning, coding, and tool support

The ERNIE 4.5 VL family supports thinking and non-thinking modes. Thinking mode is intended for requests where additional intermediate reasoning may help, while non-thinking mode is more appropriate when response latency is the priority. The supplied materials do not provide a single universal quality guarantee for either mode, so the choice should be tested against the application's prompts and latency requirements.

Editorial evaluation rates the model's reasoning at 7 out of 10 and coding at 6 out of 10. These are comparative editorial estimates, not Baidu-published benchmark scores. The coding assessment should not be interpreted as evidence that this is primarily a code-generation model. Its strongest supported role is multimodal understanding, such as interpreting a software screenshot, reading a technical diagram, or answering questions about a document that contains code and visuals.

Tool use and function-calling support were not verified for this exact model. Likewise, no provider-managed web-search capability was verified. An application can potentially build external tools around a self-hosted model, but that would be an application-level integration rather than a confirmed native capability of ERNIE-4.5-VL-28B-A3B.

Main strengths and limitations

Strengths for visual understanding

  • Broad multimodal input: the model accepts text, images, and video, allowing applications to combine written instructions with visual evidence.
  • Long context: the 131,072-token context window is well suited to lengthy documents, chart collections, and multi-part analysis tasks.
  • Efficient active computation: the approximately 3B active-parameter MoE design offers a lower active-parameter profile than its 28B total size might suggest.
  • Deployment flexibility: open weights, Apache 2.0 licensing, and support for several inference stacks enable local and self-hosted use.
  • Visual document use cases: the model is positioned for visual question answering, document understanding, chart understanding, image analysis, and video understanding.

Limitations to plan for

  • Hardware demands remain significant: the documented single-card deployment baseline requires at least 80 GB of GPU memory, so this is not a lightweight model in the everyday laptop sense.
  • Text-only output: it cannot be used as a native image, audio, or video generator.
  • Unverified serving features: no authoritative hosted API price, maximum output-token limit, JSON-mode guarantee, prompt caching, or batch API was verified for this exact model.
  • No confirmed native web search or tool calling: applications requiring provider-managed retrieval or function execution should verify those features separately or add them externally.
  • Deployment complexity: self-hosting requires operators to manage compatible GPUs, inference software, processor behavior, scaling, and security.

Pricing and access

No authoritative hosted API token price was verified for ERNIE-4.5-VL-28B-A3B in the supplied research. It should not be assigned a per-input-token or per-output-token price without a published endpoint and pricing page for this exact model.

The model's open-weight Apache 2.0 release changes the cost question rather than making deployment free. There may be no model-license fee under the stated license, but users still incur costs for GPUs, storage, electricity, cloud instances, engineering, monitoring, and any surrounding services. Teams comparing this model with a hosted multimodal API should compare total operating cost and maintenance effort, not only nominal token pricing.

Best use cases

ERNIE-4.5-VL-28B-A3B is a strong candidate for teams that need self-hosted multimodal analysis and can provide the required hardware. Practical applications include:

  • Visual question-answering systems for images, screenshots, or photographed materials.
  • Document assistants that interpret text alongside page layout, figures, and tables.
  • Chart and diagram analysis for research, reporting, and business workflows.
  • Image inspection and classification workflows that need natural-language explanations.
  • Video-understanding pipelines that ask questions about supplied video content.
  • Private or controlled deployments where sending documents and visual data to a third-party hosted endpoint is undesirable.
  • Research and fine-tuning projects using Baidu's supported ERNIE ecosystem and deployment tools.

Fine-tuning is supported according to the supplied model metadata, although the research does not specify a universal fine-tuning recipe, dataset size, or hardware requirement. Those details should be checked against the selected framework and model card before planning a training project.

When to choose this model

Choose ERNIE-4.5-VL-28B-A3B when the central requirement is multimodal understanding, open-weight access, and control over deployment. It is particularly attractive when a long context window and visual document or video analysis matter more than a turnkey hosted API. Its sparse architecture may also offer a useful speed-and-compute trade-off relative to a much larger dense vision-language model, although real performance depends on hardware and serving configuration.

A hosted multimodal model may be more appropriate when a team wants predictable per-request billing, managed scaling, minimal infrastructure work, or confirmed tool and web-search integrations. A smaller vision-language model may be preferable for edge devices, low-memory environments, or very high-throughput workloads where the documented 80 GB single-card baseline is impractical. A separate image or video generation model is the correct choice when the required output is media rather than text.

For difficult reasoning tasks, users should compare the standard model's thinking and non-thinking modes with the separately named ERNIE-4.5-VL-28B-A3B-Thinking variant rather than assuming they are interchangeable. The current model remains the better-defined choice when the requirement is an open-weight, text-output vision-language model that can be deployed through established inference frameworks.

Bottom line

ERNIE-4.5-VL-28B-A3B combines a 28B total-parameter MoE architecture with approximately 3B active parameters per token, a 131,072-token context window, and text, image, and video input. Its main value is controlled, self-hosted multimodal understanding under the Apache 2.0 license. Its main trade-offs are substantial hardware requirements, text-only output, and the absence of verified hosted pricing and several managed API features. For organizations prepared to operate the infrastructure, it offers a practical open-weight option for visual documents, charts, images, and video analysis.


Answers to Frequently Asked Questions

What are the best use cases for ERNIE-4.5-VL-28B-A3B?
The model is well suited to self-hosted multimodal applications such as visual question answering, document and chart analysis, image inspection, screenshot interpretation, diagram understanding, and video analysis. It is especially useful when organizations need deployment control or want to keep sensitive visual data away from third-party hosted APIs.
Is ERNIE-4.5-VL-28B-A3B open source, and how can it be deployed?
Baidu released the model weights and inference code under the Apache 2.0 license. It can be deployed locally or in a self-hosted environment using supported frameworks including Hugging Face Transformers, vLLM, PaddlePaddle, and FastDeploy.
How many parameters does ERNIE-4.5-VL-28B-A3B activate, and what hardware does it require?
The model has 28 billion total parameters and activates approximately 3 billion parameters per token through its Mixture-of-Experts architecture. Baidu's deployment documentation lists at least 80 GB of GPU memory for single-card deployment, although requirements can vary with quantization, context length, batch size, and serving configuration.
What is ERNIE-4.5-VL-28B-A3B?
ERNIE-4.5-VL-28B-A3B is Baidu's open-weight vision-language model for analyzing text, images, and video and producing text responses. It is designed for tasks such as visual question answering, document understanding, chart analysis, image interpretation, and video understanding.
What inputs and outputs does ERNIE-4.5-VL-28B-A3B support?
The model supports text, image, and video inputs and produces text outputs. It does not natively generate images, audio, or video. Its verified context window is 131,072 tokens.


Sources 6
Provider

About Baidu