What is ERNIE-4.5-VL-28B-A3B?
ERNIE-4.5-VL-28B-A3B is Baidu's open-weight vision-language model for applications that need to interpret text together with visual information. It can process text, images, and video, then respond with text. Typical tasks include asking questions about an image, extracting meaning from documents, interpreting charts, analyzing visual scenes, and examining video content.
The model was released on June 30, 2025, as part of Baidu's ERNIE 4.5 family. It is not a general-purpose image, video, or audio generation system: its verified output is text. Baidu released the model weights and inference code under the Apache 2.0 license, making it suitable for local or self-hosted deployment subject to the license and the operator's infrastructure.
The name describes its architecture and scale. The model has 28 billion total parameters, but its Mixture-of-Experts (MoE) design activates approximately 3 billion parameters per token. In practical terms, the model retains a larger pool of learned capacity while using a smaller active portion for each part of an inference request. This is the main reason it can occupy a position between compact vision-language models and much larger dense models.
Where it fits in Baidu's ERNIE lineup
ERNIE-4.5-VL-28B-A3B is the lightweight 28B-total-parameter, approximately 3B-active-parameter vision-language member of the ERNIE 4.5 family. Its focus is multimodal understanding rather than consumer chat features or native media generation. The family also includes separately named variants, including ERNIE-4.5-VL-28B-A3B-Thinking. That Thinking model should not be treated as the same model record, even though the standard model family supports thinking and non-thinking modes.
This distinction matters when selecting weights or comparing deployment requirements. A reference to ERNIE-4.5-VL-28B-A3B does not automatically establish the specifications, behavior, or resource requirements of another ERNIE 4.5 variant.
Inputs, outputs, and context window
The verified input modalities are text, images, and video. The official processor examples expose video processing in addition to the model documentation's text-and-vision descriptions. The model produces text only; it does not directly return images, audio, or video.
| Capability | Verified status |
|---|---|
| Text input | Supported |
| Image input | Supported |
| Video input | Supported through the official processing examples |
| Text output | Supported |
| Image, audio, or video output | Not supported as native model output |
| Context window | 131,072 tokens |
| Maximum output tokens | Not verified in the supplied official materials |
A 131,072-token context window is useful for long documents, extended visual-grounding tasks, and prompts that combine instructions with substantial text and media references. It does not mean that every deployment can accept an arbitrarily large file or video: practical limits may also depend on the processor, image or video representation, available GPU memory, and the serving framework. No authoritative maximum output-token limit for this exact model was verified.
Architecture and computational trade-off
The model's sparse MoE architecture is its most important technical distinction. A conventional dense model uses the full parameter set for every token. In an MoE model, different expert components can be selected for different tokens. ERNIE-4.5-VL-28B-A3B has 28 billion parameters in total but activates approximately 3 billion per token, according to the supplied model information.
This design can reduce active computation compared with a dense model of the same total size, although it does not eliminate the need for substantial memory. Baidu's deployment documentation states that single-card deployment requires at least 80 GB of GPU memory. The actual hardware requirement can vary with quantization, framework, batch size, context length, and serving configuration; the 80 GB figure should therefore be read as the documented baseline for the referenced deployment approach rather than a universal requirement for every possible setup.
The model is supported in deployment documentation for Hugging Face Transformers, vLLM, PaddlePaddle, and FastDeploy. These options make it more practical for teams that want to run the model themselves instead of relying on a provider-managed endpoint. They also place responsibility for GPU capacity, scaling, monitoring, access control, and model updates on the operator.
Reasoning, coding, and tool support
The ERNIE 4.5 VL family supports thinking and non-thinking modes. Thinking mode is intended for requests where additional intermediate reasoning may help, while non-thinking mode is more appropriate when response latency is the priority. The supplied materials do not provide a single universal quality guarantee for either mode, so the choice should be tested against the application's prompts and latency requirements.
Editorial evaluation rates the model's reasoning at 7 out of 10 and coding at 6 out of 10. These are comparative editorial estimates, not Baidu-published benchmark scores. The coding assessment should not be interpreted as evidence that this is primarily a code-generation model. Its strongest supported role is multimodal understanding, such as interpreting a software screenshot, reading a technical diagram, or answering questions about a document that contains code and visuals.
Tool use and function-calling support were not verified for this exact model. Likewise, no provider-managed web-search capability was verified. An application can potentially build external tools around a self-hosted model, but that would be an application-level integration rather than a confirmed native capability of ERNIE-4.5-VL-28B-A3B.
Main strengths and limitations
Strengths for visual understanding
- Broad multimodal input: the model accepts text, images, and video, allowing applications to combine written instructions with visual evidence.
- Long context: the 131,072-token context window is well suited to lengthy documents, chart collections, and multi-part analysis tasks.
- Efficient active computation: the approximately 3B active-parameter MoE design offers a lower active-parameter profile than its 28B total size might suggest.
- Deployment flexibility: open weights, Apache 2.0 licensing, and support for several inference stacks enable local and self-hosted use.
- Visual document use cases: the model is positioned for visual question answering, document understanding, chart understanding, image analysis, and video understanding.
Limitations to plan for
- Hardware demands remain significant: the documented single-card deployment baseline requires at least 80 GB of GPU memory, so this is not a lightweight model in the everyday laptop sense.
- Text-only output: it cannot be used as a native image, audio, or video generator.
- Unverified serving features: no authoritative hosted API price, maximum output-token limit, JSON-mode guarantee, prompt caching, or batch API was verified for this exact model.
- No confirmed native web search or tool calling: applications requiring provider-managed retrieval or function execution should verify those features separately or add them externally.
- Deployment complexity: self-hosting requires operators to manage compatible GPUs, inference software, processor behavior, scaling, and security.
Pricing and access
No authoritative hosted API token price was verified for ERNIE-4.5-VL-28B-A3B in the supplied research. It should not be assigned a per-input-token or per-output-token price without a published endpoint and pricing page for this exact model.
The model's open-weight Apache 2.0 release changes the cost question rather than making deployment free. There may be no model-license fee under the stated license, but users still incur costs for GPUs, storage, electricity, cloud instances, engineering, monitoring, and any surrounding services. Teams comparing this model with a hosted multimodal API should compare total operating cost and maintenance effort, not only nominal token pricing.
Best use cases
ERNIE-4.5-VL-28B-A3B is a strong candidate for teams that need self-hosted multimodal analysis and can provide the required hardware. Practical applications include:
- Visual question-answering systems for images, screenshots, or photographed materials.
- Document assistants that interpret text alongside page layout, figures, and tables.
- Chart and diagram analysis for research, reporting, and business workflows.
- Image inspection and classification workflows that need natural-language explanations.
- Video-understanding pipelines that ask questions about supplied video content.
- Private or controlled deployments where sending documents and visual data to a third-party hosted endpoint is undesirable.
- Research and fine-tuning projects using Baidu's supported ERNIE ecosystem and deployment tools.
Fine-tuning is supported according to the supplied model metadata, although the research does not specify a universal fine-tuning recipe, dataset size, or hardware requirement. Those details should be checked against the selected framework and model card before planning a training project.
When to choose this model
Choose ERNIE-4.5-VL-28B-A3B when the central requirement is multimodal understanding, open-weight access, and control over deployment. It is particularly attractive when a long context window and visual document or video analysis matter more than a turnkey hosted API. Its sparse architecture may also offer a useful speed-and-compute trade-off relative to a much larger dense vision-language model, although real performance depends on hardware and serving configuration.
A hosted multimodal model may be more appropriate when a team wants predictable per-request billing, managed scaling, minimal infrastructure work, or confirmed tool and web-search integrations. A smaller vision-language model may be preferable for edge devices, low-memory environments, or very high-throughput workloads where the documented 80 GB single-card baseline is impractical. A separate image or video generation model is the correct choice when the required output is media rather than text.
For difficult reasoning tasks, users should compare the standard model's thinking and non-thinking modes with the separately named ERNIE-4.5-VL-28B-A3B-Thinking variant rather than assuming they are interchangeable. The current model remains the better-defined choice when the requirement is an open-weight, text-output vision-language model that can be deployed through established inference frameworks.
Bottom line
ERNIE-4.5-VL-28B-A3B combines a 28B total-parameter MoE architecture with approximately 3B active parameters per token, a 131,072-token context window, and text, image, and video input. Its main value is controlled, self-hosted multimodal understanding under the Apache 2.0 license. Its main trade-offs are substantial hardware requirements, text-only output, and the absence of verified hosted pricing and several managed API features. For organizations prepared to operate the infrastructure, it offers a practical open-weight option for visual documents, charts, images, and video analysis.

