Granite Vision

Granite 4.0 3B Vision

by IBM watsonx · Current and downloadable; newer Granite Vision 4.1 4B version available

IBM Granite 4.0 3B Vision is a compact Apache 2.0 vision-language model for extracting charts, tables, and semantic key-value pairs from enterprise document images. It supports PNG and JPEG inputs, structured text outputs, and local deployment through tools such as Transformers, vLLM, SGLang, and Docker Model Runner. No official hosted-token pricing or exact-model managed API was verified.

Text Reasoning Coding
Granite 4.0 3B Vision is a compact vision-language model from IBM Research for turning document images into structured data. It accepts English instructions with PNG or JPEG images and can extract chart data, table structures, and semantic key-value pairs in formats such as CSV, JSON, HTML, OTSL, Python code, or plain text. Released under Apache 2.0, it is intended primarily for local or self-hosted deployment, so users should plan for their own hardware and inference infrastructure rather than expecting a standard hosted API with token pricing.
Outputs

What Granite 4.0 3B Vision can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Structured output
Model profile

Performance characteristics

5/10 Reasoning
4/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Granite Vision
Model type Multimodal
Release date 2026-03-27
Status Current and downloadable; newer Granite Vision 4.1 4B version available
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was published for this exact model.

Model notes

IBM Research model released under Apache 2.0. The model is distributed as a LoRA adapter on Granite 4.0 Micro, described as approximately 3.5B base-model parameters plus 0.5B adapter parameters. It supports English instructions with PNG or JPEG images and specialized task tags for chart, table, and KVP extraction. Outputs can include CSV, Python code, JSON, HTML, OTSL, or natural-language text. The model card identifies Granite Vision 4.1 4B as a newer version. No official IBM hosted-token pricing or provider-managed API endpoint was found for this exact model.

Cost

Model pricing

Input No official hosted API price found; open-weight model intended for self-hosted deployment
Output No official hosted API price found; open-weight model intended for self-hosted deployment
Model guide

Granite 4.0 3B Vision: IBM’s Open Model for Structured Document Extraction

Granite 4.0 3B Vision is an open-weight IBM vision-language model focused on extracting structured information from charts, tables, invoices, and other enterprise documents rather than serving as a general-purpose visual chatbot.

What Granite 4.0 3B Vision is

Granite 4.0 3B Vision is an open-weight vision-language model developed by IBM Research. A vision-language model accepts both visual input and text instructions, then produces text or structured representations based on what it sees. In this case, the model is specialized for document understanding: it is designed to read charts, tables, invoices, and other enterprise documents and convert their contents into data that software can use.

This specialization is important. Granite 4.0 3B Vision is not positioned primarily as an unrestricted visual assistant for long conversations, creative image interpretation, or broad multilingual reasoning. Its strongest use case is a document-processing pipeline in which an image needs to become a CSV file, JSON object, HTML table, chart summary, or set of extracted fields.

The model was released on March 27, 2026. It remains independently available and downloadable, while IBM's model documentation identifies Granite Vision 4.1 4B as a newer version in the same general family. That newer model is a successor or alternative, not another name for Granite 4.0 3B Vision.

Document tasks and output formats

Granite 4.0 3B Vision uses task-specific prompt tags to select different extraction modes. This gives an application a more explicit way to request the form of output it needs instead of relying only on an open-ended instruction.

  • Chart extraction: convert chart content into CSV data, produce a written chart summary, or generate Python code that recreates the chart.
  • Table extraction: return a table as JSON, HTML, or OTSL markup.
  • Semantic key-value extraction: identify fields and values across varying document layouts, such as invoice numbers, dates, totals, or company names.
  • Image-to-text description: provide a general textual description of an image.

These output options make the model relevant to document ingestion, visual retrieval-augmented generation, chart analysis, and automated back-office workflows. For example, a finance system could submit an image of a report and request chart data as CSV, while an invoice pipeline could provide a schema and request semantic key-value pairs.

Structured output support here should be understood as a model behavior enabled by prompts and task formats. The supplied research does not verify a separate provider-managed JSON mode, function-calling interface, or guaranteed schema-validation service.

Architecture and model size

IBM describes Granite 4.0 3B Vision as a combination of the Granite 4.0 Micro language model, a SigLIP2 vision encoder, and a LoRA adapter. LoRA, or Low-Rank Adaptation, is a way to add task-specific model behavior through additional trainable weights rather than replacing every parameter in the base model.

The vision system uses 384-by-384 image tiling, windowed Q-Former projectors, and DeepStack-style visual feature injection into the language model. In practical terms, these components convert visual features into information the language model can use while generating an answer. IBM describes the configuration as approximately 3.5 billion base-model parameters plus approximately 0.5 billion LoRA-adapter parameters.

The published model is distributed as a LoRA adapter on top of the Granite 4.0 Micro base model. Text-only requests can use the base model without loading the vision adapter, while image requests apply the multimodal adaptation. This design is relevant to deployment planning because the model is not simply a single conventional checkpoint with an independently stated parameter count.

Supported inputs and outputs

The verified input combination is an English instruction together with a PNG or JPEG image. The model produces text-based results, including natural-language descriptions and structured text formats. It does not generate images, audio, or video.

CapabilityVerified status
Text inputSupported
Image inputSupported for PNG and JPEG images
Audio inputNot supported
Video inputNot supported
Text outputSupported
Image, audio, or video outputNot supported
Context lengthNot published in the supplied research
Maximum output tokensNot published in the supplied research

The absence of a published context or output limit means applications should not assume that the model can process arbitrarily large documents or return arbitrarily long results. Large pages may need to be resized, cropped, split into sections, or processed with a document-layout pipeline before inference. The supplied research does not establish a specific image-resolution limit beyond the documented 384-by-384 tiling approach.

Deployment, licensing, and pricing

Granite 4.0 3B Vision is released under the Apache 2.0 license and is available for download from IBM's Granite organization on Hugging Face. IBM documents usage paths involving Transformers, vLLM, SGLang, Docker Model Runner, and compatible local tooling.

This is an important distinction from a typical commercial multimodal API. No official IBM hosted-token price was found for this exact model, and the research does not identify a provider-managed IBM endpoint that bills per input or output token. The model card also indicated that no hosted inference provider was deploying it through Hugging Face's inference-provider system at the time of verification.

Consequently, the effective cost depends on the user's deployment. Hardware, memory, storage, electricity, orchestration, maintenance, and engineering time all matter. A self-hosted model can be attractive when document data must remain within a controlled environment or when workloads are large and predictable, but it is not automatically cheaper than an API once operational costs are included.

The Apache 2.0 license is a useful advantage for organizations that need to inspect, adapt, and deploy an open model, subject to the license terms and any obligations associated with the surrounding software stack. It does not remove the need to evaluate security, data handling, and output quality in the intended environment.

Strengths and limitations

The model's central strength is its narrow alignment with structured document extraction. Task tags and multiple output formats can make it easier to connect the model to downstream software than a general visual assistant that returns only conversational prose. Its relatively compact configuration may also be more practical for local experimentation and controlled enterprise deployments than much larger multimodal models, although the supplied research does not provide hardware requirements or benchmark results.

Another strength is deployment flexibility. Users can download the model and operate it through several established local-serving paths rather than depending on a single hosted provider. This can support document workflows where data residency, network isolation, or operational control is more important than access to a managed API.

There are also significant limitations:

  • English emphasis: IBM identifies degraded performance on non-English documents, so multilingual extraction should be validated separately.
  • Specialized scope: the model is optimized for charts, tables, semantic fields, and related document tasks, not unrestricted visual reasoning.
  • Possible hallucinations: extracted values may be incorrect or fabricated, particularly when an image is unclear or the requested structure is ambiguous.
  • Unknown hard limits: the supplied research does not publish a context window, maximum output length, or verified hardware requirement.
  • Self-hosting responsibility: users must provide and maintain inference infrastructure because no official hosted price or exact-model managed API was verified.

These limitations are especially important for invoices, financial reports, compliance records, and other high-stakes documents. Extracted values should be checked against the source image or passed through validation rules before they trigger business actions.

Reasoning, coding, and tool support

Granite 4.0 3B Vision can perform task-oriented visual reasoning: it must identify content in a chart or document and map that content into a requested structure. However, the supplied research does not establish a general reasoning mode, a reasoning-token feature, or benchmarked performance for complex multi-step reasoning.

Its coding capability is similarly specialized. Chart-to-code extraction can generate Python code intended to recreate a chart, which is useful when chart information needs to move into an analysis workflow. That should not be confused with broad software-engineering capability or a verified coding-agent feature.

Tool calling, function calling, streaming, fine-tuning, caching, and batch API support are not verified for this exact model in the supplied research. The model can generate structured text for an application to parse, but that is different from directly invoking external tools or guaranteeing that generated JSON conforms to an application schema.

When to choose Granite 4.0 3B Vision

Choose Granite 4.0 3B Vision when the primary problem is converting document images into structured information and you are prepared to run an open model yourself. It is a particularly reasonable candidate for:

  • chart-to-CSV or chart-to-Python workflows;
  • table extraction into JSON, HTML, or OTSL;
  • invoice and form field extraction using semantic key-value schemas;
  • visual retrieval pipelines that need document content in machine-readable form;
  • private or locally controlled enterprise document processing;
  • experimentation with IBM's Granite Vision model family.

Its compact, specialized design may offer a better speed-and-cost trade-off than a much larger general-purpose vision model when the task is narrowly defined. That is an editorial deployment consideration, not a published benchmark claim: actual throughput and cost depend on hardware, quantization, serving software, image size, and workload volume.

A hosted multimodal API may be more appropriate when a team wants predictable usage-based billing, managed scaling, an official endpoint, or built-in operational features. A larger general-purpose vision model may be preferable for open-ended visual reasoning, difficult layouts, broader language coverage, or conversations that go well beyond extraction. If staying within IBM's Granite Vision lineup is more important than using this exact release, Granite Vision 4.1 4B is the newer model identified by IBM's model card and should be evaluated as a possible successor.

Overall assessment

Granite 4.0 3B Vision is best understood as a focused document-understanding component rather than a general-purpose multimodal assistant. Its Apache 2.0 distribution, local deployment options, chart and table extraction modes, and semantic key-value support make it useful for organizations building controlled document pipelines.

Its value depends on matching the model to that role. It can reduce the work required to turn visual documents into structured text, but it does not eliminate validation, document preprocessing, or deployment engineering. Teams that need multilingual coverage, broad visual reasoning, guaranteed managed APIs, or verified advanced tool support should compare it with other options instead of treating its structured extraction features as evidence of general capability.


Answers to Frequently Asked Questions

What are the main limitations of Granite 4.0 3B Vision?
The model has an English emphasis and may perform worse on non-English documents. It is optimized for structured document extraction rather than unrestricted visual reasoning, and its outputs can contain errors or hallucinated values. The supplied research does not publish a context window, maximum output length, or verified hardware requirements, so extracted data should be validated before use in high-stakes workflows.
Does Granite 4.0 3B Vision have an official hosted API or token pricing?
No official IBM hosted-token price or provider-managed endpoint was verified for this exact model. Users generally need to deploy it themselves, so costs depend on hardware, storage, electricity, infrastructure, maintenance, and engineering requirements.
How is Granite 4.0 3B Vision deployed and licensed?
Granite 4.0 3B Vision is released under the Apache 2.0 license and can be downloaded from IBM's Granite organization on Hugging Face. It can be used with Transformers, vLLM, SGLang, Docker Model Runner, and compatible local tooling. The published model is a LoRA adapter applied to the Granite 4.0 Micro base model.
What is Granite 4.0 3B Vision designed to do?
Granite 4.0 3B Vision is an open-weight vision-language model from IBM Research specialized in converting charts, tables, invoices, forms, and other document images into structured text such as CSV, JSON, HTML, OTSL, or semantic key-value pairs.
What input and output formats does Granite 4.0 3B Vision support?
The model accepts English instructions together with PNG or JPEG images and produces text-based results. It supports chart extraction, table extraction, semantic key-value extraction, and image descriptions, but it does not support audio or video input or generate images, audio, or video.


Sources 3
Provider

About IBM watsonx