Granite Vision

Granite Vision 3.3 2B

by IBM watsonx · Available; legacy relative to Granite Vision 4.0 3B Vision

IBM Granite Vision 3.3 2B is an Apache 2.0 open-weight vision-language model for visual document understanding. It analyzes English instructions and PNG or JPEG images, supports chart and table extraction, offers experimental segmentation and doctags capabilities, and provides up to 131,072 tokens of context. It is compact enough for local or dedicated deployment, while model-specific hosted pricing and maximum output limits remain unverified.

Text Reasoning Coding
Granite Vision 3.3 2B is a compact IBM Granite vision-language model released on June 11, 2025. It combines a SigLIP2 image encoder with a Granite language model to interpret visual documents and answer questions about their contents. The model is designed less as a general conversational assistant and more as a focused document-understanding component that can run locally, in compatible inference frameworks, or through supported IBM watsonx.ai deployment environments.
Outputs

What Granite Vision 3.3 2B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

5/10 Reasoning
2/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Granite Vision
Model type Multimodal
Context window 131K tokens
Release date 2025-06-11
Status Available; legacy relative to Granite Vision 4.0 3B Vision
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was found in the available IBM or model-card documentation.

Model notes

Canonical Hugging Face model ID: ibm-granite/granite-vision-3.3-2b. The model accepts English instructions and PNG or JPEG images and returns text. IBM documents experimental image segmentation, doctags generation, and multi-page support for up to eight pages. It is released under Apache 2.0 and can be run locally with Transformers or compatible inference frameworks. IBM documentation lists watsonx.ai deployment availability with deployment-specific hourly billing, but no model-specific public token price was verified. The Hugging Face model card identifies Granite Vision 4.0 3B Vision as a newer version.

Model guide

Granite Vision 3.3 2B: IBM’s Compact Model for Document Understanding

Granite Vision 3.3 2B is IBM’s open-weight 2-billion-parameter vision-language model for analyzing business documents, charts, tables, diagrams, forms, and other images. It accepts English instructions and PNG or JPEG images and produces text, making it a practical option for document question answering, OCR-oriented analysis, information extraction, and multimodal retrieval-augmented generation.

What Granite Vision 3.3 2B is

Granite Vision 3.3 2B is an open-weight vision-language model from IBM. A vision-language model accepts both text and images, then produces a text response based on the two inputs. In this case, the image might be a scanned page, financial chart, table, technical diagram, form, infographic, or another business document.

The model has approximately 2 billion parameters, which places it in a relatively compact part of the vision-language model market. Its smaller size is relevant for teams that want to run a model on their own infrastructure or use a dedicated deployment without the resource requirements associated with much larger multimodal systems.

Granite Vision 3.3 2B is primarily a document-understanding model. It is not an image-generation, audio, or video model, and it is not intended to replace a text-only Granite model for ordinary language workloads.

Primary purpose and capabilities

The model is intended to extract and interpret information whose meaning depends on both the words in a document and their visual arrangement. For example, a user can ask it to identify a value in a table, explain a chart, locate information in a form, or answer a question about a diagram.

Supported and documented use cases include:

  • Visual question answering about images and documents
  • Document question answering
  • OCR-oriented analysis and text extraction
  • Chart, plot, and table interpretation
  • Information extraction from complex layouts
  • Understanding diagrams and infographics
  • Experimental image segmentation
  • Experimental doctags generation
  • Multi-page document question answering for up to eight pages

The model accepts English instructions and PNG or JPEG images. Its output is text. The documented capabilities therefore support workflows such as extracting structured facts from visual documents, but they do not establish native structured-output enforcement, function calling, or direct generation of images and other media.

Architecture and context limits

Granite Vision 3.3 2B uses a SigLIP2 vision encoder, a two-layer multilayer-perceptron vision-language connector, and a Granite language model derived from Granite 3.1 2B Instruct. In simple terms, the vision encoder turns image content into representations that the language model can use when generating an answer.

IBM describes the model as being built on the LLaVA family of multimodal architectures. It uses multi-layer vision features and a denser image grid intended to help with detailed document content. This design is particularly relevant when the important information is small, distributed across a page, or embedded in a chart or table rather than presented as a single block of text.

The documented context limit is 131,072 tokens. The underlying language component is also described as supporting a 128K context window. The available research does not specify a separate maximum output-token limit, so applications should not assume a particular generation ceiling beyond the limits imposed by the selected runtime or deployment.

IBM documents multi-page question answering for up to eight pages. This page limit is a model-specific documented capability, not a general guarantee that every deployment will accept eight pages in every image size, format, or memory configuration. Image preprocessing, resolution, and the serving framework can affect practical performance.

Reported performance and safety

IBM reports improvements over earlier Granite Vision versions on several document-focused evaluations. The reported scores include 0.91 on DocVQA, 0.80 on TextVQA, 0.68 on InfoVQA, and 0.79 on OCRBench. These are provider-reported benchmark results rather than independent editorial ratings, and benchmark performance should not be treated as a guarantee for a particular document collection.

The benchmark mix helps explain the model’s positioning. DocVQA, TextVQA, InfoVQA, and OCRBench measure different aspects of answering questions about visual documents and recognizing or interpreting text in images. They are more relevant to document extraction than a general conversational-vision score would be.

IBM also reports safety alignment improvements over earlier versions on RTVLM and VLGuard evaluations. As with other generative models, Granite Vision 3.3 2B can still produce inaccurate, biased, or unwanted outputs. Extracted values from financial, legal, medical, or operational documents should be checked against the source image before they are used in an automated decision or stored as authoritative data.

Deployment, licensing, and pricing

Granite Vision 3.3 2B is available as an Apache 2.0 open-weight model through IBM Granite’s Hugging Face organization. Its canonical model identifier is ibm-granite/granite-vision-3.3-2b. The model can be run locally with Transformers and can also be deployed with compatible inference frameworks such as vLLM, subject to the framework’s support and the available hardware.

IBM provides a fine-tuning notebook for adapting the model to specialized tasks. Fine-tuning may be useful when a document workflow uses a consistent form layout, domain-specific terminology, or a narrow extraction objective. The supplied research does not specify hardware requirements, training costs, or guaranteed fine-tuning outcomes.

IBM documentation lists the model as available in watsonx.ai deployment environments. No verified model-specific public input-token or output-token price was found in the supplied sources. Hosted use is described as deployment-specific hourly billing rather than a clearly published token price. Actual cost will therefore depend on the selected IBM environment, deployment configuration, runtime duration, and infrastructure. Local deployment avoids a per-token hosted price but transfers infrastructure, maintenance, and operational costs to the user.

Strengths and trade-offs

The main strength of Granite Vision 3.3 2B is specialization. A model built around document understanding can be a sensible choice when the task involves tables, charts, forms, diagrams, or scanned pages rather than open-ended image conversation. Its compact parameter count can also make local or dedicated deployment more practical than using a much larger vision-language model.

The Apache 2.0 license and open-weight distribution give organizations more deployment flexibility than a hosted-only model. Teams can evaluate the model in their own environment, integrate it into a document pipeline, or adapt it for a specialized task. IBM’s enterprise ecosystem provides an additional route through watsonx.ai for organizations that prefer managed deployment.

There are important limitations. The model returns text only and does not generate images, audio, or video. The research does not verify built-in tool use, function calling, web search, streaming, caching, batch API access, or a dedicated JSON mode. Developers can build surrounding application logic, but those application features should not be confused with native model capabilities.

Its experimental segmentation and doctags features may require careful prompting, image preparation, and downstream processing. A document-understanding model can also make plausible extraction errors, especially when an image is low resolution, text is very small, layouts are unusual, or the answer depends on information that is difficult to read visually.

Reasoning, coding, and speed characteristics

Granite Vision 3.3 2B is not documented as a dedicated reasoning model. Its reasoning capability is best understood as task-focused visual interpretation: it can connect an instruction with visual evidence and produce an answer about that evidence, but the supplied research does not establish advanced reasoning modes or a special reasoning budget.

Coding is not a primary use case. The model may be able to produce short text or code-like responses when prompted, but the available information does not verify specialized coding training or coding-agent features. For ordinary programming tasks, a text-focused code model is likely to be more appropriate.

As an editorial assessment, the model’s compact size suggests a favorable speed and cost position compared with larger vision-language systems, especially for local or dedicated inference. That is a practical trade-off rather than a provider-published universal speed claim: real throughput depends on hardware, image resolution, quantization, batching, context usage, and serving software. The smaller model may be attractive when document throughput and deployment cost matter more than maximum general-purpose visual capability.

Best use cases

Granite Vision 3.3 2B is a good candidate for workflows such as:

  • Extracting line items, totals, labels, or dates from business documents
  • Answering questions about charts, plots, and tables
  • Creating text descriptions of diagrams and infographics
  • Building a document-grounded assistant for internal records
  • Adding visual retrieval-augmented generation to an enterprise search system
  • Processing forms or recurring document layouts on controlled infrastructure
  • Prototyping multimodal document pipelines with an open-weight model

For production extraction, the model should normally be paired with validation. A useful pipeline can preserve the source image, record the model’s response, apply field-level checks, and route uncertain results for human review. This is especially important when extracted information affects payments, compliance, eligibility, or customer records.

When to choose this model

Choose Granite Vision 3.3 2B when the central problem is understanding visual documents and you value an open-weight Apache 2.0 model, compact deployment, and IBM-oriented enterprise integration. It is particularly compelling when running the model locally or in a controlled environment is more important than using the largest available multimodal model.

Consider another option when the workload is primarily text-only, requires advanced coding assistance, depends on verified tool or function calling, needs image or audio generation, or demands consistently strong performance on broad visual reasoning tasks outside document analysis. IBM’s newer Granite Vision 4.0 3B Vision is identified in the model card as a newer version, so teams starting a new evaluation should compare it with Granite Vision 3.3 2B rather than assuming the older model is the default choice.

Overall, Granite Vision 3.3 2B is best viewed as a focused visual document component rather than a universal multimodal assistant. Its value comes from combining document-oriented capabilities, relatively compact open-weight deployment, and support for enterprise extraction scenarios, while its limitations are most visible in general-purpose reasoning, native application tools, and non-text media generation.


Answers to Frequently Asked Questions

What are the main limitations of Granite Vision 3.3 2B?
The model produces text only and is not intended for image, audio, or video generation. It is not documented as a dedicated reasoning or coding model, and the supplied research does not verify native function calling, tool use, web search, streaming, caching, batch APIs, or dedicated JSON mode. Because it can make visual extraction errors, important outputs should be checked against the original document.
How can Granite Vision 3.3 2B be deployed and licensed?
Granite Vision 3.3 2B is distributed as an Apache 2.0 open-weight model through IBM Granite’s Hugging Face organization. Its canonical identifier is ibm-granite/granite-vision-3.3-2b. It can be run locally with Transformers, deployed with compatible frameworks such as vLLM, or used in supported watsonx.ai environments.
What is Granite Vision 3.3 2B?
Granite Vision 3.3 2B is an open-weight vision-language model from IBM designed primarily for understanding documents, scanned pages, charts, tables, forms, diagrams, and other business images. It accepts English instructions and PNG or JPEG images, then produces text responses.
What can Granite Vision 3.3 2B be used for?
It can answer questions about documents and images, extract text and structured facts, interpret charts and tables, understand diagrams and infographics, and support OCR-oriented workflows. IBM also documents multi-page document question answering for up to eight pages, along with experimental image segmentation and doctags generation.


Sources 5
Provider

About IBM watsonx