Granite Vision 4.1

Granite Vision 4.1 4B

by IBM watsonx · Current open-weight model

IBM Granite Vision 4.1 4B is a compact Apache 2.0 vision-language model focused on structured extraction from charts, tables, invoices, forms, and other enterprise documents. It supports image-to-text, JSON, CSV, HTML, OTSL, chart-to-code, and schema-based key-value workflows, with local deployment through compatible open-source runtimes.

Text Reasoning Coding
Granite Vision 4.1 4B is a compact vision-language model from IBM Research that focuses on document intelligence rather than open-ended visual conversation. It combines a 3.4-billion-parameter Granite language model with a 0.6-billion-parameter vision encoder and projector stack. Released under the Apache 2.0 license, it can be downloaded and deployed with compatible open-source runtimes for workflows such as table extraction, chart conversion, invoice processing, and multimodal retrieval.
Outputs

What Granite Vision 4.1 4B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming Fine-tuning Structured output
Model profile

Performance characteristics

3/10 Reasoning
4/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite Vision 4.1
Model type Multimodal
Maximum output 4K tokens
Release date 2026-04-29
Status Current open-weight model
Knowledge cutoff notes

No exact knowledge cutoff date is published in the verified IBM or IBM Granite model documentation.

Model notes

Granite Vision 4.1 4B is an IBM Research vision-language model with approximately 4B parameters: a 3.4B language model plus a 0.6B vision encoder and projector stack. It is fine-tuned from Granite 4.1 3B and released under Apache 2.0. Supported workflows include chart-to-CSV, chart-to-summary, chart-to-code, table extraction to JSON/HTML/OTSL, image-to-text description, and schema-based semantic key-value extraction. The model card reports 94.2% zero-shot exact-match accuracy on the VAREX key-value extraction benchmark. The model is specialized for English document understanding and may generalize less effectively to open-ended vision tasks or non-English documents. Structured JSON extraction is documented, but a separate provider-level JSON mode is not documented. The model is available for local deployment through Transformers, vLLM, MLX VLM, Docker-compatible runtimes, and Docling.

Cost

Model pricing

Input No official hosted API price specified; downloadable open-weight model
Output No official hosted API price specified; downloadable open-weight model
Model guide

Granite Vision 4.1 4B: IBM’s Compact Model for Structured Document Extraction

Granite Vision 4.1 4B is an open-weight IBM vision-language model designed primarily for extracting structured information from enterprise document images. It can process charts, tables, invoices, forms, and similar documents, producing text, JSON, CSV, HTML, OTSL, summaries, Python code, and schema-based key-value data.

What is Granite Vision 4.1 4B?

Granite Vision 4.1 4B is an open-weight vision-language model developed by IBM Research. A vision-language model accepts both visual information and text instructions, then generates a text-based response. In this case, the model is specialized for understanding business documents and converting their visual content into structured, machine-readable information.

The model belongs to IBM’s Granite model collection and is positioned as a relatively small alternative to much larger general-purpose vision-language systems. Its approximately 4 billion parameters are divided between a 3.4-billion-parameter Granite language model and a 0.6-billion-parameter vision encoder and projector stack. This design prioritizes document extraction while keeping the model suitable for local or controlled deployment.

Granite Vision 4.1 4B was released on April 29, 2026. The official downloadable model identifier is ibm-granite/granite-vision-4.1-4b, and the model is released under the Apache 2.0 license.

What is the model designed to do?

The model’s main purpose is to turn visual documents into useful structured outputs. Instead of treating an uploaded page only as an image to describe, it can be instructed to recover the information contained in charts, tables, forms, and other layouts.

Supported workflows documented for the model include:

  • Extracting tables as JSON, HTML, or OTSL
  • Converting charts into CSV data
  • Summarizing the content of a chart
  • Generating Python code that recreates a chart
  • Extracting semantic key-value pairs according to a supplied schema
  • Describing document images in text
  • Producing summaries and other text-based representations

Semantic key-value extraction is particularly useful for forms, invoices, applications, and business records. A user can provide a schema containing fields such as invoice number, supplier, date, and total. The model then attempts to return the corresponding values as a JSON object, even when the fields appear in different positions on different documents.

Where Granite Vision 4.1 4B is strongest

Granite Vision 4.1 4B’s main advantage is specialization. Its documented tasks are closely aligned with common document-processing requirements, including recovering table structure, extracting chart data, and mapping document content to predefined fields.

The model card reports a 94.2% zero-shot exact-match score on the VAREX semantic key-value extraction benchmark. This is a provider-reported benchmark result for a specific evaluation setup, not a guarantee that every document will produce an exact result. Real-world accuracy can vary with image quality, layout complexity, handwriting, language, and the requested schema.

The model is also comparatively compact for a multimodal system. Its size may make local deployment more practical than deployment of a much larger vision-language model, particularly when the task does not require broad world knowledge or extended visual reasoning. IBM documents integrations and compatible runtimes including Transformers, vLLM, MLX VLM, Docker-based environments, and Docling.

Architecture and supported inputs

The vision stack uses a SigLIP2 vision encoder, windowed Q-Former projectors, and several visual feature injection points into the language model. Images are processed as tiled 384-by-384 patches, with visual features compressed before they are passed into the Granite language model.

The documented input format consists of English instructions together with PNG or JPEG images. The model primarily targets English-language document workflows. IBM notes that performance can decline on documents written in other languages, so multilingual extraction should be tested rather than assumed.

Granite Vision 4.1 4B supports image input and text input. Its output is text-based, although that text can represent JSON, CSV, HTML, OTSL, Python code, summaries, or ordinary descriptions. It does not directly generate images, audio, or video.

Context, output, reasoning, and tools

No verified context-window length is specified in the supplied IBM documentation, so a precise maximum input size cannot be stated. In practice, image dimensions, tiling behavior, runtime memory, prompt length, and the number of images processed will affect how much information can be handled reliably.

The recorded maximum output-token value is 4,096. This should be treated as the documented model limit in the supplied specification data, while the effective output may also depend on the runtime and generation settings.

This is not primarily a reasoning model for difficult multi-step questions. Its useful reasoning behavior is task-specific: it interprets a visual layout, identifies relevant content, and maps that content to a requested representation or schema. Editorial capability data rates its general reasoning as modest and its coding capability somewhat higher, reflecting its ability to produce chart-recreation Python code. Those ratings are editorial assessments, not IBM-published benchmark scores.

No separate provider-level tool or function-calling capability is documented for Granite Vision 4.1 4B. It can generate code and structured text, but that should not be confused with executing tools or taking actions. Applications that need validation, database writes, OCR post-processing, or workflow automation must provide those surrounding components themselves.

Deployment and pricing

Granite Vision 4.1 4B is downloadable open-weight software rather than a model with a specified IBM-hosted per-token price. The supplied research does not identify an official hosted API price, provider-managed batch service, or standard subscription price for this exact model.

Organizations can run the model locally or in controlled infrastructure using supported software such as Hugging Face Transformers, vLLM, MLX VLM, Docker-compatible runtimes, or Docling integrations. The Apache 2.0 license permits broad use subject to the license terms, but operating costs still include compute, storage, engineering, and maintenance.

The model’s compact size can improve the speed-and-cost trade-off compared with larger multimodal models, but it is not cost-free to run. The vision encoder, image tiling, model weights, and generation process all require memory and compute. Actual speed depends on hardware, quantization, batching, image resolution, and the selected inference runtime.

Important limitations

Granite Vision 4.1 4B should not be treated as a universal visual assistant. Its strongest evidence and intended use cases concern structured document extraction. IBM identifies limited generalization outside document-focused tasks, possible hallucinations, and weaker performance on non-English documents as limitations.

Extraction errors can include missing fields, incorrect values, damaged table relationships, or plausible but unsupported answers. This matters especially for invoices, financial records, compliance documents, and applications. A production workflow should validate required fields, check numerical totals, preserve the original image, and route uncertain results for review.

The model also has no verified context-length value in the supplied research. Large multi-page documents may need to be split into images or processed page by page. Users should test how their chosen runtime handles multiple images, long instructions, and unusually complex layouts.

When to choose Granite Vision 4.1 4B

Choose Granite Vision 4.1 4B when the central problem is extracting structured information from document images and you want an open-weight model that can be deployed under your own infrastructure. It is a good fit for:

  • Invoice and receipt field extraction
  • Table conversion for analytics or document migration
  • Chart-to-CSV or chart-to-code workflows
  • Form and application processing
  • Visual document search and multimodal retrieval-augmented generation
  • Document intelligence systems that require controlled or self-hosted deployment

A larger general-purpose vision-language model may be more appropriate for open-ended image questions, broad multilingual work, complex visual reasoning, or conversations that move well beyond document extraction. A conventional OCR and document-layout pipeline may be preferable when a workflow needs highly predictable text recognition, explicit confidence scores, or mature field-level validation. A hosted multimodal API may also be easier for teams that do not want to manage model files, GPUs, inference servers, and upgrades.

Overall assessment

Granite Vision 4.1 4B is best understood as a focused document-extraction model, not as a general consumer chatbot with vision. Its combination of structured outputs, chart and table workflows, schema-based extraction, open weights, and Apache 2.0 licensing makes it useful for organizations building controlled document-processing systems.

Its trade-off is equally clear: the model’s compact, specialized design does not remove the need for testing, validation, and suitable infrastructure. Buyers and developers should select it for document intelligence tasks where local control and extraction efficiency matter more than broad visual conversation, multilingual coverage, or unrestricted reasoning.


Answers to Frequently Asked Questions

What are the main limitations of Granite Vision 4.1 4B?
The model is specialized for document extraction and may perform less reliably on open-ended visual tasks, non-English documents, complex layouts, poor-quality images, and handwriting. It can hallucinate or return incorrect, missing, or damaged fields, so production workflows should validate outputs, check numerical values, preserve source images, and route uncertain results for human review.
How can Granite Vision 4.1 4B be deployed?
Granite Vision 4.1 4B can be deployed locally or in controlled infrastructure using software such as Hugging Face Transformers, vLLM, MLX VLM, Docker-compatible runtimes, and Docling integrations. It is downloadable open-weight software rather than a model with a documented IBM-hosted per-token price.
Can Granite Vision 4.1 4B extract invoice fields and tables?
Yes. The model can extract semantic key-value pairs from invoices, forms, and applications according to a supplied schema, including fields such as invoice number, supplier, date, and total. It can also convert tables into JSON, HTML, or OTSL and convert charts into CSV data.
What is Granite Vision 4.1 4B designed to do?
Granite Vision 4.1 4B is an open-weight vision-language model from IBM Research designed to extract structured information from business documents. It can process tables, charts, forms, invoices, applications, and other document images, producing outputs such as JSON, CSV, HTML, OTSL, Python code, and text summaries.
What is the official model identifier and license for Granite Vision 4.1 4B?
The official downloadable model identifier is ibm-granite/granite-vision-4.1-4b. The model is released under the Apache 2.0 license.


Sources 3
Provider

About IBM watsonx