Granite 3.3

Granite-3.3-2B-Instruct

by IBM watsonx · Available; open-weight model with local deployment and dedicated IBM watsonx deployment options

IBM Granite-3.3-2B-Instruct is a 2-billion-parameter Apache 2.0 language model with a 131,072-token context window. It targets reasoning, instruction following, coding, RAG, summarization, multilingual dialogue, function calling, and local or dedicated deployment.

Text Reasoning Coding
IBM Granite-3.3-2B-Instruct is a compact open-weight language model released on April 16, 2025. It improves on earlier Granite 2B instruct models with stronger reasoning, mathematics, coding, instruction-following, long-context, and fill-in-the-middle capabilities while remaining suitable for local or dedicated deployment.
Outputs

What Granite-3.3-2B-Instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Fine-tuning
Model profile

Performance characteristics

6/10 Reasoning
6/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite 3.3
Model type General Purpose
Context window 131K tokens
Maximum output 16K tokens
Release date 2025-04-16
Status Available; open-weight model with local deployment and dedicated IBM watsonx deployment options
Knowledge cutoff notes

The reviewed IBM and IBM-maintained model-card materials do not publish an exact knowledge-cutoff date for this checkpoint. The model's 2025 release date must not be used as a proxy for its training-data cutoff.

Model notes

Canonical Hugging Face identifier is ibm-granite/granite-3.3-2b-instruct. The model has 2 billion parameters, uses a dense architecture, and is licensed under Apache 2.0. IBM documents structured reasoning with <think> and <response> tags, plus fill-in-the-middle support for code completion. The model card identifies function-calling tasks as a capability, but the reviewed sources do not establish a separate legacy JSON-mode feature or constrained JSON-schema output for this exact open-weight checkpoint. IBM's foundation-model documentation lists a 131,072-token combined context limit and a 16,384-token maximum generated output for the 2B model. No exact knowledge-cutoff date or public per-token price was verified.

Model guide

Granite-3.3-2B-Instruct: IBM’s Compact Open-Weight Model for Long-Context Work

Granite-3.3-2B-Instruct is IBM's Apache 2.0-licensed, 2-billion-parameter instruction-tuned language model with a 131,072-token context window. It is designed for instruction following, reasoning, coding, multilingual dialogue, retrieval-augmented generation, summarization, text extraction, question answering, and function-calling workflows.

What is Granite-3.3-2B-Instruct?

Granite-3.3-2B-Instruct is a 2-billion-parameter dense language model from IBM's Granite family. It is the instruction-tuned version of Granite-3.3-2B-Base, designed to follow natural-language requests rather than simply continue raw training text. In practical terms, it can answer questions, summarize documents, extract information, generate and revise code, classify text, and support conversational or retrieval-augmented applications.

The model was released on April 16, 2025, and is distributed under the Apache 2.0 license. That licensing model makes it suitable for many research, commercial, and self-hosted scenarios, subject to the license and the applicable model-card conditions. Its open-weight availability is important: users can download the checkpoint and operate it with compatible infrastructure instead of relying exclusively on a provider-managed, per-token endpoint.

IBM positions Granite-3.3-2B-Instruct for enterprise-oriented language workflows that benefit from relatively low resource requirements, deployment control, and a long context window. IBM also lists the model as available for dedicated deployment through watsonx. This is different from saying that the model has a single public hosted API price: no exact public per-token price was verified for this checkpoint.

Where it fits in IBM's Granite lineup

Granite-3.3-2B-Instruct is one of IBM's smaller general-purpose Granite language models. Its 2-billion-parameter size places it below larger models in raw capacity, but also makes it more practical for local inference, specialized servers, and applications where speed or operating cost matters more than frontier-level performance.

The model is best understood as a text-focused open-weight component rather than a complete multimodal assistant. It can be integrated into IBM watsonx or operated with external tools such as Transformers, vLLM, SGLang, and compatible quantized runtimes. These deployment options allow teams to choose between managed enterprise infrastructure and greater control over their own hardware and data path.

Core capabilities and supported tasks

Granite-3.3-2B-Instruct is intended for text input and text output. The reviewed specifications do not document image, audio, or video input, and it does not natively generate images, audio, or video. Its main capabilities include:

  • Instruction following: answering direct requests, transforming text, extracting fields, and following multi-step prompts.
  • Reasoning and mathematics: working through reasoning-oriented questions and mathematical problems, with IBM reporting improvements over earlier Granite 3.1 and 3.2 2B instruction-tuned models.
  • Summarization: condensing long documents, meetings, and other text-heavy sources.
  • Question answering and RAG: responding to questions over supplied documents or retrieved passages. Retrieval-augmented generation, or RAG, gives the model external context at request time rather than requiring all information to be stored in the model's parameters.
  • Text extraction and classification: turning unstructured text into useful fields or assigning categories.
  • Coding: generating, explaining, repairing, and refactoring code, including fill-in-the-middle completion.
  • Function calling: supporting workflows in which the model identifies and requests an external function or tool. The model card identifies function-calling tasks as a capability, but the supplied research does not verify a separate constrained JSON-schema mode for this exact checkpoint.
  • Multilingual dialogue: supporting English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese according to IBM's documented language list.

Granite 3.3 also separates intermediate reasoning from the final answer with <think> and <response> tags. This can be useful in workflows that need to distinguish a model's internal-style reasoning section from the answer presented to an end user, although applications should define their own parsing and safety rules rather than assuming every generated response will follow a desired format perfectly.

Context window and maximum output

The verified context limit is 131,072 tokens. A token is a unit of text used by the model; it may represent a whole short word, part of a longer word, punctuation, or another fragment. A 131K-token context allows the model to process substantially longer documents or conversation histories than a small-context model, making it relevant to document question answering, meeting analysis, codebase-oriented prompts, and long RAG inputs.

IBM's watsonx foundation-model documentation lists a maximum generated output of 16,384 tokens for the 2B model. The context and output limits are not a guarantee that every request will produce a useful answer at the maximum size. Long prompts can increase memory requirements and latency, while long outputs can be repetitive or less reliable. Applications should reserve space for the requested answer and avoid including irrelevant material merely because the context window is large.

Performance, speed, and cost trade-offs

IBM's model card reports the following evaluation results for Granite-3.3-2B-Instruct: 43.45 on AlpacaEval-2.0, 28.86 on Arena-Hard, 72.48 on HumanEval, 58.09 on MATH-500, and 44.33 on DROP. These are provider-reported benchmark results under the documented evaluation setup. They are useful for understanding the model's tested profile, but they are not universal production guarantees; results can change with prompts, sampling settings, hardware, task distribution, and evaluation methodology.

The model's relatively small parameter count is its main practical advantage. Compared with larger language models, a 2B checkpoint generally requires less memory and can be faster and less expensive to operate when deployed on suitable local or dedicated hardware. The supplied research gives an editorial speed score of 8 out of 10 and a cost score of 9 out of 10, but those scores are evaluations rather than IBM-published specifications.

The trade-off is capacity. A compact model may be less dependable than larger contemporary models on difficult multi-step reasoning, complex software engineering, broad factual questions, nuanced multilingual tasks, and autonomous agent workflows. The long context window expands the amount of text the model can receive; it does not ensure perfect retrieval, attention, or reasoning across every token.

Pricing and deployment options

Granite-3.3-2B-Instruct has no verified public per-input-token or per-output-token price in the supplied research. It is an open-weight model available through IBM's Granite Hugging Face organization, so users can download it and account for their own hardware, hosting, storage, and engineering costs. IBM also documents dedicated deployment through watsonx, where pricing may depend on the selected IBM service, deployment configuration, region, and usage terms rather than on a single universally advertised model rate.

For local use, compatible options include Hugging Face Transformers, vLLM, SGLang, and other runtimes that support the Granite architecture. Quantization can reduce memory requirements, although the appropriate method and resulting quality depend on the runtime and workload. A self-hosted deployment may provide stronger control over data handling and predictable infrastructure ownership, but it transfers responsibilities such as capacity planning, monitoring, upgrades, access control, and reliability to the deploying team.

Best use cases

This model is a good candidate when an application needs a text model that is relatively economical to run and can be deployed with more control than a closed hosted service. Suitable examples include:

  • Internal enterprise assistants grounded in company documents.
  • RAG pipelines for manuals, policies, technical documentation, or meeting records.
  • Long-document summarization and question answering.
  • Text extraction and classification at scale.
  • Code explanation, completion, repair, and refactoring support.
  • Multilingual text workflows covering the languages documented by IBM.
  • Dedicated or local deployments where data handling, network boundaries, or operating-cost control matter.
  • Function-calling prototypes and workflow automation that provide external tools around the model.

Limitations and when to choose another model

Granite-3.3-2B-Instruct may not be the right choice for every workload. Its small size can limit difficult reasoning, complex coding, specialist knowledge, and high-stakes decision support. It should be evaluated with representative prompts and checked by appropriate safeguards before being used in regulated or consequential workflows. The model's exact knowledge cutoff was not published in the reviewed IBM materials, so its 2025 release date should not be treated as a training-data cutoff.

Choose a larger model when answer quality on complex reasoning, demanding code generation, broad knowledge, or agentic planning is more important than local speed and infrastructure efficiency. Choose a managed hosted model when a team does not want to operate inference infrastructure or needs a provider-managed availability and billing layer. Choose a multimodal model when the application must interpret images, audio, or video, because Granite-3.3-2B-Instruct is documented as a text-only input and output model.

Conversely, Granite-3.3-2B-Instruct is compelling when a smaller open-weight model can meet the quality requirement. It offers a combination of Apache 2.0 licensing, a 131K-token context, coding and reasoning support, multilingual text handling, and local or dedicated deployment that can be difficult to obtain from a simple consumer chatbot.

Bottom line

Granite-3.3-2B-Instruct is a practical compact language model for organizations and developers who value deployment control, long-context text processing, and lower resource requirements. Its strengths are most relevant to RAG, summarization, extraction, coding assistance, multilingual dialogue, and enterprise text workflows. Its open-weight status does not eliminate hosting costs, and its small size means that larger models can remain preferable for the hardest reasoning and generation tasks. The best fit is a measured, task-specific deployment in which speed, cost, and control are balanced against the quality ceiling of a 2-billion-parameter model.


Answers to Frequently Asked Questions

What are the main use cases and limitations of Granite-3.3-2B-Instruct?
The model is well suited to enterprise assistants, RAG pipelines, long-document summarization, text extraction and classification, coding support, multilingual workflows, and function-calling prototypes. Its 2-billion-parameter size can limit performance on highly complex reasoning, demanding software engineering, specialist knowledge, and autonomous agent tasks. It is also a text-only model and does not natively process images, audio, or video.
What languages does Granite-3.3-2B-Instruct support?
IBM documents support for English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese.
How can Granite-3.3-2B-Instruct be deployed?
The model can be downloaded as an open-weight checkpoint and run with compatible tools such as Hugging Face Transformers, vLLM, SGLang, and quantized runtimes. IBM also lists dedicated deployment through watsonx. Self-hosting provides more control over data and infrastructure but requires the deploying team to manage hardware, monitoring, security, and reliability.
What is Granite-3.3-2B-Instruct?
Granite-3.3-2B-Instruct is IBM’s 2-billion-parameter, instruction-tuned language model for tasks such as question answering, summarization, information extraction, coding, classification, multilingual dialogue, and retrieval-augmented generation. It is an open-weight, text-only model released under the Apache 2.0 license.
What is the context window of Granite-3.3-2B-Instruct?
Granite-3.3-2B-Instruct supports a context window of up to 131,072 tokens. IBM’s watsonx documentation lists a maximum generated output of 16,384 tokens for the 2B model.


Sources 5
Provider

About IBM watsonx