Granite 3.1

Granite-3.1-8B-Instruct

by IBM watsonx · Available as an open-weight model; superseded by Granite-3.3-8B-Instruct for newer deployments

IBM Granite-3.1-8B-Instruct is an Apache 2.0 open-weight, approximately 8.1B-parameter text model with a 128K-token context window. It supports multilingual generation, long-document analysis, RAG, extraction, coding tasks, and function calling. It has no native image, audio, or video capabilities, and self-hosting costs depend on infrastructure rather than a universal per-token price.

Text Reasoning Coding
Granite-3.1-8B-Instruct is an open-weight language model released by IBM on December 18, 2024. It is designed for practical text workloads rather than image, audio, or video processing, with a particular emphasis on long documents, enterprise assistants, retrieval-augmented generation, multilingual dialogue, and function-calling workflows. The model can be downloaded and deployed under the Apache 2.0 license, although IBM’s newer Granite-3.3-8B-Instruct is the more current choice for many new deployments.
Outputs

What Granite-3.1-8B-Instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming Fine-tuning
Model profile

Performance characteristics

5/10 Reasoning
6/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite 3.1
Model type General Purpose
Context window 131K tokens
Release date 2024-12-18
Status Available as an open-weight model; superseded by Granite-3.3-8B-Instruct for newer deployments
Knowledge cutoff notes

IBM's public model card and Granite 3.1 documentation identify the release date and training details but do not provide a direct, authoritative knowledge-cutoff date for this exact model.

Model notes

Canonical Hugging Face identifier: ibm-granite/granite-3.1-8b-instruct. The model is an approximately 8.1B-parameter dense decoder-only transformer released under Apache 2.0. IBM documents 128K context, multilingual support, function-calling improvements, and fine-tuning use cases. The model is open-weight, so infrastructure-dependent self-hosting costs replace a universal model API price. IBM's Granite 3.1 repository was archived in 2026, and the official model page points users to Granite-3.3-8B-Instruct as a newer version.

Model guide

Granite-3.1-8B-Instruct: IBM’s Open-Weight Model for Long-Context Enterprise Text

Granite-3.1-8B-Instruct is IBM’s Apache 2.0-licensed, approximately 8.1-billion-parameter instruction-tuned language model. Its 128K-token context window, multilingual support, function calling, and open-weight distribution make it suitable for self-hosted enterprise assistants, retrieval-augmented generation, document question answering, summarization, extraction, and code-related text tasks.

What is Granite-3.1-8B-Instruct?

Granite-3.1-8B-Instruct is IBM’s instruction-tuned version of an approximately 8.1-billion-parameter dense decoder-only transformer. In practical terms, it is a text-generation model that has been adapted to follow user instructions, answer questions, summarize material, extract information, generate code-related content, and participate in tool or function-calling workflows.

IBM released the model on December 18, 2024, under the permissive Apache 2.0 license. Its canonical Hugging Face identifier is ibm-granite/granite-3.1-8b-instruct. Because the weights are available for download, organizations can run the model on infrastructure they control rather than depending exclusively on a provider-hosted, per-token endpoint.

The model is part of IBM’s Granite 3.1 family and was fine-tuned from Granite-3.1-8B-Base. The Granite 3.1 repository has since been archived, and IBM’s model materials point to Granite-3.3-8B-Instruct as a newer successor. Granite-3.1-8B-Instruct remains relevant when its Apache 2.0 licensing, reproducible weights, existing deployment tooling, or established 128K-context behavior is more important than using the newest Granite generation.

Context window and architecture

The verified context length is 131,072 tokens, commonly described as 128K tokens. A context window is the amount of text the model can consider in one request, including the prompt and supplied documents. This capacity is useful for long reports, meeting transcripts, policy collections, technical documentation, and retrieval workflows that need to place substantial source material in context.

The model uses 40 transformer layers, 32 attention heads, and 8 key-value heads. Its documented architecture includes grouped-query attention, rotary position embeddings, SwiGLU activation, RMSNorm, and shared input/output embeddings. These implementation details matter primarily to people selecting hardware or inference software; ordinary users will mainly notice the model’s long input capacity and text-focused behavior.

The research identifies a 128K maximum context length but does not provide a verified maximum output-token limit. Output length therefore depends on the serving framework, available memory, and the remaining space within the model’s context window. Deployers should check the configuration of their chosen runtime rather than assume a universal output limit.

What the model can do

Granite-3.1-8B-Instruct is intended for instruction following and enterprise-oriented language tasks. Its documented uses include:

  • Conversational assistants and question answering
  • Summarization of long documents or transcripts
  • Document extraction and classification
  • Retrieval-augmented generation, or RAG
  • Multilingual dialogue and text transformation
  • Code-related generation and analysis
  • Function calling and tool-oriented workflows

In a RAG system, a separate retrieval component finds relevant passages and supplies them to the model. Granite-3.1-8B-Instruct can then answer questions using that supplied context. This makes it a candidate for internal knowledge assistants, document question-answering systems, compliance research, and support tools where the organization wants to keep documents and inference inside its own environment.

IBM documents support for 12 languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese. Performance is not necessarily equal across these languages; IBM indicates that English has the strongest instruction-tuning coverage. Multilingual support should therefore be validated against the specific language, terminology, and formatting requirements of a deployment.

Tool use, coding, and reasoning profile

The model supports function-calling workflows. Function calling allows an application to describe available operations, such as searching a database or creating a ticket, and lets the model produce a structured request for the application to execute. The model does not independently perform those external operations; the surrounding software must validate and run them.

Granite-3.1-8B-Instruct can be used for code-related tasks, including code generation, transformation, explanation, and extraction of structured information from technical material. The supplied research does not establish a particular programming-language ranking or benchmark result, so coding quality should be tested with representative repositories and prompts.

It is not presented as a dedicated frontier reasoning model. Its approximately 8.1B parameter size is a practical fit for applications that value deployment efficiency and controllable infrastructure over maximum reasoning performance. The database’s reasoning and coding scores are editorial evaluation fields, not IBM-published specifications, and should not be treated as benchmark results.

Supported modalities and important limitations

Granite-3.1-8B-Instruct is a text model. It accepts text input and produces text output. The supplied model information records no native image, audio, or video input, and no image, audio, video, music, or other non-text output. It therefore cannot directly inspect an image, listen to a recording, or generate media without separate systems that convert those modalities into text or handle the media output.

Its other limitations follow from its age and deployment model. Newer Granite generations may perform better on some current tasks, but the supplied research does not provide a direct benchmark comparison. The model can also produce inaccurate, biased, or unsafe responses, so production applications should use evaluation, input and output controls, and appropriate human or automated review.

Open weights provide control but also transfer responsibility to the deployer. Hardware selection, quantization, inference serving, scaling, monitoring, security, upgrades, and reliability are not automatically handled by IBM when the model is self-hosted. Function-calling output also requires application-side validation before any external action is taken.

Deployment and fine-tuning

The model can be loaded with the Hugging Face Transformers library and served with compatible inference systems such as vLLM or SGLang. These options allow a team to choose between a straightforward local experiment and a more scalable serving architecture. Actual throughput and memory requirements depend on hardware, precision, quantization, batching, sequence length, and the number of simultaneous requests; the supplied research does not specify a universal hardware requirement.

IBM describes fine-tuning as an intended use. Fine-tuning can adapt the model to specialized instructions, enterprise terminology, additional languages, or a narrow workflow. It does not remove the need for evaluation: a tuned model can still produce incorrect or unsafe answers, and tuning data should be selected and governed carefully.

The Apache 2.0 license is a significant deployment advantage for organizations that need to inspect, reproduce, or commercially use the model subject to the license terms. Teams should still review the model card, any applicable third-party components, and their own legal and security requirements before distribution or production use.

Pricing and cost

There is no universal official per-token price for Granite-3.1-8B-Instruct in the supplied research. The model is open-weight, so self-hosted cost is determined by compute infrastructure, storage, quantization, serving software, traffic, and operational requirements rather than by a single IBM model endpoint price.

This can make the model attractive for predictable internal workloads or deployments that require data and inference control. It does not automatically make every deployment inexpensive. A low-volume experiment may run on modest infrastructure, while a high-throughput service may require multiple accelerators, redundancy, monitoring, and engineering support. Hosted services may offer easier scaling, but their prices and availability are provider- and platform-specific and are not established here.

When to choose Granite-3.1-8B-Instruct

Choose Granite-3.1-8B-Instruct when you need an open-weight text model with a large context window and a permissive license. It is a sensible candidate for:

  • Private enterprise assistants that process internal documents
  • Long-context summarization and document question answering
  • RAG systems where the retrieval and inference stack is self-managed
  • Structured extraction, classification, and multilingual text workflows
  • Function-calling applications that need an instruction-tuned model
  • Teams that prefer deployment control over a fully managed API

Its 8B scale can offer a practical balance between capability, latency, and infrastructure cost compared with much larger models, although the supplied research does not provide universal speed or cost benchmarks. It may be preferable to a larger hosted model when data residency, reproducibility, license flexibility, or predictable self-hosting matters more than maximum reasoning quality.

Another option may be more appropriate when the application needs native image, audio, or video understanding; media generation; a fully managed provider endpoint with guaranteed scaling; or frontier-level reasoning. For new IBM Granite deployments, Granite-3.3-8B-Instruct is the documented newer successor and should be evaluated alongside this model. Granite-3.1-8B-Instruct is most compelling when its specific weights, Apache 2.0 licensing, long-context support, or compatibility with an existing deployment is the deciding factor.


Answers to Frequently Asked Questions

How can Granite-3.1-8B-Instruct be deployed?
Granite-3.1-8B-Instruct can be loaded with the Hugging Face Transformers library and served using compatible systems such as vLLM or SGLang. Deployment requirements vary according to hardware, precision, quantization, batching, sequence length, traffic, and concurrency. IBM’s newer documented successor is Granite-3.3-8B-Instruct, but Granite-3.1-8B-Instruct may remain useful when its Apache 2.0 license, available weights, 128K context, or existing deployment compatibility are important.
Can Granite-3.1-8B-Instruct be used for function calling and RAG?
Yes. The model supports function-calling workflows and retrieval-augmented generation. In function calling, the model produces a structured request that surrounding software must validate and execute. In a RAG system, a separate retrieval component supplies relevant passages, which the model uses to answer questions or generate responses.
Under what license is Granite-3.1-8B-Instruct released?
IBM released Granite-3.1-8B-Instruct under the permissive Apache 2.0 license on December 18, 2024. Its canonical Hugging Face identifier is ibm-granite/granite-3.1-8b-instruct. Organizations can download and self-host the weights subject to the license terms and their own legal, security, and compliance requirements.
What is Granite-3.1-8B-Instruct?
Granite-3.1-8B-Instruct is IBM’s instruction-tuned, approximately 8.1-billion-parameter text-generation model for tasks such as question answering, summarization, information extraction, code-related work, multilingual dialogue, retrieval-augmented generation, and function-calling workflows.
What is the context window of Granite-3.1-8B-Instruct?
Granite-3.1-8B-Instruct has a verified maximum context length of 131,072 tokens, commonly described as 128K tokens. This supports long documents, reports, transcripts, technical documentation, and retrieval-augmented generation workflows. The maximum output length depends on the serving framework, available memory, and remaining context capacity.


Sources 5
Provider

About IBM watsonx