Granite Embedding

Granite-Embedding-311M-Multilingual-R2

by IBM watsonx · Current open-weight model

IBM Granite-Embedding-311M-Multilingual-R2 is an open-weight multilingual embedding model for semantic search, RAG, cross-lingual retrieval, long-document search, similarity, and code retrieval. It supports more than 200 languages, produces 768-dimensional vectors, accepts up to 32,768 tokens, and supports Matryoshka reduction to smaller vector sizes. It is not a generative or conversational model, and no standard IBM per-token price was verified.

Embeddings Reasoning Coding
Granite-Embedding-311M-Multilingual-R2 is IBM's 311-million-parameter, ModernBERT-based bi-encoder for multilingual text and code retrieval. Unlike a conversational language model, it does not generate answers. Instead, it represents queries and documents as numerical vectors so that semantically related content can be found efficiently. Its long input limit, multilingual coverage, open Apache 2.0 license, and flexible deployment options make it suited to enterprise search, multilingual RAG, long-document retrieval, and code search.
Outputs

What Granite-Embedding-311M-Multilingual-R2 can produce

Embeddings
Inputs

What it can understand

Text
Model profile

Performance characteristics

1/10 Reasoning
3/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite Embedding
Model type Other
Context window 33K tokens
Release date 2026-04-29
Status Current open-weight model
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was published. As an embedding model, it is intended for encoding supplied input text rather than answering from a stated knowledge base.

Model notes

Open-weight Apache 2.0 embedding model developed by IBM's Granite Embedding team. Uses a ModernBERT-based bi-encoder and produces 768-dimensional vectors. Supports Matryoshka truncation to 512, 384, 256, or 128 dimensions. Supports more than 200 languages, with enhanced retrieval support for 52 languages and code retrieval for Python, Go, Java, JavaScript, PHP, Ruby, SQL, C, and C++. IBM documents a 32,768-token maximum sequence length. The model is intended for self-hosted or third-party deployment; no standard IBM per-token price for this exact open-weight model was verified. Editorial scores reflect its role as an embedding model rather than a generative language model.

Model guide

IBM Granite Embedding 311M Multilingual R2 for Long-Context Cross-Lingual Retrieval

IBM Granite-Embedding-311M-Multilingual-R2 is an open-weight multilingual embedding model for converting text, documents, queries, and supported programming code into vectors used for semantic search, retrieval-augmented generation, similarity, clustering, recommendation, and cross-lingual retrieval. It produces 768-dimensional embeddings, supports sequences up to 32,768 tokens, covers more than 200 languages, and offers Matryoshka dimension reduction to 512, 384, 256, or 128 dimensions.

What is Granite Embedding 311M Multilingual R2?

Granite-Embedding-311M-Multilingual-R2 is an open-weight embedding model from IBM's Granite Embedding family. An embedding model converts text into a list of numbers, commonly called a vector, that captures aspects of the text's meaning. A search system can compare the vector for a user's query with vectors created from documents and return passages that are semantically related, even when the wording is different.

The model is designed for queries, passages, documents, and supported programming code. It can support semantic search, retrieval-augmented generation (RAG), cross-lingual retrieval, similarity matching, clustering, recommendation, classification, duplicate-content detection, and code search. It is not a chatbot or answer-generation model: a separate retrieval system, vector database, reranker, or generative model is needed to build a complete question-answering application.

IBM identifies this model as part of the current Granite Embedding R2 collection. The supplied model information describes it as a 311-million-parameter model based on the ModernBERT architecture and released under the Apache 2.0 license. Those characteristics allow organizations to deploy the model using their own infrastructure or compatible third-party hosting rather than depending on a standard hosted IBM per-token endpoint for this exact model.

Key specifications at a glance

SpecificationDetails
ProviderIBM
Model familyGranite Embedding
Model typeMultilingual text and code embedding model
ParametersApproximately 311 million
ArchitectureModernBERT-based bi-encoder
Output768-dimensional embeddings
Maximum sequence length32,768 tokens
Supported languagesMore than 200, with enhanced retrieval support for 52 languages
LicenseApache 2.0
Reduced dimensions512, 384, 256, or 128 through Matryoshka truncation

The sequence length is an input limit, not an output limit. The model produces vectors rather than generated text, so a maximum output-token allowance is not applicable. Its primary output is a fixed-size embedding representation.

Multilingual and code retrieval capabilities

Granite-Embedding-311M-Multilingual-R2 supports more than 200 languages through its multilingual pretraining data. IBM reports enhanced retrieval training for 52 languages. This distinction matters: broad language support does not mean that retrieval quality will be identical in every language. Low-resource languages, specialized terminology, unusual document formats, and domain-specific writing should be tested with representative data before deployment.

The model is also intended for cross-lingual retrieval. For example, a user could submit a query in one supported language while the indexed document is written in another. This can help organizations search multilingual knowledge bases without maintaining a completely separate retrieval pipeline for every language pair.

For software teams, the model supports code retrieval for Python, Go, Java, JavaScript, PHP, Ruby, SQL, C, and C++. This can be used to find similar functions, locate examples in a codebase, retrieve documentation for an implementation pattern, or support a coding assistant's retrieval layer. The embedding model itself does not write or edit code and does not provide native code-generation behavior.

32K input length and flexible vector dimensions

The model accepts sequences of up to 32,768 tokens. This is substantially longer than the 512-token context associated with IBM's previous-generation Granite multilingual embedding model, according to the supplied research. The larger limit can reduce the need to split long documents into many small pieces and can help with multi-passage or long-document retrieval.

Long input support does not automatically make whole-document indexing the best design. Very large inputs can increase inference time and memory use, and a single vector for an entire document may be less precise than vectors created for meaningful sections. In practice, teams should compare document-level, passage-level, and hierarchical chunking strategies against their own search queries.

Granite-Embedding-311M-Multilingual-R2 supports Matryoshka dimension reduction. It produces a full 768-dimensional vector, but applications can truncate that representation to 512, 384, 256, or 128 dimensions. Smaller vectors require less storage and reduce the cost of similarity comparisons in a vector index. The trade-off is that reduced dimensions can lower retrieval quality, so the appropriate size should be selected through evaluation rather than assumed in advance.

Architecture and deployment options

The model uses a bi-encoder design. Queries and documents are encoded independently, which allows documents to be embedded once and indexed before users search. At query time, the system embeds the new query and compares it with stored vectors using cosine similarity or another distance function.

This design is efficient for large-scale retrieval, but it is not the same as a cross-encoder reranker that reads a query and document together. A production RAG system may therefore use this model for fast candidate retrieval and add a separate reranking stage when greater precision is needed.

IBM provides the model through its Granite documentation and official model resources. The supplied research also identifies compatibility or deployment information for ONNX, OpenVINO, vLLM, and related inference workflows. Flash Attention 2 is optional and may improve efficiency on compatible hardware. Exact throughput depends on hardware, batch size, sequence length, precision, runtime configuration, and vector-index design; no universal speed figure is established by the supplied information.

Performance, speed, and cost trade-offs

IBM positions the model for multilingual retrieval, English retrieval, code retrieval, long-document search, conversational multi-turn retrieval, and reasoning-as-retrieval evaluations. These are provider-reported positioning and evaluation areas, not a guarantee of a particular result for every dataset.

At approximately 311 million parameters, this model favors retrieval quality and coverage over the smallest possible memory footprint. The supplied comparison states that Granite-Embedding-311M-Multilingual-R2 generally prioritizes quality over throughput and resource efficiency when compared with the smaller Granite-Embedding-97M-Multilingual-R2. The 97M sibling may be more appropriate when latency, memory, or edge deployment is the primary constraint. The 311M model is a better candidate when multilingual coverage, long inputs, code retrieval, or accuracy is more important than minimizing infrastructure requirements.

There is no verified standard IBM per-token price for this exact open-weight model. Pricing therefore depends on how it is deployed. Self-hosted users incur infrastructure, storage, and operational costs; users of third-party hosting may pay according to that provider's compute or endpoint pricing. Vector-database storage and retrieval infrastructure are additional costs. The absence of a model-specific hosted price should not be interpreted as zero cost.

Supported inputs and outputs

The primary input is text, including natural-language queries, passages, documents, and supported source code. The model does not have verified image, audio, or video input support in the supplied specifications. Its direct output is an embedding vector rather than text, an image, audio, video, a structured answer, or a tool action.

It has no native conversational generation, web search, function calling, tool use, streaming response, or JSON-mode capability documented in the supplied model record. A surrounding application can add those features by combining the embeddings with a search engine, tools, an agent framework, or a separate generative model, but those capabilities would belong to the application rather than to Granite-Embedding-311M-Multilingual-R2 itself.

Best use cases

  • Multilingual enterprise search: Index internal policies, manuals, tickets, and knowledge-base content for semantic retrieval across languages.
  • Retrieval-augmented generation: Retrieve relevant passages before passing them to a separate language model that generates an answer.
  • Cross-lingual discovery: Match queries and documents written in different supported languages.
  • Long-document search: Encode large sections or carefully selected chunks from contracts, technical manuals, reports, and research material.
  • Code search: Find related code and documentation across the supported programming languages.
  • Similarity and deduplication: Detect related, near-duplicate, or semantically similar content.
  • Recommendation and clustering: Group documents or recommend content based on vector similarity.

Limitations to consider

The most important limitation is that this is an embedding model, not a generative language model. It cannot independently explain search results, answer questions, summarize retrieved passages, or carry on a conversation. Those functions require additional components.

Language coverage is broad but uneven retrieval quality remains possible. The enhanced training coverage for 52 languages should not be treated as a guarantee for all more than 200 supported languages. Benchmarking should include the languages, terminology, query styles, and document types used by the intended application.

Long context also involves a quality and infrastructure trade-off. Feeding an entire 32K-token document may be less effective than creating focused chunks, and longer sequences can increase memory consumption and latency. Similarly, reducing vectors to 128 or 256 dimensions can lower storage costs but may affect ranking quality.

Finally, open weights provide deployment flexibility but transfer more responsibility to the operator. Teams must choose an inference runtime, provision hardware, manage model updates, secure stored documents and vectors, monitor retrieval quality, and account for infrastructure expenses.

When to choose Granite Embedding 311M Multilingual R2

Choose this model when the application needs multilingual semantic retrieval, cross-lingual search, code retrieval, or long-input support and can accommodate a larger encoder than a compact embedding model. Its Apache 2.0 licensing and open-weight deployment options are useful for organizations that need more control over hosting and data movement than a closed hosted embedding API provides.

Choose a smaller embedding model, such as IBM's Granite-Embedding-97M-Multilingual-R2, when low latency, low memory use, or edge deployment matters more than the potential retrieval-quality advantage of the 311M model. Choose a separate reranker when the initial vector search needs more precise query-document matching, and choose a generative model when the system must produce natural-language answers. Granite-Embedding-311M-Multilingual-R2 is best understood as the retrieval foundation in such a system, not as the complete system itself.


Answers to Frequently Asked Questions

Is Granite Embedding 311M Multilingual R2 a chatbot or answer-generation model?
No. It is a bi-encoder embedding model that outputs vectors rather than text, so it cannot independently answer questions, summarize documents, or conduct conversations. A complete application needs additional components such as a vector database or search engine, optional reranker, and a separate generative model.
What are the context length and embedding dimensions of Granite Embedding 311M Multilingual R2?
The model accepts sequences of up to 32,768 tokens and produces 768-dimensional embeddings. Through Matryoshka truncation, applications can use reduced vector sizes of 512, 384, 256, or 128 dimensions to lower storage and similarity-search costs, with a possible trade-off in retrieval quality.
How many languages and programming languages does Granite Embedding 311M Multilingual R2 support?
The model supports more than 200 languages, with enhanced retrieval training reported for 52 languages. It also supports code retrieval for Python, Go, Java, JavaScript, PHP, Ruby, SQL, C, and C++.
What is IBM Granite Embedding 311M Multilingual R2 used for?
IBM Granite Embedding 311M Multilingual R2 is an open-weight embedding model for semantic search, multilingual and cross-lingual retrieval, RAG, code search, similarity matching, clustering, recommendation, classification, and duplicate-content detection. It converts queries, documents, passages, or supported source code into vectors that can be compared in a retrieval system.


Sources 4
Provider

About IBM watsonx