What is Granite Embedding 311M Multilingual R2?
Granite-Embedding-311M-Multilingual-R2 is an open-weight embedding model from IBM's Granite Embedding family. An embedding model converts text into a list of numbers, commonly called a vector, that captures aspects of the text's meaning. A search system can compare the vector for a user's query with vectors created from documents and return passages that are semantically related, even when the wording is different.
The model is designed for queries, passages, documents, and supported programming code. It can support semantic search, retrieval-augmented generation (RAG), cross-lingual retrieval, similarity matching, clustering, recommendation, classification, duplicate-content detection, and code search. It is not a chatbot or answer-generation model: a separate retrieval system, vector database, reranker, or generative model is needed to build a complete question-answering application.
IBM identifies this model as part of the current Granite Embedding R2 collection. The supplied model information describes it as a 311-million-parameter model based on the ModernBERT architecture and released under the Apache 2.0 license. Those characteristics allow organizations to deploy the model using their own infrastructure or compatible third-party hosting rather than depending on a standard hosted IBM per-token endpoint for this exact model.
Key specifications at a glance
| Specification | Details |
|---|---|
| Provider | IBM |
| Model family | Granite Embedding |
| Model type | Multilingual text and code embedding model |
| Parameters | Approximately 311 million |
| Architecture | ModernBERT-based bi-encoder |
| Output | 768-dimensional embeddings |
| Maximum sequence length | 32,768 tokens |
| Supported languages | More than 200, with enhanced retrieval support for 52 languages |
| License | Apache 2.0 |
| Reduced dimensions | 512, 384, 256, or 128 through Matryoshka truncation |
The sequence length is an input limit, not an output limit. The model produces vectors rather than generated text, so a maximum output-token allowance is not applicable. Its primary output is a fixed-size embedding representation.
Multilingual and code retrieval capabilities
Granite-Embedding-311M-Multilingual-R2 supports more than 200 languages through its multilingual pretraining data. IBM reports enhanced retrieval training for 52 languages. This distinction matters: broad language support does not mean that retrieval quality will be identical in every language. Low-resource languages, specialized terminology, unusual document formats, and domain-specific writing should be tested with representative data before deployment.
The model is also intended for cross-lingual retrieval. For example, a user could submit a query in one supported language while the indexed document is written in another. This can help organizations search multilingual knowledge bases without maintaining a completely separate retrieval pipeline for every language pair.
For software teams, the model supports code retrieval for Python, Go, Java, JavaScript, PHP, Ruby, SQL, C, and C++. This can be used to find similar functions, locate examples in a codebase, retrieve documentation for an implementation pattern, or support a coding assistant's retrieval layer. The embedding model itself does not write or edit code and does not provide native code-generation behavior.
32K input length and flexible vector dimensions
The model accepts sequences of up to 32,768 tokens. This is substantially longer than the 512-token context associated with IBM's previous-generation Granite multilingual embedding model, according to the supplied research. The larger limit can reduce the need to split long documents into many small pieces and can help with multi-passage or long-document retrieval.
Long input support does not automatically make whole-document indexing the best design. Very large inputs can increase inference time and memory use, and a single vector for an entire document may be less precise than vectors created for meaningful sections. In practice, teams should compare document-level, passage-level, and hierarchical chunking strategies against their own search queries.
Granite-Embedding-311M-Multilingual-R2 supports Matryoshka dimension reduction. It produces a full 768-dimensional vector, but applications can truncate that representation to 512, 384, 256, or 128 dimensions. Smaller vectors require less storage and reduce the cost of similarity comparisons in a vector index. The trade-off is that reduced dimensions can lower retrieval quality, so the appropriate size should be selected through evaluation rather than assumed in advance.
Architecture and deployment options
The model uses a bi-encoder design. Queries and documents are encoded independently, which allows documents to be embedded once and indexed before users search. At query time, the system embeds the new query and compares it with stored vectors using cosine similarity or another distance function.
This design is efficient for large-scale retrieval, but it is not the same as a cross-encoder reranker that reads a query and document together. A production RAG system may therefore use this model for fast candidate retrieval and add a separate reranking stage when greater precision is needed.
IBM provides the model through its Granite documentation and official model resources. The supplied research also identifies compatibility or deployment information for ONNX, OpenVINO, vLLM, and related inference workflows. Flash Attention 2 is optional and may improve efficiency on compatible hardware. Exact throughput depends on hardware, batch size, sequence length, precision, runtime configuration, and vector-index design; no universal speed figure is established by the supplied information.
Performance, speed, and cost trade-offs
IBM positions the model for multilingual retrieval, English retrieval, code retrieval, long-document search, conversational multi-turn retrieval, and reasoning-as-retrieval evaluations. These are provider-reported positioning and evaluation areas, not a guarantee of a particular result for every dataset.
At approximately 311 million parameters, this model favors retrieval quality and coverage over the smallest possible memory footprint. The supplied comparison states that Granite-Embedding-311M-Multilingual-R2 generally prioritizes quality over throughput and resource efficiency when compared with the smaller Granite-Embedding-97M-Multilingual-R2. The 97M sibling may be more appropriate when latency, memory, or edge deployment is the primary constraint. The 311M model is a better candidate when multilingual coverage, long inputs, code retrieval, or accuracy is more important than minimizing infrastructure requirements.
There is no verified standard IBM per-token price for this exact open-weight model. Pricing therefore depends on how it is deployed. Self-hosted users incur infrastructure, storage, and operational costs; users of third-party hosting may pay according to that provider's compute or endpoint pricing. Vector-database storage and retrieval infrastructure are additional costs. The absence of a model-specific hosted price should not be interpreted as zero cost.
Supported inputs and outputs
The primary input is text, including natural-language queries, passages, documents, and supported source code. The model does not have verified image, audio, or video input support in the supplied specifications. Its direct output is an embedding vector rather than text, an image, audio, video, a structured answer, or a tool action.
It has no native conversational generation, web search, function calling, tool use, streaming response, or JSON-mode capability documented in the supplied model record. A surrounding application can add those features by combining the embeddings with a search engine, tools, an agent framework, or a separate generative model, but those capabilities would belong to the application rather than to Granite-Embedding-311M-Multilingual-R2 itself.
Best use cases
- Multilingual enterprise search: Index internal policies, manuals, tickets, and knowledge-base content for semantic retrieval across languages.
- Retrieval-augmented generation: Retrieve relevant passages before passing them to a separate language model that generates an answer.
- Cross-lingual discovery: Match queries and documents written in different supported languages.
- Long-document search: Encode large sections or carefully selected chunks from contracts, technical manuals, reports, and research material.
- Code search: Find related code and documentation across the supported programming languages.
- Similarity and deduplication: Detect related, near-duplicate, or semantically similar content.
- Recommendation and clustering: Group documents or recommend content based on vector similarity.
Limitations to consider
The most important limitation is that this is an embedding model, not a generative language model. It cannot independently explain search results, answer questions, summarize retrieved passages, or carry on a conversation. Those functions require additional components.
Language coverage is broad but uneven retrieval quality remains possible. The enhanced training coverage for 52 languages should not be treated as a guarantee for all more than 200 supported languages. Benchmarking should include the languages, terminology, query styles, and document types used by the intended application.
Long context also involves a quality and infrastructure trade-off. Feeding an entire 32K-token document may be less effective than creating focused chunks, and longer sequences can increase memory consumption and latency. Similarly, reducing vectors to 128 or 256 dimensions can lower storage costs but may affect ranking quality.
Finally, open weights provide deployment flexibility but transfer more responsibility to the operator. Teams must choose an inference runtime, provision hardware, manage model updates, secure stored documents and vectors, monitor retrieval quality, and account for infrastructure expenses.
When to choose Granite Embedding 311M Multilingual R2
Choose this model when the application needs multilingual semantic retrieval, cross-lingual search, code retrieval, or long-input support and can accommodate a larger encoder than a compact embedding model. Its Apache 2.0 licensing and open-weight deployment options are useful for organizations that need more control over hosting and data movement than a closed hosted embedding API provides.
Choose a smaller embedding model, such as IBM's Granite-Embedding-97M-Multilingual-R2, when low latency, low memory use, or edge deployment matters more than the potential retrieval-quality advantage of the 311M model. Choose a separate reranker when the initial vector search needs more precise query-document matching, and choose a generative model when the system must produce natural-language answers. Granite-Embedding-311M-Multilingual-R2 is best understood as the retrieval foundation in such a system, not as the complete system itself.

