What IBM Granite Embedding 97M Multilingual R2 is
IBM Granite-Embedding-97M-Multilingual-R2 is an open-weight dense bi-encoder for turning text and code into embeddings. An embedding is a fixed-length numerical representation of meaning. When an application embeds a search query and a collection of documents, it can compare the resulting vectors—commonly with cosine similarity—to find passages that are conceptually related even when they do not use the same words.
This makes the model a retrieval component rather than a chatbot. It does not produce natural-language answers, images, audio, or video. A typical retrieval-augmented generation (RAG) system could use Granite Embedding 97M R2 to locate relevant passages and then send those passages to a separate generative model that writes the final response.
The model is provided by IBM as part of the Granite Embedding family. Its compact size is intended to reduce inference cost and latency compared with larger embedding models, while retaining multilingual and code-retrieval capabilities.
Key specifications at a glance
| Specification | Detail |
|---|---|
| Provider | IBM |
| Model type | Open-weight multilingual embedding model |
| Parameters | Approximately 97 million |
| Embedding size | 384 dimensions |
| Maximum input length | 32,768 tokens |
| Language coverage | More than 200 languages through the underlying encoder, with enhanced retrieval support for 52 languages |
| Code support | Enhanced retrieval training for Python, Go, Java, JavaScript, PHP, Ruby, SQL, C, and C++ |
| License | Apache 2.0 |
| Output | 384-dimensional embeddings |
| Hosted pricing | No provider-hosted per-token price published for the downloadable model |
What it is designed to do
The model's primary purpose is semantic retrieval: finding text or code based on meaning rather than exact keyword matches. The 384-dimensional output is small enough to make vector indexes relatively economical, which is useful when an application must store and search a large document collection.
Supported use cases include semantic site search, enterprise document retrieval, question-answering pipelines, recommendation and similarity matching, multilingual knowledge bases, and RAG. It can also be used for long-document retrieval because its maximum sequence length is 32,768 tokens, substantially longer than the 512-token limit of the previous Granite multilingual embedding generation described in the supplied research.
Code retrieval is another specific focus. A developer could embed a natural-language request such as “find code that validates a JSON Web Token” and compare it with embeddings of source files, functions, or documentation. The enhanced code training covers several common programming languages, although the model remains a retrieval model rather than a code-generation system.
Languages, inputs, and context limit
The underlying multilingual encoder supports more than 200 languages. IBM reports enhanced retrieval training for 52 named languages, covering major European, Asian, Middle Eastern, and African languages. Examples include English, Spanish, French, German, Arabic, Chinese, Japanese, Korean, Hindi, Russian, Turkish, Ukrainian, and Vietnamese.
The distinction between broad support and enhanced support matters. A model may accept text in many languages, but retrieval quality is not necessarily uniform across them. Languages in the enhanced set are the better-supported choice for production retrieval, while lower-resource languages may depend more heavily on cross-lingual transfer and should be evaluated with representative queries and documents.
Each input can contain up to 32,768 tokens. Longer inputs are truncated, so applications handling large documents should still consider sensible chunking. Embedding an entire long document can make retrieval less precise if the document contains several unrelated topics. Splitting content into meaningful sections can help an index return more targeted passages, while the long context limit provides flexibility for unusually long passages or documents.
Architecture and deployment options
Granite-Embedding-97M-Multilingual-R2 uses a ModernBERT-based encoder with 12 transformer layers, 12 attention heads, a 180,000-token vocabulary, SiLU activations, rotary position embeddings, and alternating attention lengths. IBM created the compact model through layer pruning and vocabulary selection from the larger Granite-Embedding-311M-Multilingual-R2 model, followed by knowledge distillation and contrastive fine-tuning.
These choices position the 97M model for deployments where throughput, memory use, and operational simplicity matter. The weights are released under the Apache 2.0 license, making the model suitable for research and commercial use subject to that license.
IBM documents deployment paths using Sentence Transformers and Hugging Face Transformers. ONNX and OpenVINO variants are available for CPUs, GPUs, and other compatible hardware through ONNX Runtime. The model can also be served with vLLM using an embedding task. The supplied research notes that Ollama does not currently support this model's ModernBERT architecture, so Ollama should not be treated as a verified deployment route.
Performance and speed-versus-accuracy trade-offs
IBM reports a score of 60.3 on the Multilingual MTEB Retrieval benchmark and approximately 2,534 documents per second in its published benchmark configuration. These are provider-reported results, and actual throughput will vary with hardware, batch size, sequence lengths, software versions, and indexing or preprocessing overhead.
The smaller parameter count is the model's main practical trade-off. Compared with a larger embedding model such as Granite-Embedding-311M-Multilingual-R2, the 97M variant is intended to use fewer resources and deliver higher throughput, but it is not the choice for maximum retrieval accuracy in every language or domain. The best option depends on the cost of missed results, available hardware, and the size and language mix of the corpus.
The 384-dimensional output also reduces vector-storage and similarity-search costs compared with models that produce much larger vectors. However, smaller vectors and a smaller encoder can involve an accuracy trade-off. Teams should test the model on their own search queries, document types, languages, and relevance judgments rather than treating the published benchmark as a guarantee.
Capabilities and limitations
- Embedding output: The model returns 384-dimensional vectors for text and code.
- Multilingual retrieval: It supports more than 200 languages through its encoder, with enhanced retrieval training for 52 languages.
- Long inputs: It accepts up to 32,768 tokens; inputs beyond that limit are truncated.
- Code retrieval: It has explicit retrieval training for Python, Go, Java, JavaScript, PHP, Ruby, SQL, C, and C++.
- Generative reasoning: It does not provide conversational reasoning or natural-language answer generation. A separate generative model is required for that role.
- Tool use: It has no documented built-in function calling or tool-use capability. An application can use its vectors inside a larger retrieval or agent system, but orchestration must be implemented separately.
- Modalities: The supplied specifications identify text and code input with embedding output. There is no verified image, audio, or video input or output capability.
- Fine-tuning and hosted features: The supplied research does not verify a provider-hosted fine-tuning service, streaming mode, batch API, caching feature, or managed per-token endpoint for this downloadable model.
Pricing and operating costs
IBM does not publish a provider-hosted per-token price for Granite-Embedding-97M-Multilingual-R2 in the supplied research. It is a downloadable open-weight model, so there is no verified standard monthly or per-request price to report for the model itself.
Using it still has infrastructure costs. A self-hosted deployment may require CPU or GPU capacity, storage, vector-database resources, monitoring, and engineering time. ONNX, OpenVINO, and the compact 97M parameter count can help organizations target more economical hardware, but the actual cost depends on throughput requirements, batch sizes, sequence lengths, availability targets, and the selected inference platform. If the model is accessed through a third-party service, that service's pricing applies rather than an IBM-published model price.
When to choose this model
Choose Granite-Embedding-97M-Multilingual-R2 when a project needs a relatively small multilingual embedding model that can be deployed under an Apache 2.0 license. It is particularly suitable when the application needs:
- Low-latency semantic search on cost-sensitive infrastructure.
- Multilingual retrieval across languages represented in its enhanced-support set.
- RAG retrieval for enterprise documents, manuals, policies, or knowledge bases.
- Longer input handling than conventional 512-token embedding models.
- Cross-lingual search, where a query and relevant document may use different languages.
- Similarity matching or recommendation based on text meaning.
- Retrieval across source code and natural-language programming documentation.
- Self-hosted, CPU, GPU, ONNX, or OpenVINO deployment rather than dependence on a paid hosted endpoint.
It is less appropriate when the application needs the highest possible retrieval accuracy and can afford a larger model, when the task requires generated answers, or when the input is primarily images, audio, or video. It is also not a replacement for a reranker, generative language model, speech model, or multimodal encoder. A larger sibling model may be worth evaluating when benchmark results on the target languages show that the 97M model misses too many relevant documents.
How to evaluate it in a real application
Start with representative queries and documents rather than relying only on general multilingual benchmarks. Include short keyword-like searches, natural-language questions, long passages, code queries, spelling variation, and queries written in each important language. Measure retrieval recall at practical values such as the top 5, top 10, or top 20 results, and inspect whether the returned passages contain the information needed by the downstream application.
Also measure operational behavior: embedding throughput, memory consumption, index size, query latency, and performance at the intended batch size. For RAG, evaluate the complete pipeline because a slightly faster embedder may not improve the overall system if chunking, vector search, reranking, or answer generation dominates latency.
Overall, Granite-Embedding-97M-Multilingual-R2 is best understood as a compact retrieval engine. Its combination of 384-dimensional vectors, long input support, multilingual coverage, code-retrieval training, and open licensing makes it a practical option for teams prioritizing deployment efficiency. Its clear boundary is equally important: it finds and represents information, but it does not reason through a conversation or write the final answer.

