Granite Embedding

Granite-Embedding-107M-Multilingual

by IBM watsonx · Retired from IBM watsonx.ai; open-weight checkpoint remains available through IBM's Hugging Face organization

IBM Granite-Embedding-107M-Multilingual is a compact Apache 2.0 embedding model for multilingual semantic search, cross-lingual retrieval, similarity matching, and RAG. It supports 12 languages, produces 384-dimensional vectors, and accepts up to 512 tokens. The open-weight checkpoint remains available through IBM’s Hugging Face organization, while its watsonx.ai deployment was withdrawn on November 12, 2025.

Embeddings Reasoning Coding
Granite-Embedding-107M-Multilingual converts text into numerical vectors that represent meaning rather than generating conversational answers. Those vectors can be stored in a vector database and compared to find semantically related queries, passages, or documents across 12 supported languages. The model is relatively small, produces 384-dimensional embeddings, and is intended for efficient local or third-party deployment. IBM’s hosted watsonx.ai deployment has been retired, so new users should treat the Hugging Face checkpoint as the relevant access path and plan their own inference environment.
Outputs

What Granite-Embedding-107M-Multilingual can produce

Embeddings
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family Granite Embedding
Model type Embedding
Context window 512 tokens
Release date 2024-12-18
Status Retired from IBM watsonx.ai; open-weight checkpoint remains available through IBM's Hugging Face organization
Deprecation date 2025-08-13
Shutdown date 2025-11-12
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for this embedding checkpoint. Embedding models do not expose a generative knowledge cutoff in the same way as chat or language-generation models.

Model notes

IBM's model card identifies this as a 107M-parameter dense biencoder producing 384-dimensional vectors. It supports English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese. The model uses an encoder-only XLM-RoBERTa-like architecture with CLS pooling and has a 512-token maximum sequence length. IBM lists the release date as December 18, 2024. IBM watsonx.ai lifecycle records list availability on January 6, 2025, deprecation on August 13, 2025, and withdrawal on November 12, 2025, with Granite-Embedding-278M-Multilingual-Embedding listed as the recommended replacement. Comparative scores are editorial estimates for an embedding model and do not measure generative reasoning or coding ability.

Cost

Model pricing

Input No official hosted API price verified; open-weight model intended for local or third-party deployment
Output Not applicable to token generation; produces 384-dimensional embeddings
Model guide

Granite-Embedding-107M-Multilingual: IBM’s Fast, Low-Cost Model for Multilingual Retrieval

Granite-Embedding-107M-Multilingual is a 107-million-parameter IBM Granite encoder model for multilingual semantic search, cross-lingual retrieval, similarity matching, recommendation, and retrieval-augmented generation. It produces 384-dimensional text embeddings, supports 12 languages, accepts up to 512 tokens, and is released under Apache 2.0. Its watsonx.ai deployment was deprecated in August 2025 and withdrawn on November 12, 2025, although the open-weight checkpoint remains available through IBM’s Hugging Face organization.

What is Granite-Embedding-107M-Multilingual?

Granite-Embedding-107M-Multilingual is an open-weight text embedding model from IBM’s Granite Embeddings family. Rather than writing text, it transforms a query, sentence, passage, or document into a fixed-length numerical representation called an embedding. Texts with similar meanings tend to produce vectors that are close together, allowing software to retrieve or rank relevant content.

The model is designed for multilingual information-retrieval workloads. A typical application might embed a user’s question, compare that vector with vectors for a document collection, and return the closest passages. The retrieved passages can then be supplied to a separate generative model in a retrieval-augmented generation (RAG) system. The embedding model itself does not answer the question or generate the final response.

IBM identifies the checkpoint as a dense biencoder with approximately 107 million parameters. It uses an encoder-only, XLM-RoBERTa-like transformer architecture with CLS pooling and produces 384-dimensional vectors. The model is released under the Apache 2.0 license.

Supported languages and input limit

The model was fine-tuned for English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese. IBM’s model card also states that users may fine-tune it for additional languages, but performance for an added language should be validated rather than assumed.

The maximum input length is 512 tokens. Tokens are pieces of text used by the model and do not correspond exactly to words. If a passage exceeds the limit, it is truncated. For long documents, an application should split the source into smaller chunks before embedding them. Chunking is therefore a core implementation requirement rather than an optional optimization.

What the model produces

SpecificationVerified detail
Model typeMultilingual text embedding encoder
Parameter countApproximately 107 million
Embedding size384 dimensions
Maximum input512 tokens
Primary languages12, including English, Chinese, Japanese, Arabic, and several European languages
LicenseApache 2.0
Text generationNot supported; this is not a generative language model

The output is a vector, not readable text. The model therefore has no maximum output-token setting in the usual language-model sense. Its useful output is the 384-value embedding representation.

Main use cases

  • Multilingual semantic search: find documents related to a query even when the wording differs from the source text.
  • Cross-lingual retrieval: search content in one supported language using a query written in another supported language.
  • RAG retrieval: select relevant document chunks before passing them to a separate text-generation model.
  • Similarity matching: compare support tickets, product descriptions, questions, or passages by meaning.
  • Recommendation and ranking: identify related items or rank candidates using vector similarity.
  • Clustering and classification: use embeddings as features in downstream machine-learning workflows.

For example, a company could embed a multilingual knowledge base once, store the vectors, and then embed each incoming question at search time. A vector-search system can use the resulting representations to identify relevant material without relying only on exact keyword matches.

Speed, cost, and performance positioning

At approximately 107 million parameters, this is a relatively compact embedding checkpoint. The supplied model research gives it an editorial speed score of 9 out of 10 and a cost score of 10 out of 10. Those scores are comparative editorial assessments, not IBM-published benchmark ratings. They reflect the model’s small size and suitability for efficient inference, not a guarantee of a particular latency or infrastructure cost.

IBM reported that the model was approximately twice as fast as other models with similar embedding dimensions in its comparisons. That is a provider-reported positioning claim, and actual performance depends on hardware, batching, sequence length, concurrency, quantization, and inference software. Users should benchmark their own workload, especially when comparing local CPU inference with GPU deployment or a managed embedding service.

The trade-off is capacity and scope. A compact model can reduce memory and processing requirements, but it may not match larger or newer embedding models on every language, domain, or retrieval benchmark. The 384-dimensional output also uses less storage than higher-dimensional alternatives, which can matter when a database contains millions of vectors.

Deployment and pricing

There is no verified current official hosted API price for this model in the supplied research. It was intended for local or third-party deployment as an open-weight model, so the practical cost depends on the infrastructure and serving method selected by the user. The Apache 2.0 license permits broad use subject to the license terms, but hosting, storage, database, and operational costs still apply.

The model can be loaded through commonly used open-model tooling such as Transformers or SentenceTransformers, according to IBM’s documentation and model materials. The supplied research does not establish a current IBM-managed inference endpoint for new deployments.

Its IBM watsonx.ai lifecycle history is important. IBM lists January 6, 2025 as the hosted availability date, August 13, 2025 as the deprecation date, and November 12, 2025 as the withdrawal date. As of the supplied research date, the watsonx.ai deployment is retired. The checkpoint remains identifiable through IBM’s Hugging Face organization, but hosted availability, maintenance, and infrastructure support should not be assumed from the former watsonx.ai listing.

Capabilities and limitations

Granite-Embedding-107M-Multilingual accepts text input and returns embedding output. It does not generate text, images, audio, or video. It does not provide conversational reasoning, code generation, web search, tool calling, function execution, streaming responses, or structured JSON generation. It can support a coding or RAG application as a retrieval component, but it is not itself a coding model or an agent.

The model also does not have a meaningful generative knowledge cutoff. Unlike a chat model, it is not intended to answer questions from a stored body of world knowledge. Its role is to represent the semantic content of supplied text.

The 512-token limit is the most important practical restriction. Long documents must be divided into chunks, and chunk size and overlap can affect retrieval quality. The model’s multilingual support is also not a guarantee of identical quality across all 12 languages. Domain-specific terminology, unusual scripts, translation direction, and document style can all influence results.

When to choose this model

Choose Granite-Embedding-107M-Multilingual when you need a compact, multilingual embedding model for semantic retrieval and want the flexibility of an open-weight Apache 2.0 checkpoint. It is particularly suitable when:

  • your search or RAG system covers several of its supported languages;
  • you need cross-lingual retrieval rather than English-only matching;
  • you want to run inference locally or through infrastructure you control;
  • storage efficiency matters because the model produces 384-dimensional vectors;
  • low inference cost and throughput are more important than using the newest or largest embedding model.

A newer or larger embedding model may be more appropriate when you need improved retrieval quality for a specialized domain, broader language coverage, longer inputs, or a currently supported managed API. IBM’s later Granite-Embedding-97M-Multilingual-R2 and Granite-Embedding-311M-Multilingual-R2 models are named successors in the supplied research, but their suitability should be evaluated against the exact languages, data, and deployment requirements of the project.

Bottom line

Granite-Embedding-107M-Multilingual is a focused retrieval component rather than a general-purpose AI assistant. Its main attractions are multilingual coverage, a small parameter footprint, 384-dimensional output, Apache 2.0 licensing, and support for efficient semantic-search and RAG pipelines. Its main disadvantages are the 512-token input limit, the absence of generative or tool-use capabilities, uncertain hosted availability after the watsonx.ai withdrawal, and the possibility that newer embedding checkpoints will provide better quality for some workloads.


Answers to Frequently Asked Questions

Is Granite-Embedding-107M-Multilingual available through a hosted API, and what does it cost?
The supplied research does not verify a current official hosted API price. The model is an open-weight Apache 2.0 checkpoint intended for local or third-party deployment, so costs depend on hosting, inference, storage, and vector-database infrastructure. Its former IBM watsonx.ai deployment was retired after the listed withdrawal date of November 12, 2025.
What are the input limit and embedding dimensions of Granite-Embedding-107M-Multilingual?
The model accepts a maximum of 512 tokens per input and produces a 384-dimensional embedding. Text longer than 512 tokens is truncated, so long documents should be split into smaller chunks before embedding.
What is Granite-Embedding-107M-Multilingual used for?
Granite-Embedding-107M-Multilingual converts text into 384-dimensional numerical vectors for multilingual semantic search, cross-lingual retrieval, RAG pipelines, similarity matching, recommendation, ranking, clustering, and classification. It is a retrieval component and does not generate answers or other text.
Which languages does Granite-Embedding-107M-Multilingual support?
The model was fine-tuned for English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese. IBM states that it can be fine-tuned for additional languages, but performance should be validated for each added language.


Sources 4
Provider

About IBM watsonx