What IBM slate-30m-english-rtrvr-v2 does
IBM slate-30m-english-rtrvr-v2 is an embedding model available through IBM watsonx.ai. Instead of writing an answer, it turns an input such as a search query, document passage, or support question into a numerical representation called a vector. Texts with related meanings should produce vectors that are close together when compared with a similarity measure such as cosine similarity.
This makes the model a retrieval component rather than a conversational model. A typical system might embed a collection of documents, store the resulting vectors in a vector index, embed a user's question, and retrieve the passages whose vectors are most similar. A separate generative model can then use those passages to produce an answer in a retrieval-augmented generation (RAG) workflow.
The canonical watsonx.ai model identifier is ibm/slate-30m-english-rtrvr-v2. IBM is the provider, and the model belongs to the Slate family of retrieval models.
Technical specifications and input limits
| Specification | Verified value |
|---|---|
| Model type | English retrieval embedding model |
| Architecture | Bi-encoder |
| Approximate parameter count | 30 million |
| Embedding size | 384 dimensions |
| Maximum input length | 512 tokens |
| Primary language | English |
| Output | Vector embeddings |
The 512-token limit applies to each encoded input. In practice, documents longer than that usually need to be divided into passages before indexing. Chunk size and overlap affect retrieval quality, so a deployment should test its document structure rather than assuming that one fixed chunking strategy will work for every collection.
The model has no generated-text output limit because it does not generate text. Its output is an embedding vector, and IBM does not separately price output tokens for this model.
Architecture and training background
IBM describes the underlying slate.30m.english model as a small RoBERTa-style transformer with six layers and approximately 30 million parameters. The v2 retrieval model was distilled from IBM's larger slate.125m.english.rtrvr-v2 model and adapted for retrieval tasks. Distillation transfers useful behavior from a larger model into a smaller one, generally creating a model that is cheaper and faster to run but may offer less capacity than its larger source.
According to IBM's model-card description, training included retrieval-oriented pretraining, knowledge distillation, contrastive learning, mined query-passage pairs, and supervised retrieval data. The listed sources include Natural Questions, SQuAD, Stack Exchange, SPECTER, S2ORC, SearchQA, HotpotQA, FEVER, MIRACL, IBM Docs, and IBM Software Support data. These details are provider-reported training information, not a guarantee that the model will perform equally well on every organization's terminology or document domain.
Strengths, speed, and reported performance
The main practical strength of slate-30m-english-rtrvr-v2 is its small size. A 30-million-parameter encoder can be a sensible choice when an application needs to embed many passages, keep memory use under control, or prioritize economical inference over the highest possible retrieval quality. Its 384-dimensional vectors are also smaller than vectors from many larger embedding models, which can reduce vector-index storage and comparison costs.
IBM's model card reports a BEIR-15 nDCG@10 score of 49.06 for the June 30, 2024 version and a LongNQ nDCG@10 score of 62.07. It also reports approximately 0.20 seconds per query for reranking under the stated A100 40 GB GPU test setup. These are provider-reported results from particular evaluation and hardware conditions; they should not be treated as universal latency or quality guarantees.
For this model, the relevant trade-off is retrieval efficiency versus breadth and capacity. A larger or multilingual embedding model may perform better on difficult, diverse, or non-English data, while slate-30m-english-rtrvr-v2 can be attractive when English-only retrieval and a compact footprint are sufficient.
Supported modalities and capabilities
The model accepts text and returns numerical embeddings. It does not accept image, audio, or video inputs, and it does not directly produce text, images, audio, video, or speech. It is therefore not a multimodal generation model and should not be evaluated using chatbot, image-generation, or voice-assistant expectations.
- Text input: Supported for English queries, passages, and documents.
- Embedding output: Supported as 384-dimensional vectors.
- Text generation: Not supported.
- Reasoning: No standalone reasoning capability is provided; the model scores semantic relatedness for retrieval.
- Code generation: Not supported. It may embed code-related text, but it is not a code-generation model.
- Tools and functions: No native tool calling or function execution is identified for this embedding model.
- Streaming and structured output: Not applicable to the documented embedding output.
These limitations are not defects if the model is used as one stage in a larger search or RAG system. They do mean that a separate model is required to formulate answers, call tools, or perform application actions.
Pricing and lifecycle status
The supplied watsonx.ai pricing information lists an input price of USD 0.0001 per 1,000 input tokens. Because the model returns embeddings rather than generated tokens, there is no separately priced output-token category in the supplied data. Actual billing can still depend on IBM account, region, platform terms, and current watsonx.ai pricing conditions.
IBM announced the model on August 15, 2024 and later marked it deprecated on May 8, 2026. The lifecycle schedule is region-specific. IBM scheduled withdrawal in Dallas for September 8, 2026. Frankfurt, London, Tokyo, Sydney, Toronto, AWS Mumbai, and AWS GovCloud were scheduled for withdrawal on January 12, 2027. A project created after a regional withdrawal date should not assume that the model remains available there.
For that reason, slate-30m-english-rtrvr-v2 is better viewed as a legacy or migration case than as a recommended starting point for a new production system. IBM lists granite-embedding-278m-multilingual and multilingual-e5-large as recommended alternatives. They are migration targets, not aliases that automatically preserve identical vectors or ranking behavior. Re-indexing documents and retesting retrieval quality may be required when changing models.
Best use cases
- English semantic search: Find passages by meaning rather than exact keyword overlap.
- Dense vector retrieval: Build a compact vector index for an English document collection.
- RAG retrieval: Select relevant passages before sending them to a separate answer-generating model.
- Question matching: Detect similar support requests, duplicate questions, or related FAQ entries.
- Document indexing: Encode enterprise documents, support content, or technical passages for later retrieval.
- Resource-conscious deployments: Use a smaller encoder when memory, storage, or inference cost is more important than maximum retrieval quality.
For long documents, split the source into meaningful passages, embed each passage, and retain enough metadata to reconstruct the source location. For a RAG application, retrieved passages should still be checked for relevance and freshness before they are shown to a user or passed to a generative model.
When to choose this model—and when not to
Choose slate-30m-english-rtrvr-v2 only when its compact English retrieval profile fits the deployment and the relevant watsonx.ai region still supports it. It can make sense for an existing IBM workload that already uses the model, for controlled migration testing, or for a narrow English search task where its reported efficiency and low token price are valuable.
A different option is more appropriate when the application needs multilingual retrieval, longer inputs, higher retrieval quality on difficult domains, or a supported model for a new production deployment. IBM's listed alternatives, granite-embedding-278m-multilingual and multilingual-e5-large, should be evaluated when broader language coverage or a current migration path matters. A generative language model is required when the application must write answers, summarize retrieved passages, generate code, or interact with tools.
The most important operational limitation is its deprecated status. Even if its quality and cost are suitable, a new system built around it may require an avoidable migration soon. Teams should compare candidate replacement models on their own queries, measure recall and ranking quality, account for the different vector dimensions, and rebuild the index rather than assuming that existing embeddings are interchangeable.

