Slate

slate-30m-english-rtrvr-v2

by IBM watsonx · Deprecated; withdrawn in Dallas on 2026-09-08 and scheduled for withdrawal in other listed regions on 2027-01-12.

A compact IBM bi-encoder for English semantic retrieval. It produces 384-dimensional embeddings from queries and passages, accepts up to 512 tokens, and suits vector search, document matching, and RAG retrieval. Its small footprint and low listed input price are balanced by English-only coverage and deprecated, region-specific availability.

Embeddings
IBM slate-30m-english-rtrvr-v2 converts English queries, passages, and documents into 384-dimensional vector embeddings. Those vectors can be compared to find semantically similar content, making the model useful for search indexes, knowledge bases, duplicate-question detection, and retrieval-augmented generation pipelines. Its small 30-million-parameter architecture favors efficient inference, but the model is limited to English retrieval, a 512-token input length, and embedding output. It has also been deprecated by IBM, with availability ending on different dates depending on the watsonx.ai region.
Outputs

What slate-30m-english-rtrvr-v2 can produce

Embeddings
Inputs

What it can understand

Text
Model profile

Performance characteristics

8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Slate
Model type Other
Context window 512 tokens
Release date 2024-08-15
Status Deprecated; withdrawn in Dallas on 2026-09-08 and scheduled for withdrawal in other listed regions on 2027-01-12.
Deprecation date 2026-05-08
Shutdown date 2026-09-08 in Dallas; 2027-01-12 in Frankfurt, London, Tokyo, Sydney, Toronto, AWS Mumbai, and AWS GovCloud
Knowledge cutoff notes

IBM does not publish a conventional knowledge cutoff for this retrieval embedding model. Its training data and evaluation data are described in the model card, but no single authoritative cutoff date is specified.

Model notes

This is an embedding-only model, not a generative language model. It is a bi-encoder retrieval model that produces 384-dimensional vectors from queries, passages, and documents. The underlying model has approximately 30 million parameters and uses a small RoBERTa-style six-layer architecture. IBM describes it as a distilled version of slate-125m-english-rtrvr-v2. The model is English-focused and accepts up to 512 tokens. IBM deprecated it on 2026-05-08. Withdrawal dates vary by region: Dallas was scheduled for 2026-09-08, while Frankfurt, London, Tokyo, Sydney, Toronto, AWS Mumbai, and AWS GovCloud were scheduled for 2027-01-12. IBM recommends granite-embedding-278m-multilingual or multilingual-e5-large as alternatives.

Cost

Model pricing

Input USD 0.0001 per 1,000 input tokens on watsonx.ai
Output Not separately priced; the model returns embedding vectors rather than generated output tokens
Model guide

IBM slate-30m-english-rtrvr-v2: Compact English Embeddings for Semantic Retrieval

IBM slate-30m-english-rtrvr-v2 is a compact, English-only bi-encoder embedding model for semantic search, dense retrieval, document matching, and retrieval-augmented generation. It produces 384-dimensional vectors, accepts up to 512 tokens, and is designed to offer economical, low-latency retrieval rather than text generation. IBM deprecated the model in 2026, with region-specific withdrawal dates, so new deployments should consider IBM's recommended alternatives.

What IBM slate-30m-english-rtrvr-v2 does

IBM slate-30m-english-rtrvr-v2 is an embedding model available through IBM watsonx.ai. Instead of writing an answer, it turns an input such as a search query, document passage, or support question into a numerical representation called a vector. Texts with related meanings should produce vectors that are close together when compared with a similarity measure such as cosine similarity.

This makes the model a retrieval component rather than a conversational model. A typical system might embed a collection of documents, store the resulting vectors in a vector index, embed a user's question, and retrieve the passages whose vectors are most similar. A separate generative model can then use those passages to produce an answer in a retrieval-augmented generation (RAG) workflow.

The canonical watsonx.ai model identifier is ibm/slate-30m-english-rtrvr-v2. IBM is the provider, and the model belongs to the Slate family of retrieval models.

Technical specifications and input limits

SpecificationVerified value
Model typeEnglish retrieval embedding model
ArchitectureBi-encoder
Approximate parameter count30 million
Embedding size384 dimensions
Maximum input length512 tokens
Primary languageEnglish
OutputVector embeddings

The 512-token limit applies to each encoded input. In practice, documents longer than that usually need to be divided into passages before indexing. Chunk size and overlap affect retrieval quality, so a deployment should test its document structure rather than assuming that one fixed chunking strategy will work for every collection.

The model has no generated-text output limit because it does not generate text. Its output is an embedding vector, and IBM does not separately price output tokens for this model.

Architecture and training background

IBM describes the underlying slate.30m.english model as a small RoBERTa-style transformer with six layers and approximately 30 million parameters. The v2 retrieval model was distilled from IBM's larger slate.125m.english.rtrvr-v2 model and adapted for retrieval tasks. Distillation transfers useful behavior from a larger model into a smaller one, generally creating a model that is cheaper and faster to run but may offer less capacity than its larger source.

According to IBM's model-card description, training included retrieval-oriented pretraining, knowledge distillation, contrastive learning, mined query-passage pairs, and supervised retrieval data. The listed sources include Natural Questions, SQuAD, Stack Exchange, SPECTER, S2ORC, SearchQA, HotpotQA, FEVER, MIRACL, IBM Docs, and IBM Software Support data. These details are provider-reported training information, not a guarantee that the model will perform equally well on every organization's terminology or document domain.

Strengths, speed, and reported performance

The main practical strength of slate-30m-english-rtrvr-v2 is its small size. A 30-million-parameter encoder can be a sensible choice when an application needs to embed many passages, keep memory use under control, or prioritize economical inference over the highest possible retrieval quality. Its 384-dimensional vectors are also smaller than vectors from many larger embedding models, which can reduce vector-index storage and comparison costs.

IBM's model card reports a BEIR-15 nDCG@10 score of 49.06 for the June 30, 2024 version and a LongNQ nDCG@10 score of 62.07. It also reports approximately 0.20 seconds per query for reranking under the stated A100 40 GB GPU test setup. These are provider-reported results from particular evaluation and hardware conditions; they should not be treated as universal latency or quality guarantees.

For this model, the relevant trade-off is retrieval efficiency versus breadth and capacity. A larger or multilingual embedding model may perform better on difficult, diverse, or non-English data, while slate-30m-english-rtrvr-v2 can be attractive when English-only retrieval and a compact footprint are sufficient.

Supported modalities and capabilities

The model accepts text and returns numerical embeddings. It does not accept image, audio, or video inputs, and it does not directly produce text, images, audio, video, or speech. It is therefore not a multimodal generation model and should not be evaluated using chatbot, image-generation, or voice-assistant expectations.

  • Text input: Supported for English queries, passages, and documents.
  • Embedding output: Supported as 384-dimensional vectors.
  • Text generation: Not supported.
  • Reasoning: No standalone reasoning capability is provided; the model scores semantic relatedness for retrieval.
  • Code generation: Not supported. It may embed code-related text, but it is not a code-generation model.
  • Tools and functions: No native tool calling or function execution is identified for this embedding model.
  • Streaming and structured output: Not applicable to the documented embedding output.

These limitations are not defects if the model is used as one stage in a larger search or RAG system. They do mean that a separate model is required to formulate answers, call tools, or perform application actions.

Pricing and lifecycle status

The supplied watsonx.ai pricing information lists an input price of USD 0.0001 per 1,000 input tokens. Because the model returns embeddings rather than generated tokens, there is no separately priced output-token category in the supplied data. Actual billing can still depend on IBM account, region, platform terms, and current watsonx.ai pricing conditions.

IBM announced the model on August 15, 2024 and later marked it deprecated on May 8, 2026. The lifecycle schedule is region-specific. IBM scheduled withdrawal in Dallas for September 8, 2026. Frankfurt, London, Tokyo, Sydney, Toronto, AWS Mumbai, and AWS GovCloud were scheduled for withdrawal on January 12, 2027. A project created after a regional withdrawal date should not assume that the model remains available there.

For that reason, slate-30m-english-rtrvr-v2 is better viewed as a legacy or migration case than as a recommended starting point for a new production system. IBM lists granite-embedding-278m-multilingual and multilingual-e5-large as recommended alternatives. They are migration targets, not aliases that automatically preserve identical vectors or ranking behavior. Re-indexing documents and retesting retrieval quality may be required when changing models.

Best use cases

  • English semantic search: Find passages by meaning rather than exact keyword overlap.
  • Dense vector retrieval: Build a compact vector index for an English document collection.
  • RAG retrieval: Select relevant passages before sending them to a separate answer-generating model.
  • Question matching: Detect similar support requests, duplicate questions, or related FAQ entries.
  • Document indexing: Encode enterprise documents, support content, or technical passages for later retrieval.
  • Resource-conscious deployments: Use a smaller encoder when memory, storage, or inference cost is more important than maximum retrieval quality.

For long documents, split the source into meaningful passages, embed each passage, and retain enough metadata to reconstruct the source location. For a RAG application, retrieved passages should still be checked for relevance and freshness before they are shown to a user or passed to a generative model.

When to choose this model—and when not to

Choose slate-30m-english-rtrvr-v2 only when its compact English retrieval profile fits the deployment and the relevant watsonx.ai region still supports it. It can make sense for an existing IBM workload that already uses the model, for controlled migration testing, or for a narrow English search task where its reported efficiency and low token price are valuable.

A different option is more appropriate when the application needs multilingual retrieval, longer inputs, higher retrieval quality on difficult domains, or a supported model for a new production deployment. IBM's listed alternatives, granite-embedding-278m-multilingual and multilingual-e5-large, should be evaluated when broader language coverage or a current migration path matters. A generative language model is required when the application must write answers, summarize retrieved passages, generate code, or interact with tools.

The most important operational limitation is its deprecated status. Even if its quality and cost are suitable, a new system built around it may require an avoidable migration soon. Teams should compare candidate replacement models on their own queries, measure recall and ranking quality, account for the different vector dimensions, and rebuild the index rather than assuming that existing embeddings are interchangeable.


Answers to Frequently Asked Questions

Is IBM slate-30m-english-rtrvr-v2 still available for new production deployments?
IBM marked the model as deprecated on May 8, 2026, with region-specific withdrawal dates beginning September 8, 2026, in Dallas and January 12, 2027, in several other regions. It is generally better treated as a legacy or migration case. IBM lists granite-embedding-278m-multilingual and multilingual-e5-large as recommended alternatives, but switching requires re-indexing and retrieval-quality testing because their embeddings are not interchangeable.
How much does IBM slate-30m-english-rtrvr-v2 cost?
The supplied watsonx.ai pricing lists an input price of USD 0.0001 per 1,000 input tokens. There is no separately priced output-token category because the model returns embeddings rather than generated text. Actual billing may vary by account, region, platform terms, and current IBM pricing.
Can IBM slate-30m-english-rtrvr-v2 generate text or handle images and audio?
No. The model only accepts text and returns numerical embedding vectors. It does not generate text, process images, audio, or video, perform tool calls, or provide standalone reasoning. A separate generative model is needed to write answers or perform application actions.
What is IBM slate-30m-english-rtrvr-v2 used for?
IBM slate-30m-english-rtrvr-v2 is an English text embedding model for semantic search, dense vector retrieval, document indexing, question matching, and retrieval-augmented generation (RAG). It converts queries and passages into vectors that can be compared for semantic similarity.
What are the main technical specifications of ibm/slate-30m-english-rtrvr-v2?
The model is a bi-encoder with approximately 30 million parameters. It produces 384-dimensional embeddings, supports English text input, and has a maximum input length of 512 tokens per encoded input.


Sources 5
Provider

About IBM watsonx