Model catalog

IBM watsonx Models

Browse the AI models associated with IBM watsonx. Compare current and historical models by family, capabilities, context window, availability and intended use.

56 models tracked
56 Total models
26 Model families
8 Model types
50 Current / accessible
All models

IBM watsonx model catalog

Zero-shot forecasting of regularly sampled numerical time series across changing temporal resolutions

Type Other
Context 4K
Reasoning 1/10
Speed 8/10
Status

Current open-weight model; r1.1 revision available

Input No hosted API pricing documented; self-hosted open weights
Output No hosted API pricing documented; self-hosted open weights
View model →

Japanese text generation, summarization, classification, extraction, question answering, and Japanese-English translation in legacy IBM enterprise deployments

Type General Purpose
Context 4K
Reasoning 4/10
Speed 7/10
Status

Deprecated; withdrawn from IBM Software Hub 5.2.0 and removed from watsonx.ai in 2.2.0

View model →

Multilingual enterprise question answering, retrieval-augmented generation, summarization, extraction, classification, and text generation in English, German, Spanish, French, and Portuguese

Type General Purpose
Context 8K
Reasoning 4/10
Speed 3/10
Status

Retired; deprecated January 15, 2025 and withdrawn from standard watsonx.ai availability April 16, 2025

View model →

Fine-tuning, domain adaptation, text generation, summarization, extraction, classification, and question answering

Type General Purpose
Context 4K
Reasoning 4/10
Speed 7/10
Fine-tuning
Status

Available for deployment on demand in IBM watsonx.ai; also available as an Apache 2.0 open-weight model through IBM's Hugging Face organization.

Input $0.0006 per 1,000 tokens
Output $0.0006 per 1,000 tokens
View model →

Fine-tuning, long-context document processing, enterprise text classification, extraction, summarization, question answering, and self-managed deployment

Type General Purpose
Context 131K
Reasoning 5/10
Speed 6/10
Fine-tuning
Status

Available; base model intended for fine-tuning and dedicated deployment

View model →

Self-hosted enterprise assistants, long-context document analysis, RAG, summarization, extraction, multilingual text workflows, and function calling

Type General Purpose
Context 131K
Reasoning 5/10
Speed 7/10
Tool use Fine-tuning Streaming
Status

Available as an open-weight model; superseded by Granite-3.3-8B-Instruct for newer deployments

View model →

Long-context enterprise RAG, document and meeting summarization, information extraction, multilingual dialogue, code-related tasks, function calling, and self-hosted deployments

Type Reasoning
Context 131K
Reasoning 7/10
Speed 7/10
Tool use Fine-tuning Streaming
Status

Available; legacy in some IBM watsonx catalogs

View model →

Local or dedicated enterprise assistants, RAG, long-document summarization, multilingual text workflows, coding assistance, reasoning, and resource-conscious inference

Type General Purpose
Context 131K
Reasoning 6/10
Speed 8/10
Tool use Fine-tuning Streaming
Status

Available; open-weight model with local deployment and dedicated IBM watsonx deployment options

View model →

Self-hosted enterprise assistants, long-context RAG, summarization, multilingual text tasks, coding assistance, function calling, and cost-sensitive deployments.

Type General Purpose
Context 131K
Reasoning 6/10
Speed 8/10
Tool use Fine-tuning Streaming
Status

Retired from IBM watsonx.ai deploy-on-demand service on 2026-02-22; open-weight checkpoint remains available for self-hosted and compatible third-party deployments.

Input $0.0002 per 1,000 tokens, historical IBM watsonx Developer Hub listing; no current IBM-hosted price verified after withdrawal.
Output $0.0002 per 1,000 tokens, historical IBM watsonx Developer Hub listing; no current IBM-hosted price verified after withdrawal.
View model →

Low-latency local inference, long-context text generation, RAG, multilingual assistants, coding assistance, and lightweight tool-calling agents

Type Lightweight
Context 128K
Reasoning 5/10
Speed 8/10
Tool use Fine-tuning
Status

Current open-weight instruct model

View model →

Enterprise RAG, multi-tool agents, function calling, customer-support automation, multilingual instruction following, and long-context workloads

Type General Purpose
Context 131K
Reasoning 7/10
Speed 8/10
Tool use Fine-tuning
Status

Available; open-weight instruct model

Input $0.0000636 per 1,000 tokens on IBM watsonx.ai multitenant hardware
Output $0.000265 per 1,000 tokens on IBM watsonx.ai multitenant hardware
View model →

Efficient enterprise assistants, multilingual text generation, RAG, classification, extraction, summarization, coding assistance, fill-in-the-middle completion, structured JSON, and tool-calling workflows

Type Lightweight
Context 128K
Reasoning 5/10
Speed 9/10
Tool use Fine-tuning Streaming
Status

Current; open-weight instruct model; available for download and listed for deploy-on-demand use in IBM watsonx.ai

View model →
IBM watsonx logo
Granite 4.1

Granite 4.1 3B

Efficient local or private deployment, multilingual enterprise text processing, RAG, summarization, extraction, coding assistance, function calling, and lightweight AI assistants.

Type Lightweight
Context 131K
Reasoning 5/10
Speed 9/10
Tool use Fine-tuning Streaming
Status

Currently available open-weight model; earlier-generation Granite model superseded by the newer Granite 4.2 family, with no verified deprecation or shutdown date.

Input No official hosted API token price published; downloadable weights are available under Apache 2.0.
Output No official hosted API token price published; downloadable weights are available under Apache 2.0.
View model →
IBM watsonx logo
Granite 4.1

Granite 4.1 8B

Self-hosted enterprise assistants, multilingual text generation, RAG, coding assistance, structured extraction, and tool-calling agents

Type General Purpose
Context 131K
Reasoning 6/10
Speed 8/10
Tool use Fine-tuning Streaming
Status

Current open-weight instruction model

Input Not available for official per-token hosted pricing
Output Not available for official per-token hosted pricing
View model →
IBM watsonx logo
Granite 4.1

Granite 4.1 30B

Self-hosted enterprise assistants, long-context RAG, multilingual applications, coding, structured extraction, and tool-calling agents

Type General Purpose
Context 131K
Reasoning 7/10
Speed 5/10
Tool use Fine-tuning Streaming
Status

Current open-weight instruct model; publicly available

View model →
IBM watsonx logo
Granite 4.2

Granite 4.2 3B

Efficient reasoning, coding assistance, tool calling, multilingual dialogue, local deployment, and lightweight enterprise agents

Type Reasoning
Context 131K
Reasoning 8/10
Speed 8/10
Tool use Fine-tuning
Status

Current; publicly available open-weight model

View model →
IBM watsonx logo
Granite 4.2

Granite 4.2 8B

Local or self-hosted reasoning, coding assistants, tool calling, multilingual dialogue, retrieval-augmented generation, and agentic workflows

Type Reasoning
Context 131K
Reasoning 8/10
Speed 8/10
Tool use Fine-tuning Streaming
Status

Current; open-weight and downloadable

Input No official hosted API price; downloadable Apache 2.0 weights
Output No official hosted API price; inference infrastructure costs depend on deployment
View model →
IBM watsonx logo
Granite 4.2

Granite 4.2 30B

Self-hosted enterprise reasoning, coding agents, multilingual applications, tool calling, long-context workflows, and organizations requiring Apache 2.0 licensing

Type Reasoning
Context 131K
Reasoning 8/10
Speed 6/10
Tool use Fine-tuning
Status

Current; open-weight and available for download

Input No official IBM hosted API price; model weights are available for self-hosted deployment
Output No official IBM hosted API price; model weights are available for self-hosted deployment
View model →
IBM watsonx logo
Granite 7B

Granite-7B-Lab

Self-hosted English text generation, summarization, extraction, classification, experimentation with LAB-aligned open-weight models, and resource-conscious deployments

Type General Purpose
Context 8K
Reasoning 5/10
Speed 7/10
Fine-tuning
Status

Withdrawn from IBM watsonx.ai on 2025-01-07; open-weight checkpoint remains available for self-hosted deployment

View model →
IBM watsonx logo
Granite 13B V2

granite-13b-chat-v2

English enterprise chat, retrieval-augmented generation, question answering, summarization, extraction, and classification

Type General Purpose
Context 8K
Reasoning 4/10
Speed 6/10
Status

Withdrawn; access ended January 19, 2025

Input $0.0006 per 1,000 input tokens, historical watsonx.ai rate
Output $0.0006 per 1,000 output tokens, historical watsonx.ai rate
View model →

Self-hosted coding assistants, code generation, code explanation, code conversion, repository-scale prompts, and historical or reproducible research

Type Coding
Context 128K
Reasoning 4/10
Speed 8/10
Fine-tuning
Status

Retired; withdrawn from IBM watsonx.ai on 2025-07-17

Input $0.0006 per 1,000 input tokens historically on IBM watsonx.ai
Output $0.0006 per 1,000 output tokens historically on IBM watsonx.ai
View model →

Local code generation, explanation, repair, translation, and research with an open Apache 2.0 model

Type Coding
Context 4K
Reasoning 5/10
Speed 7/10
Fine-tuning
Status

Deprecated; still publicly available for historical and scientific use

View model →

Schema linking and relevant-table or relevant-column selection in text-to-SQL pipelines

Type Coding
Context 8K
Reasoning 3/10
Speed 3/10
Status

Available; deploy on demand only

Input Not available as pay-as-you-go token pricing; deploy-on-demand infrastructure pricing applies
Output Not available as pay-as-you-go token pricing; deploy-on-demand infrastructure pricing applies
View model →

Natural-language-to-SQL generation over structured databases, analytics assistants, and the SQL-generation stage of text-to-SQL pipelines.

Type Coding
Context 8K
Reasoning 5/10
Speed 4/10
Fine-tuning
Status

Available; deploy-on-demand model in IBM watsonx.ai

Input Not publicly listed as a standalone token price; IBM lists the model as deploy-on-demand.
Output Not publicly listed as a standalone token price; IBM lists the model as deploy-on-demand.
View model →

Self-hosted coding assistants, code generation, code conversion, code explanation, and programming experiments where an Apache 2.0 open-weight model is preferred.

Type Coding
Context 8K
Reasoning 5/10
Speed 4/10
Fine-tuning Streaming
Status

Available as an open-weight model and documented for IBM watsonx.ai deploy-on-demand use; deprecated and withdrawn from the watsonx.ai multitenant offering according to IBM's 2025 lifecycle notice.

View model →

Self-hosted code generation, code explanation, code repair, code conversion, coding assistants, and legacy Granite Code compatibility

Type Coding
Context 8K
Reasoning 6/10
Speed 4/10
Fine-tuning Streaming
Status

Legacy/deprecated; available in some IBM watsonx environments and retained for historical or scientific use

View model →

Fast, low-footprint English semantic search, RAG retrieval, similarity matching, and vector indexing

Type Embedding
Context 512
Reasoning 1/10
Speed 9/10
Status

Legacy; downloadable and usable, but superseded by Granite Embedding Small English R2

View model →

Low-latency multilingual semantic search, retrieval-augmented generation, vector search, document similarity, long-document retrieval, and cross-lingual code retrieval

Type Lightweight
Context 33K
Reasoning 2/10
Speed 9/10
Status

Current; open-weight model

View model →

Low-cost multilingual semantic search, cross-lingual retrieval, vector databases, similarity matching, and RAG pipelines

Type Embedding
Context 512
Reasoning 1/10
Speed 9/10
Fine-tuning
Status

Retired from IBM watsonx.ai; open-weight checkpoint remains available through IBM's Hugging Face organization

Input No official hosted API price verified; open-weight model intended for local or third-party deployment
Output Not applicable to token generation; produces 384-dimensional embeddings
View model →

English semantic search, vector retrieval, RAG, similarity matching, enterprise knowledge-base search, and local embedding deployment

Type Embedding
Context 512
Reasoning 1/10
Speed 8/10
Status

Legacy/superseded but downloadable and usable

View model →

Multilingual semantic search, retrieval-augmented generation, cross-lingual retrieval, long-document search, similarity, and code retrieval

Type Other
Context 33K
Reasoning 1/10
Speed 7/10
Status

Current open-weight model

View model →

English semantic search, vector retrieval, retrieval-augmented generation, document similarity, clustering, and enterprise information retrieval

Type Embedding
Context 8K
Reasoning 1/10
Speed 8/10
Status

Current open-weight model

Input No official hosted API price; downloadable weights are available under the Apache 2.0 license
Output No official hosted API price; downloadable weights are available under the Apache 2.0 license
View model →

Compact English semantic search, retrieval-augmented generation, document similarity, and private vector-search deployments

Type Embedding
Context 8K
Speed 8/10
Status

Current open-weight model

View model →

Multilingual semantic search, vector retrieval, RAG, clustering, similarity matching, and text classification features

Type Embedding
Context 512
Reasoning 1/10
Speed 8/10
Fine-tuning
Status

Retired; shutdown completed 2026-08-08

View model →
IBM watsonx logo
Granite Geospatial

granite-geospatial-biomass

Above-ground biomass mapping, forest monitoring, carbon-stock estimation, ecological analysis, and satellite-based remote-sensing research

Type Other
Reasoning 1/10
Speed 6/10
Multimodal Image input Fine-tuning
Status

Available open-weight model

View model →

Canopy-height mapping, forest monitoring, vegetation analysis, carbon-cycle research, ecological assessment, and remote-sensing experimentation.

Type Other
Reasoning 1/10
Speed 5/10
Image input Fine-tuning
Status

Available as downloadable open-weight model; not deployed by an inference provider; IBM repository disclosure states that the project is not maintained as an IBM product.

View model →

Land surface temperature estimation, urban heat island analysis, satellite-based environmental monitoring, and temporal gap filling

Type Other
Reasoning 1/10
Speed 5/10
Multimodal Image input Fine-tuning
Status

Available as an open-weight model for local inference

View model →
IBM watsonx logo
Granite Geospatial

granite-geospatial-uki

Remote-sensing feature extraction, Earth-observation transfer learning, and regional flood-segmentation workflows using multispectral and SAR satellite imagery

Type Other
Reasoning 1/10
Speed 5/10
Multimodal Image input Fine-tuning
Status

Available as an open-weight research model; no hosted inference provider is currently listed

Input No official hosted API pricing; open-weight model
Output No official hosted API pricing; open-weight model
View model →

Spatial downscaling of weather forecasts, reanalysis data, and climate simulations

Type Other
Reasoning 0/10
Speed 0/10
Fine-tuning
Status

Current and available as downloadable open-weight model

View model →
IBM watsonx logo
Granite Guardian

Granite Guardian 3.8B

Prompt and response safety classification, jailbreak detection, RAG groundedness and relevance checks, hallucination evaluation, and enterprise AI guardrails

Type Other
Context 131K
Reasoning 5/10
Speed 7/10
Status

Deprecated in IBM watsonx.ai documentation; open-weight model and official model materials remain available

Input $0.0002 per 1,000 input tokens in IBM watsonx.ai documentation
Output $0.0002 per 1,000 output tokens in IBM watsonx.ai documentation
View model →
IBM watsonx logo
Granite Guardian

Granite Guardian 4.1 8B

AI safety guardrails, jailbreak detection, RAG groundedness and relevance checks, function-call hallucination detection, custom criteria evaluation, and best-of-N response ranking

Type Other
Context 8K
Reasoning 6/10
Speed 7/10
Status

Current and available as an open-weight model

Input Not applicable; no official hosted token price found for the downloadable model
Output Not applicable; no official hosted token price found for the downloadable model
View model →
IBM watsonx logo
Granite Speech 4.1

Granite-Speech-4.1-2B-Plus

Multilingual speech-to-text with speaker labels, word-level timestamps, keyword biasing, and self-hosted enterprise transcription

Type Other
Context 4K
Reasoning 2/10
Speed 7/10
Multimodal Audio input
Status

Current open-weight model

View model →
IBM watsonx logo
Granite Speech 4.1

Granite Speech 4.1 2B

Multilingual offline transcription, speech-to-text, speech translation, subtitle generation, and domain-specific recognition with keyword biasing

Type Other
Context 128K
Reasoning 2/10
Speed 8/10
Multimodal Audio input Fine-tuning
Status

Current open-weight model

View model →
IBM watsonx logo
Granite Speech 4.1

Granite Speech 4.1 2B NAR

High-throughput multilingual speech transcription where low inference latency is more important than maximum recognition accuracy.

Type Speech Recognition
Context 4K
Reasoning 2/10
Speed 10/10
Audio input
Status

Current; open-weight; Apache 2.0

Input No official hosted API pricing; model weights are available for self-hosted use.
Output No official hosted API pricing; model weights are available for self-hosted use.
View model →

High-throughput, low-latency English speech-to-text transcription on local, edge, and enterprise systems

Type Speech Recognition
Reasoning 1/10
Speed 10/10
Audio input Fine-tuning Streaming
Status

Current; open-weight model

View model →

Fast local English speech-to-text transcription, edge-device ASR, research, browser demos, and high-throughput noncommercial batch processing

Type Other
Reasoning 1/10
Speed 10/10
Audio input Streaming
Status

Current; research and noncommercial use only

Input No official hosted API price; downloadable model weights
Output No official hosted API price; downloadable model weights
View model →
IBM watsonx logo
Granite Time Series

granite-ttm-512-96-r2

Lightweight multivariate forecasting at minute- and hour-level resolutions, especially when 512 historical observations and a 96-point forecast horizon are appropriate.

Type Other
Context 512
Reasoning 0/10
Speed 9/10
Fine-tuning
Status

Available

Input $0.13 per 1,000 input data points
Output $0.38 per 1,000 output data points
View model →
IBM watsonx logo
Granite TimeSeries PatchTST-FM

Granite TimeSeries PatchTST-FM-r2

Zero-shot forecasting of demand, prices, energy loads, traffic, telemetry, and other regularly sampled numerical time series; probabilistic forecasts and uncertainty intervals; local or self-managed inference.

Type Other
Context 8K
Reasoning 0/10
Speed 8/10
Status

Current; open-weight model

View model →
IBM watsonx logo
Granite Time Series TTM

Granite-TTM-R3

High-throughput multivariate time-series forecasting, zero-shot or few-shot forecasting, probabilistic prediction, exogenous-variable forecasting, and CPU-friendly production deployment.

Type Other
Reasoning 0/10
Speed 9/10
Fine-tuning
Status

Current open-weight model family; available through the IBM Granite Hugging Face repository. No hosted inference provider deployment was listed on the reviewed model page.

View model →

Efficient multivariate forecasting for regularly sampled energy, traffic, manufacturing, network, sales, and sensor data

Type Other
Context 1K
Reasoning 1/10
Speed 9/10
Fine-tuning
Status

Available through IBM watsonx.ai; downloadable model branch

Input watsonx.ai API pricing class 14
Output watsonx.ai API pricing class 15
View model →

Multivariate forecasting when at least 1,536 historical observations per channel are available, especially demand, traffic, electricity, manufacturing, finance, and other minute- or hour-level forecasting tasks.

Type Other
Context 2K
Reasoning 0/10
Speed 9/10
Fine-tuning
Status

Available

Input $0.00013 per 1,000 data points
Output $0.00038 per 1,000 data points
View model →

Chart extraction, table parsing, semantic key-value extraction, visual document processing, and local enterprise RAG pipelines

Type Multimodal
Reasoning 5/10
Speed 7/10
Multimodal Image input
Status

Current and downloadable; newer Granite Vision 4.1 4B version available

Input No official hosted API price found; open-weight model intended for self-hosted deployment
Output No official hosted API price found; open-weight model intended for self-hosted deployment
View model →

Enterprise document understanding, chart and table extraction, OCR-oriented image analysis, visual question answering, and multimodal RAG

Type Multimodal
Context 131K
Reasoning 5/10
Speed 8/10
Multimodal Image input Fine-tuning
Status

Available; legacy relative to Granite Vision 4.0 3B Vision

View model →
IBM watsonx logo
Granite Vision 4.1

Granite Vision 4.1 4B

Structured extraction from charts, tables, invoices, forms, and enterprise document images

Type Multimodal
Reasoning 3/10
Speed 7/10
Multimodal Image input Fine-tuning
Status

Current open-weight model

Input No official hosted API price specified; downloadable open-weight model
Output No official hosted API price specified; downloadable open-weight model
View model →

English semantic search, dense retrieval, vector indexing, duplicate-question matching, and lightweight retrieval-augmented generation

Type Other
Context 512
Speed 8/10
Status

Deprecated; withdrawn in Dallas on 2026-09-08 and scheduled for withdrawal in other listed regions on 2027-01-12.

Input USD 0.0001 per 1,000 input tokens on watsonx.ai
Output Not separately priced; the model returns embedding vectors rather than generated output tokens
View model →

English semantic search, vector database indexing, retrieval-augmented generation, document matching, and query-passage retrieval

Type Embedding
Context 512
Reasoning 1/10
Speed 7/10
Status

Deprecated; scheduled for withdrawal on 2027-01-12

Input $0.0001 per 1,000 input tokens
View model →