Granite Time Series TTM

Granite-TTM-R3

by IBM watsonx · Current open-weight model family; available through the IBM Granite Hugging Face repository. No hosted inference provider deployment was listed on the reviewed model page.

Granite-TTM-R3 is IBM’s open-weight TinyTimeMixer forecasting family for multivariate numerical time series. It supports zero-shot use, few-shot adaptation, fine-tuning, exogenous variables, probabilistic multi-quantile forecasts, and high-throughput CPU or GPU inference. Because it includes multiple checkpoints, users must choose a configuration that matches their history length, forecast horizon, frequency, and model-size requirements.

Reasoning Coding
Granite-TTM-R3 is a compact IBM Granite time-series forecasting model family built for organizations that need predictions from numerical sequences without the cost and latency of large general-purpose models. It is intended for workloads such as demand planning, inventory forecasting, industrial monitoring, and large-scale batch prediction. The model is open-weight, available through IBM’s Granite Hugging Face repository, and designed to support zero-shot use as well as adaptation to domain-specific data.
Capabilities

Supported features

Fine-tuning Batch API
Model profile

Performance characteristics

0/10 Reasoning
0/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite Time Series TTM
Model type Other
Context window tokens
Maximum output tokens
Status Current open-weight model family; available through the IBM Granite Hugging Face repository. No hosted inference provider deployment was listed on the reviewed model page.
Knowledge cutoff notes

A conventional textual knowledge cutoff is not applicable to this numerical time-series forecasting model. Its training data is described as selected GIFT-Eval datasets and custom synthesized data based on KernelSynth, using non-leaking historical context.

Model notes

Granite-TTM-R3 is represented as a suite of pretrained checkpoints selected by context length, prediction length, frequency, and model-size configuration, so repository-level fields such as context_length, max_output_tokens, and parameter count are not single fixed values. The model card describes standard and Lite variants, with model sizes ranging from approximately 1M to 35M parameters for standard configurations and approximately 1M to 18M for Lite configurations. The model card reports zero-shot forecasting, few-shot adaptation with as few as approximately 1,000 samples, full fine-tuning, multivariate forecasting, exogenous or control-variable integration, and high-throughput batch inference. It is licensed under Apache 2.0. Reported throughput is approximately 7,500 samples per second on GPU and 180 samples per second on CPU for standard R3, and approximately 18,000 samples per second on GPU and 800 samples per second on CPU for R3 Lite; actual performance depends on hardware and configuration. The model page states that no inference provider deployment was available at the time of review. IBM's legacy Granite Time Series documentation directs users to Hugging Face and GitHub for current documentation.

Model guide

Granite-TTM-R3: IBM’s Fast, Open-Weight Model for Time-Series Forecasting

Granite-TTM-R3 is IBM’s open-weight TinyTimeMixer model family for efficient multivariate time-series forecasting. It supports zero-shot forecasting, few-shot adaptation, fine-tuning, exogenous variables, probabilistic multi-quantile predictions, and high-throughput CPU or GPU inference. The family includes multiple checkpoints rather than one model with a single fixed context window, forecast horizon, or parameter count.

What is Granite-TTM-R3?

Granite-TTM-R3 is IBM’s third-generation TinyTimeMixer-based time-series forecasting family. Unlike a language model, it does not generate conversational text. It receives numerical time-series data and produces forecasts for future time steps. A time series might represent sales by day, electricity demand by hour, machine temperature, traffic volume, or another measurement collected over time.

The model is designed for multivariate forecasting, meaning it can work with several related numerical signals rather than only one sequence. It also supports exogenous, or control, variables: additional information that may help explain or predict the target series. For example, a demand forecast could use historical sales together with related operational or calendar variables when the selected checkpoint and data configuration support them.

IBM publishes Granite-TTM-R3 as an open-weight model family through the Granite-TTM-R3 Hugging Face repository. The repository contains multiple checkpoint configurations selected for different context lengths, prediction lengths, frequencies, and model sizes. Therefore, “Granite-TTM-R3” does not describe one checkpoint with one universal input window or forecast horizon.

Where it fits in IBM’s Granite lineup

Granite-TTM-R3 belongs to IBM Granite’s time-series model family rather than to IBM’s conversational or code-generation models. Its role is specialized: it targets numerical forecasting and efficient deployment. IBM’s broader watsonx portfolio can be used for enterprise AI development, governance, data workflows, and model deployment, but this model itself is focused on forecasting sequences over time.

The model is also associated with IBM’s Granite Time Series Foundation Models toolkit and implementation repositories. IBM’s current model materials direct users to Hugging Face and GitHub for the most recent model and toolkit information. This positioning makes Granite-TTM-R3 more suitable for developers and data-science teams that want to run or adapt an open-weight forecasting model than for users seeking a hosted general-purpose chatbot.

How the model works

Granite-TTM-R3 builds on IBM’s TinyTimeMixer approach. Instead of relying on an expensive full self-attention architecture commonly associated with large language models, it uses lightweight mixer-based components intended to process temporal information efficiently.

The model card describes several elements of the R3 design:

  • Trend-residual decomposition: the model separates broader movement in a series from shorter-term variation, helping it represent different types of temporal behavior.
  • Multi-resolution processing: temporal patterns can be considered at different scales rather than only at one time resolution.
  • FFT-based embeddings: frequency-domain information is used as part of the representation of temporal signals.
  • GLU gating and register tokens: these architectural components help control information flow through the network.
  • Student-teacher pretraining: the training approach transfers information from a teacher model to a more compact student model.
  • Multi-quantile forecasting: the output can represent several forecast quantiles, providing more than a single point estimate when probabilistic forecasting is configured.

These design choices are intended to reduce computational requirements while retaining useful forecasting behavior across different datasets. They do not turn the model into a general reasoning system: its specialization remains numerical time-series prediction.

Forecasting capabilities and adaptation

Granite-TTM-R3 supports several ways to use a forecasting checkpoint:

  • Zero-shot forecasting: the model can be applied to an unseen dataset without task-specific training, subject to the selected checkpoint’s supported configuration.
  • Few-shot adaptation: IBM’s model materials describe adaptation with as few as approximately 1,000 samples in some settings. This is a reported capability, not a guarantee for every dataset or forecast configuration.
  • Fine-tuning: teams can adapt the model to domain-specific data when zero-shot or few-shot performance is insufficient.
  • Multivariate forecasting: multiple related series can be modeled together.
  • Exogenous-variable support: additional control or explanatory variables can be incorporated where the checkpoint and implementation support the required input structure.
  • Probabilistic forecasting: multi-quantile outputs can express a range of likely outcomes rather than only one predicted value.

In practical terms, this makes the family useful for starting with a pretrained model, testing it on a forecasting problem, and then deciding whether lightweight adaptation or full fine-tuning is worthwhile. The appropriate workflow depends on the data frequency, history length, prediction horizon, number of variables, and checkpoint configuration.

Speed, model size, and deployment efficiency

Efficiency is the main reason to consider Granite-TTM-R3. IBM reports approximately 7,500 samples per second on GPU and approximately 180 samples per second on CPU for a standard R3 configuration. For R3 Lite, IBM reports approximately 18,000 samples per second on GPU and approximately 800 samples per second on CPU.

These are vendor-reported measurements, not universal performance guarantees. Actual throughput can change with hardware, batching, preprocessing, sequence length, forecast length, and the particular checkpoint. A “sample” also represents a model-specific forecasting input, so the figures should not be treated as a direct measure of how quickly every real-world forecasting pipeline will run.

The model card describes standard configurations ranging from approximately 1 million to 35 million parameters and Lite configurations ranging from approximately 1 million to 18 million parameters. Because the repository contains a suite of checkpoints, there is no single parameter count for Granite-TTM-R3 as a whole. Smaller configurations may be preferable when CPU latency, memory use, or high-volume inference matters more than maximum model capacity.

Context length, forecast horizon, and configuration limits

A fixed context length or maximum output-token limit is not applicable to the model family in the same way it would be for a language model. Granite-TTM-R3 does not produce text tokens. Each checkpoint is associated with a particular configuration involving factors such as historical context length, prediction length, frequency, and model size.

Users must therefore select a checkpoint that matches the structure of their task. A checkpoint intended for one forecast horizon or sampling frequency may not be the right choice for another. The repository-level name alone does not identify one universal maximum history window, forecast horizon, or number of parameters.

This is an important implementation consideration. Before deployment, users should verify the checkpoint’s documented input and output shapes, sampling frequency, supported covariates, and preprocessing expectations. Treating the entire R3 repository as a single fixed model could lead to incompatible inputs or misleading comparisons.

Supported modalities and unsupported tasks

Granite-TTM-R3 is a numerical model. Its input is time-series data rather than text, images, audio, or video, and its output is a forecast rather than natural-language content or media. It does not provide native text generation, image generation, audio generation, video generation, speech output, or embeddings according to the reviewed model information.

It also is not presented as a conversational reasoning model, coding assistant, web-search system, or tool-calling agent. The model has no documented function-calling or general tool-use capability, and it does not offer a language-model-style JSON response mode. Any surrounding application logic, data retrieval, monitoring, or workflow automation must be implemented outside the forecasting model.

Pricing and access

No per-token or hosted inference price was listed for the reviewed Granite-TTM-R3 model page. The model is described as an open-weight model available through IBM Granite’s Hugging Face repository, so users can evaluate the weights and associated tooling directly rather than relying on a listed hosted API price.

That does not mean deployment is cost-free. Users may incur costs for cloud compute, storage, data preparation, monitoring, and infrastructure. The model’s small size and reported CPU throughput can reduce hardware requirements compared with larger forecasting systems, but the actual total cost depends on how many series are processed, how often forecasts are generated, and whether the model is run locally, on a private server, or in a cloud environment.

Main strengths and trade-offs

Strengths:

  • Designed specifically for numerical time-series forecasting rather than adapted from a general language model.
  • Open-weight distribution under the Apache 2.0 license, according to the supplied model information.
  • Supports zero-shot forecasting, few-shot adaptation, fine-tuning, multivariate inputs, exogenous variables, and probabilistic outputs.
  • Small model configurations and CPU-friendly inference can support resource-constrained deployments.
  • High reported throughput is useful for batch forecasting across many series.
  • A family of checkpoints provides options for different history lengths, horizons, frequencies, and model-size requirements.

Trade-offs:

  • The checkpoint family requires careful configuration selection; there is no single universal context or prediction limit.
  • Reported throughput depends on hardware and workload and should be validated using the user’s own data pipeline.
  • It is not a general-purpose AI assistant and cannot replace a language model for text, coding, conversation, or tool use.
  • Users need suitable numerical data and forecasting expertise, including decisions about frequency, covariates, preprocessing, and evaluation.
  • No hosted inference provider deployment or standard usage-based price was listed on the reviewed model page.

When to choose this model

Granite-TTM-R3 is a strong candidate when the core problem is forecasting many numerical sequences efficiently. Suitable examples include demand and sales forecasting, inventory and supply-chain planning, industrial or operational signals, financial and economic trend analysis, and real-time or batch prediction where CPU deployment is valuable.

It is particularly worth evaluating when a team wants an open-weight model that can be tested without committing to a proprietary hosted forecasting API, or when the workload requires high throughput and the model’s checkpoint configuration matches the data. Its fine-tuning and exogenous-variable support also make it more appropriate than a purely generic baseline when domain-specific forecasting signals matter.

When another option may be more appropriate

A general-purpose language model is a better choice for natural-language analysis, report generation, coding, question answering, or conversational interfaces. A broader multimodal model is more appropriate when the task combines forecasting with images, documents, audio, or video. A hosted forecasting service may be preferable when a team does not want to manage model files, hardware, preprocessing, or inference operations.

Another forecasting model may also be preferable if its documented checkpoint directly matches the required context length, frequency, forecast horizon, or covariate structure more closely. Granite-TTM-R3’s family-based design offers flexibility, but it also means that users must inspect the specific checkpoint rather than assuming that every R3 variant supports identical inputs and outputs.

Bottom line

Granite-TTM-R3 is a specialized, efficient forecasting model family rather than an all-purpose AI model. Its value lies in compact open-weight checkpoints, support for several forecasting and adaptation modes, probabilistic predictions, and high reported CPU or GPU throughput. It is best suited to teams working with numerical time-series data that need scalable inference and control over deployment. The main practical caution is that R3 is a collection of configurations, so context length, forecast horizon, model size, and input requirements must be checked at the individual checkpoint level.


Answers to Frequently Asked Questions

What is Granite-TTM-R3 used for?
Granite-TTM-R3 is designed for numerical time-series forecasting, such as predicting sales, electricity demand, inventory needs, traffic volume, machine temperature, and other measurements collected over time. It supports univariate and multivariate forecasting, and some checkpoints can use exogenous variables.
How fast and how large is Granite-TTM-R3?
IBM reports approximately 7,500 samples per second on GPU and 180 samples per second on CPU for a standard R3 configuration. For R3 Lite, IBM reports approximately 18,000 samples per second on GPU and 800 samples per second on CPU. Model sizes range from approximately 1 million to 35 million parameters for standard configurations and from approximately 1 million to 18 million parameters for Lite configurations. Actual performance depends on hardware, batching, preprocessing, sequence length, forecast length, and checkpoint.
What are the main capabilities of Granite-TTM-R3?
Granite-TTM-R3 supports zero-shot forecasting, few-shot adaptation, fine-tuning, multivariate forecasting, exogenous variables where supported, and probabilistic multi-quantile predictions. It is built for efficient inference and is available in multiple checkpoint configurations for different frequencies, context lengths, forecast horizons, and model sizes.
Is Granite-TTM-R3 a language model or chatbot?
No. Granite-TTM-R3 is a specialized forecasting model, not a conversational language model. It accepts numerical time-series data and produces forecasts, but it does not natively generate text, write code, answer questions, use tools, or process images, audio, or video.
Does Granite-TTM-R3 have one fixed context length or forecast horizon?
No. Granite-TTM-R3 is a family of checkpoints rather than one model with a universal context length or forecast horizon. Users must select a checkpoint that matches the required historical window, prediction length, sampling frequency, input and output shapes, supported covariates, and preprocessing requirements.


Sources 4
Provider

About IBM watsonx