Granite TTM

granite-ttm-1024-96-r2

by IBM watsonx · Available through IBM watsonx.ai; downloadable model branch

IBM Granite TTM 1024-96-R2 is a compact numerical forecasting model for regularly sampled multivariate time series. It uses 1,024 historical observations per channel to predict up to 96 future observations, supports zero-shot use and fine-tuning, and prioritizes efficient inference over general-purpose language, coding, multimodal, or tool-use capabilities.

Reasoning Coding
IBM Granite TTM 1024-96-R2 is a specialized time-series forecasting model, not a conversational language model. It analyzes numerical observations from one or more related channels and predicts the next values in the sequence. The model is configured for a 1,024-point input context and a forecast horizon of up to 96 points, with an emphasis on efficient inference, zero-shot use, and task-specific fine-tuning.
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Granite TTM
Model type Other
Context window 1K tokens
Maximum output 96 tokens
Release date October 2024
Status Available through IBM watsonx.ai; downloadable model branch
Knowledge cutoff notes

This is a time-series forecasting model rather than a knowledge-grounded language model. IBM's available model documentation does not specify a conventional textual knowledge cutoff.

Model notes

The canonical watsonx.ai model ID is ibm/granite-ttm-1024-96-r2. It requires at least 1,024 historical data points per channel and forecasts up to 96 future data points. The model is a branch of the IBM Granite Tiny Time Mixer R2 repository and corresponds to the TTM-E configuration described in IBM model materials. The downloadable repository is licensed under Apache 2.0. IBM documentation describes the broader Granite TTM release as containing approximately 1 million parameters and supporting zero-shot forecasting, fine-tuning, channel mixing, exogenous variables, and static categorical features through the associated time-series framework. The max_output_tokens field is represented here as the model's maximum forecast length of 96 time points, not text tokens. Pricing is expressed using IBM watsonx.ai pricing classes rather than per-token rates.

Cost

Model pricing

Input watsonx.ai API pricing class 14
Output watsonx.ai API pricing class 15
Model guide

IBM Granite TTM 1024-96-R2 for Efficient Multivariate Forecasting

IBM Granite TTM 1024-96-R2 is a compact Tiny Time Mixer foundation model for multivariate time-series forecasting. It consumes 1,024 historical observations per channel and forecasts up to 96 future observations, making it a practical choice for regularly sampled operational, energy, traffic, manufacturing, sales, and sensor data when low resource use matters more than general-purpose AI capabilities.

What IBM Granite TTM 1024-96-R2 is

IBM Granite TTM 1024-96-R2 is a Tiny Time Mixer foundation model from IBM's Granite time-series collection. Its canonical watsonx.ai model identifier is ibm/granite-ttm-1024-96-r2. The model is designed for multivariate time-series forecasting: it takes numerical measurements over time and estimates future measurements for one or more channels.

A channel can represent a single measured series, such as electricity demand, temperature, network traffic, product sales, or a machine sensor. “Multivariate” means that several channels can be considered together rather than forecasting every series in complete isolation. This is useful when the variables are related, such as demand and price, or vibration and motor load.

The 1024-96-R2 name describes the model's central configuration. It uses up to 1,024 historical observations per channel as context and forecasts up to 96 future observations per channel. The time represented by those observations depends on the sampling interval. With hourly data, 96 points correspond to four days; with 10-minute data, they represent 16 hours.

Where it fits in IBM's Granite lineup

Granite TTM 1024-96-R2 is one branch of IBM's Granite Tiny Time Mixer R2 time-series model family. It sits alongside IBM's broader foundation-model portfolio but serves a very different purpose from Granite language models. It does not generate prose, answer questions, write code, or act as a general-purpose reasoning assistant.

IBM exposes the model through watsonx.ai's time-series forecasting functionality, and the model branch is also available for download through IBM Granite's Hugging Face organization. This gives teams a choice between using IBM's managed environment and evaluating or deploying the downloadable model with supported local or custom tooling.

The broader Granite TTM materials describe a model family of approximately one million parameters. That compact scale is important to the model's positioning: it is intended to deliver useful forecasting without the hardware requirements normally associated with large generative models. The exact branch discussed here remains tied to the 1,024-observation context and 96-observation forecast configuration.

Core forecasting specifications

SpecificationVerified detail
ProviderIBM
Model familyGranite TTM, or Tiny Time Mixer
Model identifieribm/granite-ttm-1024-96-r2
Input typeNumerical time-series data
Context length1,024 historical observations per channel
Maximum forecast horizon96 future observations per channel
Primary taskMultivariate point forecasting
Fine-tuningSupported through the associated time-series tooling
Downloadable licenseApache 2.0 for the model repository

IBM's watsonx.ai documentation identifies a requirement for at least 1,024 data points per channel in an API request. In practice, that means the model is not a good fit for a newly launched sensor or business metric with only a short history. The forecast length is also a model configuration rather than a text-token limit: the value 96 refers to future time points, not generated words or tokens.

How input and forecasting work

Before inference, IBM recommends externally standard-scaling each time-series channel through the associated preprocessing utilities. Standard scaling adjusts a series so that differences in magnitude between channels do not dominate the forecast. For example, a power-consumption channel measured in thousands and a temperature channel measured in tens may need normalization before being processed together.

The model can be used in a zero-shot setting, meaning that a pretrained model is applied to a compatible dataset without first fine-tuning it for that specific task. This can be useful for rapid experimentation or for building an initial baseline. Fine-tuning is available when a team has representative historical data and needs to adapt the model to a particular domain, set of channels, or forecasting behavior.

The wider Granite TTM tooling supports channel-independent and channel-mixing approaches. Channel mixing allows relationships between series to contribute to the forecast. IBM's model materials also describe support in the associated framework for exogenous or control variables and static categorical information where the relevant workflow supports them. These capabilities should not be interpreted as unrestricted support for arbitrary text or mixed media: the core model remains a numerical time-series forecaster.

Main strengths and trade-offs

The clearest strength of Granite TTM 1024-96-R2 is specialization. It is narrowly focused on a common forecasting problem and is much smaller than general-purpose generative AI systems. That makes it a sensible candidate for repeated forecasts across many operational series, especially when CPU-only or resource-constrained inference is desirable.

  • Efficient deployment: The compact model design is intended to support inference without a large GPU deployment.
  • Useful fixed window: A 1,024-point history and 96-point horizon cover many minute-level and hourly forecasting scenarios.
  • Zero-shot starting point: Teams can evaluate a pretrained model before investing in fine-tuning.
  • Multivariate modeling: Related channels can be handled in a forecasting workflow designed for cross-series information.
  • Open downloadable branch: The Hugging Face repository is identified as Apache 2.0 licensed, supporting local evaluation subject to the repository's terms and deployment requirements.

These benefits come with significant trade-offs. The fixed configuration is restrictive for tasks that require a much shorter or substantially longer input history or forecast horizon. Forecast quality can vary with sampling frequency, scaling, data distribution, channel relationships, and similarity between the target data and the model's pretraining data. A zero-shot result should therefore be compared with domain-specific baselines before it is used in an operational decision process.

Modalities, reasoning, coding, and tools

Granite TTM 1024-96-R2 accepts numerical time-series inputs and produces numerical forecast outputs. It has no verified text, image, audio, or video input or output capability. It is not multimodal in the ordinary generative-AI sense, and it does not provide conversational responses or structured JSON generation as a model feature.

Reasoning and coding are not meaningful primary capabilities for this model. It does not reason over natural-language instructions, generate software, call functions, browse the web, or use external tools. An application can place its numerical forecast into a larger workflow that includes other software, but that orchestration should not be attributed to the model itself.

Editorially, the model's strongest scores in the supplied evaluation are its speed and cost ratings, while its reasoning and coding ratings are very low. Those are comparative editorial assessments, not IBM-published benchmark results. The practical interpretation is that this model trades broad capability for a lightweight, focused forecasting workload.

Pricing and availability

The model is available through IBM watsonx.ai and as a downloadable Granite time-series model branch. IBM's watsonx.ai foundation-model documentation assigns the model to API pricing class 14 for input and pricing class 15 for output. The supplied information does not provide a public per-request, per-token, or per-forecast monetary rate for those classes, so an exact price cannot be stated responsibly.

The pricing labels also require careful interpretation. This is not a text-generation model whose output is naturally measured in tokens. The model's output is a forecast of up to 96 numerical time points. Users considering managed watsonx.ai deployment should verify the current IBM pricing schedule, account terms, region, and usage-metering rules rather than estimating cost from the model's maximum forecast length.

For local or custom deployment, the downloadable repository provides another route, but infrastructure, engineering, storage, monitoring, and operational costs still apply. The Apache 2.0 repository license does not mean that managed watsonx.ai usage is free.

Best use cases

This model is a strong candidate when the data is regularly sampled, numerical, and long enough to provide at least the required historical context. Practical examples include:

  • Forecasting electricity demand, renewable generation, or facility load.
  • Predicting traffic volumes, network performance, or service utilization.
  • Estimating manufacturing readings such as temperature, pressure, vibration, or throughput.
  • Forecasting inventory, sales, or other regularly recorded operational metrics.
  • Running many forecasting experiments where fast, low-resource inference is more important than general-purpose model functionality.

It is particularly useful as a baseline. A team can first test zero-shot forecasts, compare them with statistical or machine-learning baselines, and then fine-tune the model if the domain provides enough representative data.

When to choose Granite TTM 1024-96-R2

Choose Granite TTM 1024-96-R2 when you need a compact, specialized forecaster for a 1,024-point historical window and a horizon of up to 96 points. It is also a reasonable choice when CPU-friendly inference, downloadable deployment, or integration with IBM's watsonx.ai forecasting workflow matters more than chat, coding, document understanding, or tool use.

Another option may be more appropriate when the data has fewer than 1,024 observations per channel, the required horizon is longer than 96 points, or the sampling pattern is irregular and does not fit the model's expected configuration. A different forecasting method may also be preferable if the business requires calibrated uncertainty estimates, a highly interpretable statistical model, or a custom architecture tailored to unusual data. For text instructions, code generation, image analysis, audio, or general reasoning, a language or multimodal model is the relevant type of tool instead.

Bottom line

IBM Granite TTM 1024-96-R2 is best understood as a focused forecasting component rather than a general AI assistant. Its defining characteristics are the 1,024-observation input requirement, the 96-observation forecast horizon, multivariate numerical processing, fine-tuning support, and compact deployment profile. Those constraints make it unsuitable for many generative-AI tasks, but they also make its role clear: it is designed to provide efficient forecasts for regularly sampled operational and sensor data.


Answers to Frequently Asked Questions

How much historical data and how many future points does IBM Granite TTM 1024-96-R2 support?
The model uses up to 1,024 historical observations per channel and forecasts up to 96 future observations per channel. The real-world time span depends on the sampling interval; for example, 96 hourly observations represent four days.
What is IBM Granite TTM 1024-96-R2 used for?
IBM Granite TTM 1024-96-R2 is a compact foundation model for multivariate numerical time-series forecasting. It can forecast data such as electricity demand, temperature, network traffic, sales, inventory, and machine-sensor readings.
What are the main limitations of IBM Granite TTM 1024-96-R2?
The model requires at least 1,024 data points per channel in the relevant API workflow and has a maximum forecast horizon of 96 points. It is intended for regularly sampled numerical time series and does not generate text, code, images, audio, or conversational responses. Forecast quality should be compared with domain-specific baselines before operational deployment.
Does IBM Granite TTM 1024-96-R2 support multivariate forecasting and fine-tuning?
Yes. The model is designed for multivariate forecasting, allowing related numerical channels to be considered together. It supports zero-shot use and fine-tuning through the associated time-series tooling, with external standard scaling recommended before inference.


Sources 5
Provider

About IBM watsonx