What is Granite TimeSeries PatchTST-FM-r2?
Granite TimeSeries PatchTST-FM-r2 is IBM’s approximately 385-million-parameter foundation model for numerical time-series analysis. It is the open-weight continuation of PatchTST-FM-r1 and is designed primarily for zero-shot forecasting: the model can be used on a new time series without task-specific retraining for that individual dataset.
A time series is a sequence of measurements recorded in order, such as hourly electricity demand, daily sales, sensor readings, traffic counts, or asset prices. The model examines a historical context and estimates future values. It can also support probabilistic forecasting, which represents uncertainty rather than returning only one predicted value, and missing-value imputation through the Granite TSFM inference pipeline.
IBM positions this model within the Granite Time Series family, rather than as a general-purpose language or multimodal model. The model card identifies ibm-granite/granite-timeseries-patchtst-fm-r2 as its canonical Hugging Face identifier. It is intended for local or self-managed forecasting workflows.
Core capabilities and supported data
PatchTST-FM-r2 works with numerical time-series data rather than natural-language prompts. Its documented uses include zero-shot forecasting, probabilistic forecasts, flexible forecast lengths, and missing-value imputation. It does not accept images, audio, video, or conversational text as model inputs, and it is not designed to generate language, images, speech, or other media.
The model supports a context of up to 8,192 time steps. In practical terms, this is the maximum historical sequence length identified in the supplied model information. The appropriate real-world duration depends on the sampling interval: 8,192 hourly observations cover a different period from 8,192 daily or minute-level observations.
Its prediction head produces 99 quantiles. Quantiles are useful for expressing a range of plausible future outcomes—for example, a central forecast alongside lower and upper estimates. This makes the model more suitable for planning and risk-sensitive decisions than a system that produces only a single point estimate, although users still need to evaluate whether the resulting uncertainty intervals are well calibrated for their own data.
| Specification | Documented detail |
|---|---|
| Provider | IBM |
| Model family | Granite TimeSeries PatchTST-FM |
| Parameters | Approximately 385 million |
| Maximum context | 8,192 time steps |
| Prediction output | Point and probabilistic forecasts with 99 quantiles |
| Primary data type | Numerical time series |
| Hosted API pricing | No official model-specific price listed |
| License | Apache License 2.0 or OpenMDW License 1.0 |
How the architecture works
The model extends the PatchTST design with Conformer-style blocks. Each block combines multi-head self-attention with temporal convolution and feed-forward layers. In accessible terms, attention helps the model relate patterns that are far apart in the sequence, while convolution focuses on local changes and short-term shapes.
PatchTST-FM-r2 divides a time series into overlapping patches instead of processing every point as an entirely separate modeling unit. The documented patch length is 16 time steps and the stride is 8, so neighboring patches overlap. This can help preserve local continuity while giving the model a structured representation of longer sequences.
The architecture contains 30 blocks and alternates convolution kernel sizes of 5 and 3 in a repeating pattern. Training uses Hamming-window loss weighting, while inference uses overlap-and-add forecasting to reduce discontinuities between adjacent predicted patches. These are implementation details that matter most to users evaluating reproducibility, behavior, or resource requirements; ordinary users do not need to configure them to use the published pipeline.
Training and reported evaluation
According to the supplied model information, pretraining combines selected GiftEvalPretrain datasets, synthetic data based on KernelSynth, a TSMixup corpus that excludes datasets from the GIFT-Eval evaluation set, and approximately 500,000 synthetic CauKer sequences of length 4,096. These sources indicate that the model was trained for broad time-series behavior rather than for one narrowly defined industry.
IBM reports that PatchTST-FM-r2 ranked second among replicable zero-shot models on GIFT-Eval as of September 8, 2026. The cited comparison reports a geometric-mean CRPS of 0.467 and MASE of 0.6846. These figures are author-reported benchmark results, not guarantees for a particular organization’s data. Benchmark performance can change substantially with sampling frequency, forecast horizon, data quality, scaling choices, and domain shift.
For a production evaluation, users should compare the model against a simple baseline and any existing statistical or machine-learning system. They should measure both point-forecast accuracy and the quality of uncertainty intervals, especially when decisions depend on inventory, staffing, capacity, or risk thresholds.
Pricing, licensing, and deployment
No official per-token, per-request, or hosted inference price is listed for this exact model. The model card states that it is not deployed by a Hugging Face Inference Provider. Consequently, the direct model price is not comparable to a typical managed language-model API tariff.
Users who run it locally or on their own infrastructure must account for the cost of the required compute, storage, operations, and engineering work. The supplied research does not specify a universal hardware requirement or a fixed self-hosting cost, so those values should not be inferred from the parameter count alone.
The model weights and implementation are dual-licensed under the Apache License 2.0 and OpenMDW License 1.0. Users may select either license. Organizations should review the full license terms and any obligations relevant to redistribution, modification, commercial deployment, and internal use before adopting it in a product.
Main strengths and trade-offs
The clearest strength is specialization. PatchTST-FM-r2 is built for numerical forecasting, so its architecture and output design directly address temporal patterns and uncertainty. Its 8,192-step context can accommodate long historical sequences, and the 99-quantile head is useful when planning requires more than a single expected value.
Zero-shot operation is another practical advantage. A team can test the pretrained model on multiple time series without first creating and maintaining a separate training pipeline for each one. This can reduce the initial modeling effort and make it easier to explore forecasting across many related signals.
There are also important trade-offs. At approximately 385 million parameters, this is not a lightweight spreadsheet function or a minimal embedded model. Self-managed deployment may require meaningful compute and operational expertise. The supplied research does not provide a hosted service with a model-specific price, so teams seeking a simple consumption-based API may prefer another option.
The model is also specialized rather than broad. It has no documented tool or function-calling support, no web search, no coding capability, and no general language reasoning interface. Its numerical output should not be confused with a written explanation, a business recommendation, or a general-purpose AI agent.
When to choose PatchTST-FM-r2
PatchTST-FM-r2 is a strong candidate when the primary problem is forecasting regularly sampled numerical data and the team wants a pretrained model that can be evaluated without task-specific training. Suitable examples include:
- Forecasting demand, sales, energy loads, traffic, prices, or operational telemetry.
- Producing prediction ranges for capacity planning and risk-aware decisions.
- Testing one foundation model across many time series before investing in individual models.
- Running inference in a local or self-managed environment where open weights and license choice are important.
- Imputing missing values through the documented Granite TSFM pipeline.
Another option may be more appropriate when the data is irregularly sampled, when the task requires natural-language interaction, or when the application needs tool calls, web access, code generation, image understanding, speech, or a managed API. A conventional statistical model or smaller forecasting model may also be preferable when low latency, minimal infrastructure, or very low operating cost matters more than broad zero-shot capability.
Limitations to check before production use
The model is intended for numerical time series and does not become a general-purpose model simply because it is part of IBM’s broader Granite portfolio. Users should not expect it to explain forecasts in natural language, retrieve external information, call business tools, or process multimedia.
Forecast quality may vary with sampling regularity, missing observations, scaling, forecast horizon, domain shift, and the characteristics of the target series. A long context window does not guarantee that every historical point will be useful, and a 99-quantile output does not automatically mean that the intervals are calibrated.
Before deployment, validate accuracy and uncertainty behavior on representative historical data. Check how the model handles the organization’s missing values, frequency, outliers, seasonal patterns, and forecast horizons. Also review memory and compute requirements, licensing, data handling, monitoring, and fallback behavior. These checks are especially important because the model is open-weight and self-managed rather than presented as a fully operated forecasting service.
Bottom line
Granite TimeSeries PatchTST-FM-r2 is a specialized IBM foundation model for zero-shot numerical time-series forecasting, probabilistic prediction, and imputation. Its defining features are the 8,192-step context, overlapping-patch Conformer-style architecture, approximately 385 million parameters, and 99-quantile prediction head. It is most compelling for teams that need a broad pretrained forecasting model and can manage local or self-hosted inference. It is less suitable for users seeking a cheap hosted endpoint, conversational reasoning, tool use, or multimodal generation.

