Granite Timeseries Ensemble R1

Granite Timeseries Ensemble R1 is an extensible framework for combining quantile forecasts from multiple pretrained time-series models. Its default configuration brings together four complementary, permissively licensed Granite checkpoints: PatchTST-FM-r1, PatchTST-FM-r2, FlowState-r1.1, and TTM-r3.

The ensemble is designed to improve overall forecasting performance and reduce the need to select a single model architecture in advance. It provides two aggregation methods: uniform probability pooling as the simple, recommended default, and adaptive interquartile-range (IQR) weighting as an alternative to evaluate on representative data.

This repository contains reusable ensemble recipes rather than merged model weights. Member checkpoints are downloaded from their own Hugging Face repositories at runtime, and the common forecasting interface can be extended with other compatible models and aggregation methods.

Why ensemble?

Forecasting models have different inductive biases, and no single architecture performs best across every dataset, context length, and forecast horizon. Granite Timeseries Ensemble combines forecasts from complementary model families, reducing reliance on the assumptions and failure modes of any one architecture.

The default Probabilistic Ensemble uses uniform aggregation: it requires no fitted combination weights or historical validation data. This makes the aggregation simple, reproducible, and less exposed to estimation error. Equal-weight combinations are often surprisingly competitive with combinations fitted to historical performance—a result known as the forecast combination puzzle and explained in part by the bias-variance trade-off in estimated weights.

Adding diverse, reasonably accurate members can dilute model-specific errors and improve robustness. Adding weak or highly correlated members is not guaranteed to help, so member selection should still be validated for the intended application. Learned aggregation may improve performance when representative training data are available. These recipes instead prioritize a strong accuracy-simplicity trade-off for zero-shot use.

Ensemble methods

Probabilistic Ensemble

The default linear_pool method gives each member equal influence, pools its quantile values in probability space, and extracts the requested empirical quantiles from the combined values. It does not learn ensemble weights.

Interquartile Range Ensemble

The iqr_weighted method estimates each member's dispersion from its interquartile range. Members with narrower intervals receive greater weight when corresponding quantiles are combined. It does not fit weights from historical forecast errors. Isotonic regression enforces nondecreasing output quantiles.

A narrower interval does not necessarily imply greater accuracy or better calibration. Validate both methods on representative application data.

Granite-only configuration

Both evaluated methods use the same four Granite checkpoints:

Member Family Checkpoint
Granite PatchTST-FM-r1 Transformer ibm-granite/granite-timeseries-patchtst-fm-r1
Granite PatchTST-FM-r2 Conformer ibm-granite/granite-timeseries-patchtst-fm-r2
Granite FlowState-r1.1 State-space model ibm-granite/granite-timeseries-flowstate-r1
Granite TTM-r3 Lightweight mixer ibm-granite/granite-timeseries-ttm-r3

The root config.json is the recommended recipe. The root recipe returns nine quantiles: 0.1, 0.2, …, 0.9, and uses the 0.5 quantile as its point forecast. TTM revisions are selected dynamically from the requested context and forecast horizon.

Installation and usage

pip install "granite-tsfm>=0.3.10"

Minimal example: recommended defaults

import pandas as pd

from tsfm_public.models.ensemble.modeling_ensemble import QuantileEnsembleForecaster


df = pd.read_csv(
    "https://raw.githubusercontent.com/zhouhaoyi/ETDataset/main/ETT-small/ETTh1.csv",
    parse_dates=["date"],
)

ensemble = QuantileEnsembleForecaster.from_pretrained(
    "ibm-granite/granite-timeseries-ensemble-r1",
)

result = ensemble(
    df.tail(1024),
    timestamp_column="date",
    target_columns=["OT"],
    context_length=1024,
    prediction_length=24,
)

print(result.predicted)            # Median forecasts
print(result.predicted_quantiles)  # Quantile forecasts

Advanced example: user-guided configuration

Evaluated alternatives using different aggregation methods and model combinations are available as recipe subfolders in this repository. Pass the desired path through the subfolder argument.

Subfolder Members Method Intended use
recipes/granite-linear Granite-only Probabilistic Ensemble Recommended default
recipes/granite-iqr Granite-only Interquartile Range Ensemble Permissively licensed alternative
recipes/expanded-linear Granite and IBM Research Probabilistic Ensemble Research comparison
recipes/expanded-iqr Granite and IBM Research Interquartile Range Ensemble Research comparison

The optional expanded recipes reference IBM Research checkpoints with separate licensing and usage conditions, including non-commercial or research-use restrictions. Review each member model card before using these recipes.

Load an evaluated alternative with subfolder, or modify the configuration before model checkpoints are loaded:

from tsfm_public.models.ensemble.configuration_ensemble import ProbabilisticEnsembleConfig
from tsfm_public.models.ensemble.modeling_ensemble import QuantileEnsembleForecaster

config = ProbabilisticEnsembleConfig.from_pretrained(
    "ibm-granite/granite-timeseries-ensemble-r1",
    subfolder="recipes/granite-iqr",
)

# Optional advanced changes.
config.iqr_weighted_options = {
    "temperature": 0.8,
    "max_weight": 0.4,
}
config.weights = [0.25, 0.25, 0.25, 0.25]
config.quantile_levels = [0.1, 0.25, 0.5, 0.75, 0.9]
config.validate()

ensemble = QuantileEnsembleForecaster.from_config(config)
config.save_pretrained("./my-ensemble-recipe")

The released IQR recipes use temperature=1.0 and max_weight=null (uncapped). Increasing temperature moves weights towards uniform; lowering max_weight limits how much influence a single member can receive.

Supported member types are patchtst, flowstate, and ttm. A member definition requires forecaster_type and model_checkpoint; FlowState may also specify model_revision, scale_factor, and batch_first. TTM revisions must remain dynamically selected and should not be pinned in the recipe.

GIFT-Eval results

The two Granite-only four-member methods were evaluated over 97 GIFT-Eval dataset configurations. MASE and the GIFT-Eval probabilistic score (mean_weighted_sum_quantile_loss, shown here as CRPS) are normalized against Seasonal Naive and geometrically aggregated across 97 configurations. Lower is better.

The comparison includes reference models whose metadata declares no test-data leakage and available replication code; entries marked fine-tuned are excluded. Models with known non-permissive licenses are also excluded, as are ensembles that depend on a non-permissively licensed member. The figure shows the ten best eligible reference models for each rank metric together with both Granite-only ensemble methods. The displayed peer set is filtered, but the rank values themselves are computed against the full GIFT-Eval leaderboard.

The public GIFT-Eval results report Granite PatchTST-FM-r2 as the strongest individual member of the collection:

Forecast MASE ↓ MASE Rank ↓ CRPS / MWQL ↓ CRPS Rank ↓
Granite PatchTST-FM-r2 0.6846 40.93 0.4672 38.40
Probabilistic Ensemble 0.6830† 37.19 0.4707 39.91
Interquartile Range Ensemble 0.6830† 37.65 0.4700 36.90

† Tied at displayed precision.

MASE and CRPS summarize aggregate point and probabilistic accuracy. MASE Rank and CRPS Rank instead aggregate each model's relative performance across the 97 dataset configurations, spanning multiple domains, frequencies, and forecast horizons. Lower rank values indicate more consistently strong performance across configurations and provide a complementary view of robustness.

Granite-only GIFT-Eval rank comparison

Within the eligible comparison, the Probabilistic and Interquartile Range Ensembles rank first and second respectively on MASE Rank. The Interquartile Range Ensemble ranks first on CRPS Rank, while the Probabilistic Ensemble ranks fourth. This indicates competitive performance across the broad mix of series and forecast settings represented by GIFT-Eval.

Within this filtered comparison, both Granite ensembles place in the top four on both metrics. The Probabilistic Ensemble ranks third on MASE and the Interquartile Range Ensemble fourth, with both moving ahead of Granite PatchTST-FM-r2, the strongest individual member of the collection on aggregate MASE. At the displayed precision, both ensembles score 0.6830. Their full-precision values differ by approximately 0.005%, which should not be interpreted as material.

The Interquartile Range Ensemble ranks third on CRPS, immediately behind Granite PatchTST-FM-r2, while the Probabilistic Ensemble ranks fourth. These results show that combining the four Granite architectures improves their aggregate point-forecast result while retaining competitive probabilistic accuracy.

The Probabilistic Ensemble remains the simpler recommended default; the IQR method offers a competitive alternative when adaptive dispersion-based weighting is desired.

The MASE Rank and CRPS Rank values, together with the displayed comparison positions, reflect standings computed on 28 September 2026. They may change as new methods are added to the leaderboard.

Extending the ensemble

The released recipes use four Granite models, but the ensemble framework is not limited to these members. Additional forecasters can be integrated by implementing the common Forecaster interface and returning predictions at the quantile levels expected by the ensemble.

This makes it possible to:

  • add other Granite or compatible external forecasting models;
  • create recipes for different accuracy, latency, or licensing requirements;
  • compare uniform pooling with alternative aggregation methods; and
  • reuse the ensemble interface without merging or retraining member checkpoints.

New members should produce compatible forecast shapes and quantile levels. Their accuracy, calibration, inference cost, and license compatibility should be evaluated before deployment.

Intended use and limitations

Granite Timeseries Ensemble R1 is intended for zero-shot time-series forecasting and probabilistic forecasting experiments.

  • Multivariate channels are forecast independently; cross-channel information is not used.
  • Loading four checkpoints increases latency and memory use.
  • Member context limits, horizons, preprocessing, and missing-value behavior differ.
  • Neither aggregation method guarantees calibration or robustness under distribution shift.
  • IQR weighting can favor an overconfident member.
  • Application use requires validation on representative data and ongoing monitoring.

License

The ensemble recipe and Granite TSFM implementation are provided under Apache-2.0. Member weights remain governed by their own repositories. The default Granite-only recipe references only Granite checkpoints. The optional expanded recipes reference separately licensed IBM Research checkpoints; review their individual model cards and license terms before use.

References

Downloads last month
90
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train ibm-granite/granite-timeseries-ensemble-r1

Collection including ibm-granite/granite-timeseries-ensemble-r1

Papers for ibm-granite/granite-timeseries-ensemble-r1