--- license: apache-2.0 pipeline_tag: time-series-forecasting library_name: tfc-t0 thumbnail: https://www.theforecastingcompany.com/og/default.png tags: - time-series - forecasting - probabilistic-forecasting - foundation-models - pretrained-models - transformer - multivariate - known-future-covariates - open-weights - covariates - pytorch - mlx - apple-silicon - safetensors - model_hub_mixin - pytorch_model_hub_mixin model-index: - name: t0-alpha results: - task: type: time-series-forecasting name: Time Series Forecasting dataset: name: fev-bench type: autogluon/fev-bench metrics: - name: Skill score type: skill-score value: 42.2 source: name: fev-bench leaderboard url: https://huggingface.co/spaces/autogluon/fev-bench - task: type: time-series-forecasting name: Time Series Forecasting dataset: name: GIFT-Eval type: Salesforce/GiftEval metrics: - name: CRPS type: crps value: 0.4941 - name: MASE type: mase value: 0.7240 source: name: GIFT-Eval leaderboard url: https://huggingface.co/spaces/Salesforce/GIFT-Eval ---

The Forecasting Company

# `t0-alpha`

arXiv

`t0-alpha` is an open-weights time-series forecasting foundation model from [The Forecasting Company](https://theforecastingcompany.com/). `t0` is a transformer-based model that produces probabilistic multi-horizon forecasts and natively operates on multiple covariates. `t0-alpha` is the first public iteration of the model. You can use `t0` on [Retrocast](https://app.retrocast.com/), The Forecasting Company's platform for forecasting on your own data and comparing forecasts across open-weight models. **Model family:** [`t0-alpha` (PyTorch/MLX)](https://huggingface.co/theforecastingcompany/t0-alpha) ยท [ONNX FP16](https://huggingface.co/theforecastingcompany/t0-alpha-onnx-fp16) ยท [ONNX INT8](https://huggingface.co/theforecastingcompany/t0-alpha-onnx-int8) ยท [Collection](https://huggingface.co/collections/theforecastingcompany/t0-alpha-model-family-6a99be18a9e3ab245fda8501) ![t0 forecasting French national electricity demand in Retrocast](https://raw.githubusercontent.com/theforecastingcompany/tfc-t0/main/assets/enedis_with_holidays.webp) _`t0` forecasting French national electricity demand in Retrocast. Data: [Enedis open data](https://data.enedis.fr/)._ ## Model Details - Model name: `t0-alpha` - Model family: `t0` - Developer: [The Forecasting Company](https://theforecastingcompany.com/) - Task: probabilistic time-series forecasting - Architecture: decoder-style patch transformer - Parameters: approximately 102M - License: Apache-2.0 - Weights: https://huggingface.co/theforecastingcompany/t0-alpha - PyTorch runtime: [`tfc-t0`](https://pypi.org/project/tfc-t0/) - MLX runtime: [`tfc-t0-mlx`](https://pypi.org/project/tfc-t0-mlx/) - Managed API: https://docs.retrocast.com/documentation/t0-alpha `t0-alpha` is an alpha release intended for research, experimentation, and applied forecasting evaluation. ## Intended Use `t0-alpha` is intended for probabilistic time-series forecasting. It can be used for univariate and multivariate forecasting, forecasting with historical or known-future covariates and multi-horizon forecasting. Known-future covariates can include calendar features, planned events, holidays, promotions, weather forecasts, or other external signals available over the forecast horizon. Forecasts should be treated as probabilistic estimates, not guarantees. ## ๐Ÿ“ˆ Forecasting With Covariates `t0` leverages covariate information, in the past and future when available, to improve its forecast. | Without covariates | With covariates | | ----------------------------------------------------------------- | ----------------------------------------------------------- | | ![t0 forecast without covariates](https://raw.githubusercontent.com/theforecastingcompany/tfc-t0/main/assets/medicam_without_cov.webp) | ![t0 forecast with covariates](https://raw.githubusercontent.com/theforecastingcompany/tfc-t0/main/assets/medicam_with_cov.webp) | _Data: [Medic'AM](https://www.assurance-maladie.ameli.fr/etudes-et-donnees/medicaments-classe-atc-medicam), monthly drug reimbursements from the French national health insurance._ The [Quickstart](#quickstart) below shows the API for both a plain univariate forecast and a multivariate forecast that conditions on historical and known-future covariates. ## Installation Choose a runtime for the same original `t0-alpha` checkpoint: | Runtime | Best for | Install | | --- | --- | --- | | PyTorch | Broad hardware support and the PyTorch ecosystem | `pip install tfc-t0` | | MLX | Local, inference-only use on Apple silicon | `pip install tfc-t0-mlx` | | ONNX FP16 | Accelerator-oriented local and edge deployments | [`t0-alpha-onnx-fp16`](https://huggingface.co/theforecastingcompany/t0-alpha-onnx-fp16) | | ONNX INT8 | CPU and in-browser inference | [`t0-alpha-onnx-int8`](https://huggingface.co/theforecastingcompany/t0-alpha-onnx-int8) | | Managed API | Hosted inference without local weights | [`theforecastingcompany` SDK](https://pypi.org/project/theforecastingcompany/) | ### PyTorch ```bash pip install tfc-t0 ``` Requirements: - Python `>=3.10` - PyTorch `>=2.4` Optional extras: ```bash pip install "tfc-t0[evaluation]" pip install "tfc-t0[plot]" ``` ### MLX on Apple silicon ```bash pip install tfc-t0-mlx ``` The MLX package uses the same model repository, loads its safetensors directly and does not install PyTorch. ## ๐Ÿš€ Quickstart These weights are public โ€” no authentication is needed to download them. The simplest path is a univariate forecast through `predict`: ```python import torch from t0 import T0Forecaster model = T0Forecaster.from_pretrained("theforecastingcompany/t0-alpha").eval() context = torch.randn(4, 512) # 4 series, 512 past timesteps out = model.predict(context, horizon=64, quantile_levels=[0.1, 0.5, 0.9]) out.quantiles # (4, 64, 3) out.median # (4, 64) ``` `predict` accepts PyTorch tensors and NumPy arrays. ### MLX Quickstart The MLX runtime deliberately follows the same forecasting interface: ```python import numpy as np from t0_mlx import T0Forecaster model = T0Forecaster.from_pretrained("theforecastingcompany/t0-alpha").eval() context = np.random.randn(4, 512).astype(np.float32) out = model.predict(context, horizon=64, quantile_levels=[0.1, 0.5, 0.9]) out.quantiles.shape # (4, 64, 3) out.median.shape # (4, 64) ``` See [T0 for MLX](https://github.com/theforecastingcompany/tfc-t0/tree/main/mlx) for feature coverage, compilation guidance and reproducible Apple-silicon benchmarks. ### Forecasting With Covariates Anything known over the past goes in `context`. Alongside the target, extra variates attend to it and are forecast together. Anything known over the future, such as calendar features, planned promotions, or weather forecasts, goes in `future_covariates`, shaped `[B, F, context + horizon]`. The model conditions on it but does not forecast it. ```python import torch from t0 import T0Forecaster model = T0Forecaster.from_pretrained("theforecastingcompany/t0-alpha").eval() context = torch.randn(2, 512) # 2 series, 512 past timesteps future_covariates = torch.randn(2, 3, 512 + 64) # 3 covariates known over context + horizon out = model.predict( context, horizon=64, quantile_levels=[0.1, 0.5, 0.9], future_covariates=future_covariates, ) out.quantiles # (2, 64, 3) out.median # (2, 64) ``` ### Batched Inference ```python import numpy as np from t0 import T0Forecaster, batch_series model = T0Forecaster.from_pretrained("theforecastingcompany/t0-alpha").eval() daily = np.random.randn(180) # one series, 180 past timesteps store = np.random.randn(2, 96) # one series of 2 variates, 96 past timesteps hourly = np.random.randn(1024) # one series, 1024 past timesteps context, mask, group_ids = batch_series([daily, store, hourly]) context.shape # (4, 1024) โ€” variates stacked, right-aligned to the longest series group_ids # [0, 1, 1, 2] โ€” `store`'s two variates are forecast jointly out = model.predict(context, horizon=24, quantile_levels=[0.1, 0.5, 0.9], mask=mask, group_ids=group_ids) out.quantiles # (4, 24, 3) out.median[0] # the 24-step median forecast for `daily` ``` Integrations that prepare complete T0 inputs, including known-future covariates, can batch the native representation directly: ```python from t0 import TimeSeries first = TimeSeries.from_array(context_1, future_covariates_1) second = TimeSeries.from_array(context_2, future_covariates_2) batch = TimeSeries.batch([first, second]) out = model.predict( batch, horizon=64, context_length=max(context_1.shape[-1], context_2.shape[-1]), ) ``` Here each context includes its batch axis, for example `[1, V, T]`, and each known-future input is `[1, F, T + horizon]`. The output is ordered by the flattened target rows in `batch`. ### Converting your data to `TimeSeries` `TimeSeries` is the model's native input. It holds target rows, known-future covariate rows, a mask and group ids, all on one width. `predict` builds one for you from a raw array. You only need to construct one yourself to batch inputs of different widths, or to call `forward` directly. ```python from t0 import TimeSeries # context only, with `horizon` marking the region to predict model_input = TimeSeries.from_array(context, horizon=24) # context: [B, V, T] # with known-future covariates, whose width sets the horizon model_input = TimeSeries.from_array(context, future_covariates) # covariates: [B, F, T + 24] out = model.predict(model_input, horizon=24, quantile_levels=[0.1, 0.5, 0.9]) ``` `predict` infers `context_length` from where the forecast region starts. Pass it explicitly when batching series of different widths. `forward` takes the same `TimeSeries` and runs a single differentiable pass over it, with no rollout. That is the entry point for fine-tuning. **For efficient inference at scale, look at [Retrocast](https://app.retrocast.com/).** ## Input Contract - `context` may be shaped `(B, T)` for batched univariate forecasting. - `context` may also be shaped `(T,)`, which is promoted to a single-row batch. - `context` may be shaped `(B, V, T)` for multiple target variates. - `future_covariates`, when provided, should be shaped `(B, F, context + horizon)`. - `mask`, when provided, holds `MaskType` values shaped like `context`: `MISSING` for an absent observation, `PAD` for a cell that only widens a shorter series out to the batch's width. - NaN in `context` is read as an absent observation. Padding is the case NaN cannot express, so a batch of unequal-length series needs a `mask` (or `batch_series`) to declare it. - Patches made entirely of `PAD` stay out of attention. - `group_ids`, when provided, holds one id per row of the context. Rows sharing an id are variates of one series and are forecast jointly. - `group_ids` cannot be combined with `future_covariates`, which are addressed per sample. - NaN in `future_covariates` is treated as missing. - `horizon` must be at least 1. - Requested quantiles must be non-empty, sorted ascending, unique, and in `(0, 1)`. - The model was trained to emit quantiles `0.1`, `0.25`, `0.5`, `0.75`, and `0.9`. - Requested levels between the trained ones are interpolated; levels beyond them follow exponential tails pinned through the outermost trained levels. - Horizons up to 1024 timesteps are decoded in one forward pass. - Longer horizons use autoregressive rollout. - Returned forecasts are finite `float32` tensors on the model's device. ## ๐Ÿ—๏ธ Architecture `t0` is a decoder-style patch transformer. It encodes each patch from values, within-patch time index, and validity mask. The transformer alternates causal time-axis self-attention with variate-axis group self-attention. Time attention uses time-aware rotary embeddings. Variate attention lets variates in the same sample attend to one another. The stack uses pre-norm RMSNorm blocks, SwiGLU feed-forward layers, and a quantile head. At inference, target and historical variates are normalized with causal running statistics. Future covariates use per-row global statistics. | Field | Value | | --- | --- | | Parameters | approximately 102M | | Layers | 24 | | Layer pattern | 2 time-attention layers, then 1 group-attention layer | | Time attention layers | 16 | | Group attention layers | 8 | | Embedding dim | 512 | | Feedforward dim | 2048 | | Attention heads | 8 | | Patch size | 32 | | Dropout | 0.1 | | Scaler | causal mean/std with `arcsinh` transform | | Native quantile levels | 0.1, 0.25, 0.5, 0.75, 0.9 | ## Evaluation `t0-alpha` is reported on the [GIFT-Eval leaderboard](https://huggingface.co/spaces/Salesforce/GIFT-Eval) and the [fev-bench leaderboard](https://huggingface.co/spaces/autogluon/fev-bench). | Benchmark | Metric | Value | | --- | --- | ---: | | GIFT-Eval | CRPS | 0.4941 | | GIFT-Eval | MASE | 0.7240 | | fev-bench | Skill score | 42.2 | Users should also evaluate `t0-alpha` on their own historical backtests. Useful checks include quantile loss, CRPS, MASE, empirical quantile coverage, calibration, and breakdowns by frequency, horizon, domain, history length, and covariate availability. ## ๐Ÿงฐ Public API - `T0Forecaster`: the model itself. - `Forecast`: the object returned by the model. - `T0Config`: the configuration of the model. - `MaskType`: the reason a time step is masked out. - `VariateType`: whether a row is a target, a historical covariate or a known-future covariate. - `batch_series`: utility to batch time series of potentially different lengths. - `TimeSeries.from_array` / `TimeSeries.batch`: build the model's native input, including known-future covariates and an explicit forecast `horizon`. `predict` accepts either a `TimeSeries` or a raw context array. ## ๐Ÿงฌ Lineage and Attributions `t0` builds on ideas from open-source forecasting models. We gratefully acknowledge: - **Toto** by Datadog ([repo](https://github.com/DataDog/toto)) and **Chronos-2** by Amazon ([repo](https://github.com/amazon-science/chronos-forecasting)) for factorizing attention in the time and variates dimension. - **TiRex** by NXAI ([repo](https://github.com/NX-AI/tirex)) for contiguous patch masking. Code-level attributions are listed in [`NOTICE`](https://huggingface.co/theforecastingcompany/t0-alpha/blob/main/NOTICE), all under Apache-2.0. ## Environmental Impact Training compute and carbon emissions are not currently reported. ## ๐Ÿ“š Citation `t0` is described in [t0: A Time-Series Foundation Model for Forecasting with Context](https://arxiv.org/abs/2609.24559). If our model is useful, please cite: ```bibtex @article{meyer2026t0, title = {$t_0$: A Time-Series Foundation Model for Forecasting with Context}, author = {Meyer, Lucas and Sole, Claudio and Xiang, Huikan and Li, Nicolas and Franceschino, Lucas and Quera-Bofarull, Arnau and Scholl, Maarten P. and Fainberg, Joachim and N{\'e}giar, Geoffrey}, journal = {arXiv preprint arXiv:2609.24559}, year = {2026}, url = {https://arxiv.org/abs/2609.24559}, } ``` ## โš–๏ธ License Apache-2.0. See [`LICENSE`](https://huggingface.co/theforecastingcompany/t0-alpha/blob/main/LICENSE) and [`NOTICE`](https://huggingface.co/theforecastingcompany/t0-alpha/blob/main/NOTICE). ## Contact For issues and bug reports, use the tracker for the relevant runtime: - PyTorch: https://github.com/theforecastingcompany/tfc-t0/issues - MLX: https://github.com/theforecastingcompany/tfc-t0/issues