Skip to main content
Bibha home

Research digest,

TimesFM-3: zero-shot multivariate time series forecasting in one pass

TimesFM-3 is the newest version of Google Research's time series foundation model and the first in the family pretrained natively for multivariate time series forecasting. It predicts several related series at once, uses past and known future covariates, and returns the whole horizon in one forward pass without task-specific fine-tuning.

By Pupinder Singh, Co-Founder & CTO, Bibha AI Labs

Warm-toned ribbons flowing side by side, linked by glowing threads, then fanning out into translucent green bands on the right
AI-generated illustration: related series moving together, then widening into a range of possible future outcomes.

What is TimesFM-3?

TimesFM-3 is a time series foundation model from Google Research built for zero-shot, multivariate time series forecasting: it predicts several related series at once on datasets it was never trained on. Ayush Jain and Rajat Sen announced it on 31 August 2026, and the post credits Yichen Zhou, Petros Mol, Abhimanyu Das and Samet Oymak as collaborators.

The family began with TimesFM in 2024, described in an ICML 2024 paper as a patched decoder-style attention model pretrained on a large time series corpus. Versions up to TimesFM-2.5, released in September 2025, forecast each series from its own history. The repository notes that 2.5 later regained covariate support through an add-on called XReg, but the model itself stayed univariate.

TimesFM-3 has 330 million parameters and was pretrained on real and synthetic data totalling more than 1 trillion time points. Its model card lists the sources: the GIFT-Eval pretraining collection minus datasets that overlap with fev-bench, Wikipedia page views up to November 2023, top Google Trends queries up to the end of 2022, and synthetic and augmented series.

Why does multivariate time series forecasting matter?

Most business forecasts depend on more than one series. Demand for a product moves with related products, store traffic and events that are already scheduled, so a model that sees a single history leaves useful signal unused. Multivariate time series forecasting lets a model learn those links and carry them into the prediction.

Google's post illustrates the idea with ice cream sales at a retail chain. History alone misses the drivers: cone and syrup sales move with the main product, store visits shift with the season, and promotions, holidays and weather forecasts are known in advance.

TimesFM-3 handles three input roles without per-task tuning.

  • Multiple targets. Several related series, such as competing ice cream brands, are forecast jointly, each with a point forecast and quantiles.
  • Past covariates. Signals observed only up to the present, such as historical foot traffic, inform the forecast without needing future values.
  • Past-future covariates. Signals known ahead of time, such as a promotion calendar or a weather forecast, steer the predicted horizon directly.

How does TimesFM-3 work?

TimesFM-3 keeps the decoder-only transformer design of earlier versions and adds a second direction of attention. Each series is cut into patches of 32 time steps and normalised on its own scale, the patches become tokens on a grid of series by time, and attention alternates between the two axes.

Tokens for target and past-covariate series hold one patch each. Tokens for past-future covariates also append the upcoming patches, a lookahead that lets the model read known future values early. After an input residual block, the tokens pass through a stack that alternates the two attention types for several layers. The model card lists 20 transformer layers, a model width of 1280 and 16 attention heads.

  • Causal temporal attention. Within one series, each token attends only to earlier tokens, so the model cannot see the future of the series it is predicting.
  • Full variate attention. At each time step, a token attends to every other series in the input, so the model can learn cross-series effects, such as a promotion in one series lifting sales in another.

Forecasting the whole horizon in one pass

Earlier TimesFM versions generated forecasts one patch at a time and fed each output back as input, a loop that adds latency and compute and lets errors compound. TimesFM-3 fills the entire horizon in a single forward pass using Contiguous Patch Masking, a training-time masking strategy introduced with the TiRex forecasting model.

In practice the model appends masked placeholder tokens for every future patch. Target and past-covariate series stay masked in the horizon because their future is unknown, while past-future covariates remain visible. The alternating attention layers then fill all masked patches at once, with no iterative loop.

For each target, the model returns 9 quantiles, from the 10th to the 90th percentile, at every horizon step, which gives a direct view of forecast uncertainty. The model card lists a context patch length of 32 and a forecast horizon patch length of 64.

Google's post shows the effect with a planned promotion schedule passed in as a past-future covariate. A univariate forecast repeats the weekly pattern and ignores the promotions, while the multivariate forecast learns the sales lift from history and anticipates a rise of about 20% on each promotion day. The post presents this as an illustration rather than a benchmark result.

How does TimesFM-3 perform on forecasting benchmarks?

The authors report that TimesFM-3 is the top-ranked pretrained foundation model on all three public benchmarks they used, GIFT-Eval, fev-bench and TIME, for both point accuracy and probabilistic forecast quality. Results are given as average rank across tasks, so the claim concerns relative order rather than error size.

The comparison covers recent foundation models with multivariate support, including Chronos-2 and models from the Toto 2.0 family, as well as the earlier TimesFM-2.5. Each benchmark plot shows TimesFM-3 twice. In univariate mode, where every target is forecast alone without covariates, it already matches or beats the other models; with covariates and cross-series information switched on, it ranks first on average for point and probabilistic metrics alike.

The repository's release notes add detail: first overall on fev-bench across 100 real-world tasks, first overall on TIME across 50 datasets and 98 tasks, and first among foundation models on GIFT-Eval. The three benchmarks test different things.

  • GIFT-Eval. Introduced by Aksu and colleagues in 2024, it spans 23 datasets with over 144,000 series and 177 million data points across seven domains and 10 frequencies, plus a separate pretraining set designed to avoid leakage.
  • fev-bench. Introduced by Shchur and colleagues in 2025, it contains 100 forecasting tasks across seven domains, 46 of them with covariates, and reports win rates and skill scores with bootstrapped confidence intervals.
  • TIME. Accepted at ICML 2026, it offers 50 new datasets and 98 tasks built for strict zero-shot evaluation, assembled through a human-in-the-loop process to protect data quality.

What are the limitations and open questions?

The biggest open question is independent verification. We found no technical report or paper for TimesFM-3 at the time of writing, so its architecture and results come from the announcement post, the model card and the code repository, and the benchmark claims rest on average ranks reported by the authors.

Our analysis: the cost of cross-series attention grows with the number of series in each input, so very wide panels, such as thousands of products, may need grouping. The sources do not state a limit on series count; the repository notes only that contexts longer than 15,360 points are truncated to the most recent values.

  • Licence limits. The TimesFM 3.0 weights use a non-commercial licence that rules out commercial or production use of downloaded or self-hosted copies. The repository says such use is allowed through authorised Google Cloud services, including BigQuery ML.
  • Ranks, not error sizes. Average ranks show ordering, not how large the accuracy gaps are. Sizing the benefit for a given use case needs the underlying scores or a local back-test.
  • Illustrative examples. The promotion example in the post demonstrates behaviour; it is not a measured accuracy result.
  • Data overlap. The model card says pretraining excludes GIFT-Eval pretraining datasets that overlap with fev-bench. Public sources such as Wikipedia page views and Google Trends are in the mix, so an evaluation built on those series should check for overlap.

Why it matters for applied and enterprise AI

Our analysis: TimesFM-3 narrows the gap between quick zero-shot baselines and the covariate-aware models that planning teams usually build by hand. If one pretrained model can take related series and known future events directly, teams can test forecasting use cases before investing in per-dataset training.

Google's post notes that time series foundation models are already used in retail, finance, observability, manufacturing, healthcare and the natural sciences. In each, the useful question is whether known future inputs, such as promotions, maintenance windows or weather forecasts, improve accuracy enough to justify the data work they require.

Three practical checks follow. Compare univariate and multivariate modes on a back-test of your own data, since the gain depends on how informative the covariates are. Evaluate the quantiles as well as the point forecast, because many planning decisions turn on the spread. And confirm the licence path before any production use.

Where can you get TimesFM-3?

TimesFM-3 is available as PyTorch weights on Hugging Face under google/timesfm-3.0-pytorch, with code in the google-research/timesfm repository on GitHub. The timesfm package installs from PyPI with either a PyTorch or an MLX extra.

The repository includes univariate and multivariate examples. A multivariate call takes target series arranged by variate and context length, optional past-only covariates over the same context and past-future covariates that extend across the horizon, and returns a forecast and 9 quantiles per target. An MLX backend runs on Apple silicon: on an M4 Max with a context of 512 and a horizon of 64, the repository reports a median latency of 11.1 ms for one series and 666 series per second at a batch size of 32.

For commercial and production use, the repository points to Google Cloud. Its September 2026 update says TimesFM 3.0 has finished rolling out in BigQuery ML; the announcement had suggested trying TimesFM-2.5 through the AI.FORECAST function in BigQuery in the meantime.

Questions and answers

What is the difference between TimesFM-2.5 and TimesFM-3?

TimesFM-2.5 forecasts each series from its own history, with covariates available only through an add-on called XReg. TimesFM-3 is pretrained for multivariate input: it forecasts several targets jointly, uses past and past-future covariates natively and decodes the full horizon in one pass instead of patch by patch. It has 330 million parameters, against 200 million for TimesFM-2.5, and its weights use a non-commercial licence rather than Apache 2.0.

Does TimesFM-3 need fine-tuning?

No. TimesFM-3 is designed for zero-shot forecasting: it is applied directly to a new dataset's history and covariates without task-specific training, and the multivariate support needs no per-task tuning either. The repository does include a LoRA fine-tuning example for TimesFM-2.5, so adapting a TimesFM model to domain data remains possible where the effort is justified.

What are past-future covariates in time series forecasting?

They are inputs whose values are known for both the history and the forecast horizon, such as a promotion calendar, public holidays or a weather forecast. Past covariates, by contrast, are observed only up to the present. TimesFM-3 builds lookahead tokens for past-future covariates and keeps them visible across the horizon, so the forecast can respond to each planned event.

Can TimesFM-3 be used commercially?

Not through the downloaded weights. The repository says the TimesFM 3.0 pretrained weights are restricted to non-commercial, non-production use, while the source code and the weights up to version 2.5 remain under Apache 2.0. Commercial and production use of TimesFM 3.0 is permitted through authorised Google Cloud services such as BigQuery ML, under Google Cloud's terms.

How does TimesFM-3 express forecast uncertainty?

Alongside the point forecast, it outputs 9 quantile levels spanning the 10th to the 90th percentile for every target at every horizon step. Planners can read the spread between the lower and upper quantiles as a range of plausible outcomes, and the authors rank the model on probabilistic metrics as well as point accuracy on all three benchmarks.

References

  1. Das, A., Kong, W., Sen, R., & Zhou, Y. (2024). A decoder-only foundation model for time-series forecasting. International Conference on Machine Learning (ICML 2024). arXiv:2310.10688. https://arxiv.org/abs/2310.10688 (external site)
  2. Google Research. (2026). TimesFM: Time Series Foundation Model [Computer software]. GitHub. https://github.com/google-research/timesfm (external site)
  3. Google. (2026). TimesFM 3.0 (PyTorch) [Model]. Hugging Face. https://huggingface.co/google/timesfm-3.0-pytorch (external site)
  4. Auer, A., Podest, P., Klotz, D., Böck, S., Klambauer, G., & Hochreiter, S. (2025). TiRex: Zero-shot forecasting across long and short horizons with enhanced in-context learning. Advances in Neural Information Processing Systems (NeurIPS 2025). arXiv:2505.23719. https://arxiv.org/abs/2505.23719 (external site)
  5. Aksu, T., Woo, G., Liu, J., Liu, X., Liu, C., Savarese, S., Xiong, C., & Sahoo, D. (2024). GIFT-Eval: A benchmark for general time series forecasting model evaluation. arXiv:2410.10393. https://arxiv.org/abs/2410.10393 (external site)
  6. Shchur, O., Ansari, A. F., Turkmen, C., Stella, L., Erickson, N., Guerron, P., Bohlke-Schneider, M., & Wang, Y. (2025). fev-bench: A realistic benchmark for time series forecasting. arXiv:2509.26468. https://arxiv.org/abs/2509.26468 (external site)
  7. Qiao, Z., Pan, S., Wang, A., Zhukova, V., Liu, Y., Jiang, X., Wen, Q., Long, M., Jin, M., & Liu, C. (2026). It's TIME: Towards the next generation of time series forecasting benchmarks. International Conference on Machine Learning (ICML 2026). arXiv:2602.12147. https://arxiv.org/abs/2602.12147 (external site)

Original article

Jain, A., & Sen, R. (2026, 31 August). TimesFM-3: A zero-shot foundation model for multivariate forecasting. Google Research Blog. https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/ (external site)

This is Bibha's independent summary of published research. Bibha is not affiliated with the authors or Google.

All news and research