On August 31, Google Research announced TimesFM-3, a model that forecasts time series such as sales by combining them with sales trends for related products and with future promotion plans. Its distinguishing feature is that it handles how multiple series move together inside the model itself, producing forecasts without additional training for each use case. According to Google, the 330-million-parameter model was pretrained on more than 1 trillion time points drawn from real and synthetic data. The model can take in more information for demand forecasting, but the released weights carry non-commercial, non-production restrictions, so there is a gap between testing performance and using it in business operations.

AD

What can it forecast when you give it promotion plans?

Google's example is forecasting ice cream sales. Looking only at past sales, a model can extend day-of-week ups and downs into the future. But sales history alone does not reveal which days next month will have discounts. TimesFM-3 accepts such plans as input.

The inputs play different roles. Sales counts for several brands are the "targets" to be forecast together. Past visitor counts are added as "auxiliary variables" to capture their relationship with sales; the technical term is covariates. In addition, promotion schedules, holidays, and weather forecasts available at forecast time are auxiliary variables that can also be supplied for the future period.

In Google's example, a multivariate forecast given the promotion schedule anticipates a sales increase of about 20% on promotion days. The model reads the relationship between past sales and promotions, and its forecast rises on days when a promotion is planned. This 20% is an increase in the forecast. It is neither the share by which sales actually rose at a store that adopted the model, nor a reduction in forecast error.

Being able to input promotion schedules is a step beyond forecasts that ignore plans. However, the example does not prove the causal effect of how much a discount drove sales growth. Nor can a sales-volume forecast be tied directly to the discount rate that maximizes profit; cost and inventory conditions have to be considered separately.

The previous version also supported auxiliary variables; what changed is how they are built in

TimesFM-2.5 also had a way to use auxiliary variables. The official repository's update history states that support via XReg was restored on October 29, 2025. Even when the model itself is designed to forecast a single series, combining a regression step around it allows external information to be brought in.

The change in TimesFM-3 is not first-time support for external variables, but that multiple forecast targets and auxiliary variables are now handled together inside the model.

Comparison item TimesFM-2.5 TimesFM-3
What the model itself forecasts A single time series Jointly forecasts multiple related time series
Handling of auxiliary variables Supported by combining the external regression XReg Natively handles information available only for the past and information that also covers the future
Standard released weights Apache-2.0 Separate license limited to non-commercial, non-production use
Source code Apache-2.0 Apache-2.0

The comparison is based on Google's README, the TimesFM-3 announcement, and the weights license, checked on September 13, 2026 and cross-referenced across the model itself, extensions, weights and code. The old version includes the XReg extension, and this is not a measured comparison of accuracy or speed. Separate contracts via APIs or the cloud are also outside the table.

In short, although the old version could also use external information, this time the mechanism for learning relationships between series is built into the model itself. The core of the design is the difference between forecasting multiple products independently and forecasting them while referring to one another's movements.

Zero-shot multivariate forecasting is not unique to Google. Amazon's official repository lists support for zero-shot forecasting with multivariate data and auxiliary variables in Chronos-2, released on October 20, 2025. TimesFM-3 can be seen as Google putting a new generation of its own model into that competition.

AD

Alternating between time and series to fill in the unknown future

TimesFM-3 groups consecutive 32 time points into "patches" and normalizes the scale of values for each series before processing. This is so that it can handle each series' movements even when, say, visitor counts and product sales differ by orders of magnitude. It uses the same decoder-style Transformer as earlier generations as its foundation.

Inside, it alternates attention along the time direction and along the series direction. In the time direction, it refers to the past of the same series to capture patterns of change. In the series direction, it also refers to other series at the same time point to capture relationships such as that between visitor counts and sales. The design moves back and forth between change over time and the co-movement of multiple variables.

For auxiliary variables that also cover the future, known upcoming patches are attached to the current patch. The unknown future parts, such as sales, are then left blank, while known promotion plans remain visible. Google describes a method that fills these blanks together in a single forward pass. The aim is to curb the accumulation of error and the repeated processing that come with methods that feed earlier forecasts one after another into the next. That said, this cannot be read as an improvement rate in processing time.

The output is also not limited to a single number. For each forecast time point, it returns nine quantiles from the 10th to the 90th percentile. A quantile is a value that divides the forecast distribution at a given proportion. Because it can show an expected value centered on the median along with the range above and below, you can see not only "how many are likely to sell" but also how widely the forecast is spread. However, how well that range holds up on real data needs to be evaluated on the user's side.

Don't read benchmark leadership as your own accuracy

Google reports that TimesFM-3 ranked first among pretrained foundation models on three public benchmarks: GIFT-Eval, FEV-Bench and TIME. The evaluation covers both point forecasting, which looks at how well a representative value hits, and probabilistic forecasting, which evaluates the forecast distribution. The figure in the announcement shows the average rank across multiple tasks.

The comparison includes Chronos-2, the Toto 2.0 series and the earlier TimesFM-2.5. It also separates TimesFM-3 run in univariate mode, where each series is independent, from multivariate mode, which uses information across series and auxiliary variables. According to Google, it matched or beat competitors even in univariate mode, and its average rank improved further in multivariate mode.

This result does not mean it was the most accurate for every forecasting method, or for every store and product. Nor can you calculate from an average rank by what percentage your own error would fall. The performance described here is based on Google's announcement and is not the result of independent replication testing.

The model card also describes the training data breakdown. It says the data includes GiftEvalPretrain with data overlapping FEV-Bench removed, Wikipedia page views through November 2023, Google Trends through the end of 2022, and added synthetic and augmented data. There is language addressing overlap with evaluation data, but that alone does not determine suitability for unseen business data.

In more practical evaluation, "when the input could have been known" matters. For example, if you forecast tomorrow's sales, you need to use the weather forecast for tomorrow that was available at that time. Feeding in the actual weather confirmed afterward gives the model information that would not have been available in operation. Precisely because the model can use future information, comparisons that separate known plans from hindsight are essential.

AD

The released weights can't simply be used for ordering decisions

The standard weights of TimesFM-3 are a public model available from GitHub and Hugging Face, but that does not mean commercial use is free. Google's README distinguishes between Apache-2.0, which applies to the code and to weights up to 2.5, and a non-commercial license, which applies to the 3.0 weights.

The license's definition of non-commercial purposes includes experiments with internal benchmarks and datasets, while conditioning use on not applying the results to commercial decision-making, customer-facing deliverables, or paid products and services. It is not determined simply by whether you run it internally. Using forecasts for actual order quantities or promotion decisions is not treated the same as mere performance evaluation, and commercial or production use requires a separate commercial license from Google.

Cloud availability also needs to be read separately. In the August 31 announcement, Google indicated plans to integrate TimesFM-3 into BigQuery over the coming weeks. The AI.FORECAST documentation checked on September 13 lists TimesFM 2.0 and 2.5, and we could not confirm from it that version 3 is available. The downloadable version's license itself also states that different terms may apply via Google's API.

TimesFM-3 makes it possible to pass the movements of related products and already scheduled promotions to the same model as the sales forecast. If you verify accuracy using data available at your own forecast time, and conditions suited to your intended use are in place, it could become an option for bringing scheduled changes that sales history alone struggles to capture into everyday demand forecasting.