Time Series Foundation Models: Forecasting Demand Without Training
Posted on: 9/30/2026 10:54:32 AM
Table of contents
- 1. Every operational decision is a forecasting problem
- 2. Three generations of forecasting: from ARIMA to foundation models
- 3. Inside a Time Series Foundation Model
- 4. The 2026 leap: from univariate to in-context learning with covariates
- 5. The 2026 model landscape
- 6. From forecast to decision: quantiles are what make money
- 7. Evaluate properly: don’t trust the leaderboard, backtest
- 8. Getting hands-on
- 9. When a foundation model is NOT the answer
- 10. Production architecture
- Conclusion
How many crates of milk should the store order tomorrow? How many agents does the contact center need for the Friday evening shift? How many nodes should the cluster add before the 11.11 sale? How much safety stock should the warehouse hold? These sound like different questions, but they are all the same problem: time series forecasting. And in most businesses, this is where AI creates the most direct operational value — every percentage point of forecast error converts straight into money: dead stock tying up capital, stockouts losing revenue, overstaffing or understaffing.
For decades, good forecasting meant having a team of data scientists building a separate model for each series and tuning it every week. In 2026 the game has changed: Time Series Foundation Models (TSFMs) — models pretrained on hundreds of billions to more than a trillion data points — can take any sequence of numbers and return a probabilistic forecast instantly, with no training. On August 31, 2026, Google released TimesFM-3, the first TimesFM model with native multivariate and covariate support. This article dissects how TSFMs work, the 2026 model landscape, how to turn forecasts into operational decisions, and the traps to avoid in production.
1. Every operational decision is a forecasting problem
Look at any operational process and you will find a "future number" sitting in the middle of it:
- Retail & supply chain: demand per SKU × store × day to place orders, allocate inventory and plan promotions.
- Workforce & contact centers: call/ticket volume per 30-minute interval to build shift schedules.
- IT infrastructure: CPU, request traffic and storage for proactive autoscaling and capacity planning.
- Energy: electricity load, solar/wind output, hourly power prices.
- Operational finance: cash flow, weekly revenue, return rates.
What they share is asymmetric error cost. Under-forecasting milk by 100 units (stockout, lost customers) does not cost the same as over-forecasting by 100 units (expired stock written off) — and the ratio differs by industry. So a good forecasting system must not just return "the average"; it must return a probability distribution: "there is a 70% chance demand stays below 1,240 units". This is exactly what the 2026 generation of foundation models does well, and we will come back to it in section 6.
2. Three generations of forecasting: from ARIMA to foundation models
To see why TSFMs are a leap, look back at the two generations before them:
flowchart LR
subgraph G1["Gen 1 — Local statistics"]
A1["ARIMA / ETS / Prophet"] --> A2["1 model PER series
refit on new data"]
end
subgraph G2["Gen 2 — Global ML"]
B1["LightGBM / DeepAR / N-BEATS / TFT"] --> B2["1 model for YOUR whole dataset
needs feature engineering"]
end
subgraph G3["Gen 3 — Foundation models"]
C1["TimesFM / Chronos / TiRex / Toto"] --> C2["Pretrained on 100s of billions of points
zero-shot, in-context learning"]
end
G1 --> G2 --> G3
style A1 fill:#f8f9fa,stroke:#e94560,color:#2c3e50
style B1 fill:#2c3e50,stroke:#fff,color:#fff
style C1 fill:#e94560,stroke:#fff,color:#fff
Generation 1 — local statistics: ARIMA, Exponential Smoothing (ETS), Prophet. Each series gets its own model, learned from its own history. Pros: interpretable, cheap. Cons: no knowledge transfer across series — a newly launched SKU with three weeks of history is nearly impossible to forecast.
Generation 2 — global ML: a single model trained on all of a company’s series. LightGBM with lag/rolling/calendar features dominated the M5 competition (Walmart retail data); DeepAR, N-BEATS and the Temporal Fusion Transformer brought deep learning into the game. Much stronger, but every company still has to build its own training pipeline, feature engineering, tuning and periodic retraining.
Generation 3 — foundation models: just as LLMs learned "language" from the internet, TSFMs learn the "grammar of time series" — trends, seasonality, spikes, nested cycles — from massive multi-domain corpora (energy, traffic, retail, weather, cloud ops) plus synthetic data. Faced with a new series, the model needs no retraining: it reads the history like a "prompt" and infers what comes next.
3. Inside a Time Series Foundation Model
The first question everyone asks: LLMs process words, so how do you "tokenize" a sequence like 1203, 1187, 1455...? There are three technical problems to solve.
3.1. Normalization: putting every series on the same scale
One store’s sales might be 50 units/day; a province’s electricity load might be 3,000,000 kW. A shared model has to "forget" absolute scale to learn shape. TSFMs normalize each series (per-series normalization) before feeding it in and invert it at the output. Chronos-2 uses robust scaling combined with an sinh-1 transform to dampen extreme values.
3.2. Patching: the "word" of a time series is a segment
Instead of making each data point a token (too long, too noisy), modern models slice the series into patches — TimesFM-3, for example, uses 32-step patches. Each patch is embedded as a vector, like a "word" carrying local shape (rising, falling, peaking). The first Chronos generation took a different path: quantizing values into discrete "bins" like an LLM vocabulary; later generations moved to patches because they scale better to long contexts.
3.3. Probabilistic output: the quantile head
Instead of a single number, modern TSFMs output many quantiles at once: TimesFM-3 outputs 9 quantiles (P10 → P90), Chronos-2 outputs 21 (P1 to P99). The model is trained with quantile (pinball) loss, and many models predict the entire horizon in a single forward pass rather than step by step — avoiding the error accumulation of autoregressive generation.
flowchart LR RAW["Raw series
1203, 1187, 1455..."] --> NORM["Per-series
normalization"] NORM --> PATCH["Patching
(e.g. 32 steps)"] PATCH --> EMB["Patch
embeddings"] EMB --> TR["Transformer / xLSTM
temporal attention
+ variate attention"] TR --> QH["Quantile head
P10 ... P90"] QH --> DENORM["De-normalize"] DENORM --> OUT["Probabilistic forecast
for the full horizon"] style RAW fill:#f8f9fa,stroke:#e94560,color:#2c3e50 style TR fill:#e94560,stroke:#fff,color:#fff style OUT fill:#2c3e50,stroke:#fff,color:#fff
Why does "zero-shot" work at all?
Real-world time series share a finite set of "patterns": linear/non-linear trends, daily/weekly/yearly seasonality, holiday effects, spikes followed by mean reversion, structural shifts. Pretrained on diverse enough data (plus synthetic data generated from AR, ETS and random-kernel processes), the model learns to recognize and extrapolate these patterns. Your series, however new, is almost certainly a combination of patterns the model has already seen.
4. The 2026 leap: from univariate to in-context learning with covariates
The first generation of TSFMs had a fatal weakness for operations: they only looked at the history of the series being forecast. But tomorrow’s sales don’t depend only on yesterday’s sales — they depend on the promotion calendar, price, holidays, weather. Electricity demand depends on temperature. Support ticket volume depends on the release schedule. These variables are called covariates (exogenous variables), and they come in three kinds:
| Covariate type | Definition | Operational examples |
|---|---|---|
| Past-only | Values known historically, not known in advance for the future | Store foot traffic, web sessions, measured temperature |
| Known future | Values known for both past and future | Promotion calendar, list price, holidays, weekends, release dates |
| Static / categorical | Attributes that don’t change over time, or labels | Region, store type, product category |
The 2025–2026 models solve this with in-context learning: the target and covariates are fed in together, and the attention mechanism learns how they influence each other — at inference time, with no training.
- Chronos-2 uses group attention: related series (target + covariates, or the variables of a multivariate series) form a group, and information is shared within the group at every patch position. Cost grows linearly with the number of variables instead of quadratically, as it would if everything were flattened into one long sequence.
- TimesFM-3 alternates two kinds of attention: causal temporal attention (along time, within each series) and full variate attention (across series). For known-future covariates, each token is also concatenated with the future patch of that signal (a "lookahead").
- Toto 2.0 and TiRex-2 also go multivariate; Toto 2.0, however, still lists covariate support as "planned for a future release".
flowchart TB
subgraph IN["Inputs in the same context group"]
T["Target: sales
(past)"]
P["Past covariate: foot traffic
(past only)"]
F["Future covariate: promo calendar
(past + future)"]
end
T --> ATT["Temporal attention
alternating with variate attention"]
P --> ATT
F --> ATT
ATT --> Q["Sales forecast quantiles
that account for promo days"]
style ATT fill:#e94560,stroke:#fff,color:#fff
style Q fill:#2c3e50,stroke:#fff,color:#fff
style T fill:#f8f9fa,stroke:#e94560,color:#2c3e50
style P fill:#f8f9fa,stroke:#e94560,color:#2c3e50
style F fill:#f8f9fa,stroke:#e94560,color:#2c3e50
Are covariates actually used? The "perfect covariate" test
A model that "supports covariates" on paper does not necessarily exploit them. An experiment published on September 28, 2026 on the AI Horizon Forecast blog tested exactly this: it generated a 4-year synthetic sales series with weekly/yearly seasonality and 48 irregular promotions (each shifting sales by about 55 units), then gave the models a perfectly accurate promotion flag plus a meaningless "decoy" flag. Event-window MSE results:
| Model | MSE without covariate | MSE with covariate | Improvement |
|---|---|---|---|
| TimesFM-3 | 154.3 | 32.4 | ~4.8x |
| Chronos-2 | 153.7 | 38.8 | ~4.0x |
| TiRex-2 | 152.7 | 55.1 | ~2.8x |
| Toto-2.0 | 153.7 | 153.6 | Essentially unchanged |
Field lesson
Before trusting a model with your promotion calendar, run the perfect covariate test: feed in a variable you know drives the target (e.g. the promo flag itself on data with a clear effect), then compare forecasts with and without it. If the forecast barely moves, the model is ignoring covariates — whatever the documentation says.
5. The 2026 model landscape
| Model | Organization | Size & architecture | Covariates / multivariate | License |
|---|---|---|---|---|
| TimesFM-3 (Aug 2026) | Google Research | 330M, decoder-only, 32-step patches, temporal + variate attention, 9 quantiles | Yes: multiple targets, past & future covariates | Downloaded weights: non-commercial. Commercial use via Google Cloud (BigQuery ML) |
| TimesFM 2.5 (Sep 2025) | Google Research | 200M, context up to 16k points | Univariate; covariates via external regression (XReg) | Apache-2.0 |
| Chronos-2 (Oct 2025) | Amazon | 120M (small: 28M), T5-style encoder-only, group attention, 8,192 context, 21 quantiles | Yes: multivariate, past & future, numeric and categorical | Apache-2.0 |
| TiRex / TiRex-2 | NX-AI | 35M, xLSTM (retains state tracking), very fast | TiRex-2: multivariate + streaming | NXAI Community License (large enterprises pay for commercial use) |
| Moirai 2.0 (Nov 2025) | Salesforce | Decoder-only, quantile loss, multi-token prediction; main release 11.4M params, 30x smaller than Moirai 1.0-Large | Univariate only — multivariate and covariate support deliberately dropped | Open weights — check the model card |
| Toto 2.0 (2026) | Datadog | 4M → 2.5B, u-μP transformer, alternating time/variate attention, 9 quantiles | Multivariate; covariates planned | Open weights — check the model card |
A few important observations when reading this table:
- Two opposing philosophies on covariates. Google, Amazon and NX-AI are betting on multivariate and covariates; Salesforce dropped the feature entirely in Moirai 2.0 after seeing "minimal benefit" on benchmarks. Both are partly right: on general-purpose benchmark data, covariates make little difference; on operational data with promotions, holidays and pricing, they make a multi-fold difference (see the experiment in section 4).
- Small models are still very strong. TiRex with 35M parameters once topped GIFT-Eval; for Moirai 2.0, the larger variants (87M, 305M) actually underperform the 11.4M small model. Toto 2.0, by contrast, reports every size beating the one below it, with no saturation at 2.5B. The "scaling law" for time series is still an open question.
- License is the first filter, not the last. TimesFM-3 leads the benchmarks, but its downloadable weights are for non-commercial, non-production use only; commercial use has to go through Google Cloud services. For self-hosted systems, Chronos-2 (Apache-2.0) is currently the safest and most mature choice.
- Speed: Chronos-2 delivers about 300 series per second on a single A10G GPU (batch 1,024, context 2,048, horizon 64); the 28M variant is nearly twice as fast at only about 1% lower accuracy. Forecasting hundreds of thousands of SKUs every night is entirely feasible on one mid-range GPU.
6. From forecast to decision: quantiles are what make money
The most common mistake when applying AI forecasting is to take the median (P50) and order straight from it. The real value of a TSFM lies in the distribution. The classic newsvendor problem gives a very compact rule:
The critical ratio formula
Let Cu be the cost of being one unit short (lost margin) and Co the cost of one unit too many (write-off, holding). The optimal order quantity is the q* = Cu / (Cu + Co) quantile of the demand distribution.
Example, a bakery chain: each loaf sold earns $0.50 margin, each unsold loaf costs $0.33 to write off → q* = 0.50 / (0.50 + 0.33) ≈ 0.6 → bake exactly the P60 of the forecast, not the P50. For short-life fresh milk (expensive write-offs), q* drops; for high-margin dry goods, q* rises to P80–P90.
The same principle applies throughout operations:
- Contact center staffing: core staff at P50, an on-call pool covering up to P90 of call volume.
- Cloud infrastructure: scale proactively to the P95 of traffic forecast for the next 30–60 minutes, instead of waiting for CPU to cross a threshold and then reacting.
- Safety stock: the P95 − P50 gap over the replenishment lead time is your safety stock, replacing formulas that assume a normal distribution.
But using quantiles for decisions is only safe when they are calibrated — that is, P90 really contains about 90% of actual values. An ICLR 2026 study found that TSFMs are consistently better calibrated than baseline models and are not systematically over- or under-confident — unlike the overconfidence often seen in other deep learning models. Good news, but you still have to measure it on your own data.
flowchart LR FC["TSFM
quantile forecast"] --> POL["Decision policy
q* = Cu / (Cu + Co)"] COST["Shortage / excess cost
per product category"] --> POL POL --> ACT["Action
order, staff, scale"] ACT --> REAL["Actual outcome"] REAL --> MON["Monitoring
P10–P90 coverage, MASE"] MON -->|"calibration drift"| FC style FC fill:#e94560,stroke:#fff,color:#fff style POL fill:#2c3e50,stroke:#fff,color:#fff style COST fill:#f8f9fa,stroke:#e94560,color:#2c3e50
7. Evaluate properly: don’t trust the leaderboard, backtest
Leaderboards like GIFT-Eval (97 tasks from 55 datasets) and fev-bench (100 tasks) are great for building a shortlist, but they have two problems:
- Data leakage: long-standing public datasets have very likely ended up in the models’ pretraining corpora. The TIME benchmark (2026) was created precisely to address this, with 50 entirely fresh datasets and 98 tasks across 8 domains (energy, transport, healthcare, finance, retail, cloud ops...).
- Aggregate rankings hide detail: TIME found that rankings per pattern (strong seasonality, non-stationarity, high noise...) diverge significantly from the overall ranking. The #1 model on average is not necessarily #1 for your data.
The good news from TIME: leading TSFMs (Chronos-2, TimesFM-2.5, TiRex) consistently beat Seasonal Naive, especially on non-stationary and strongly seasonal series. But the only way to know for sure is a rolling-origin backtest on your own data:
- Pick several "cutoff points" in the past (e.g. the last 8–12 weeks); at each one, show the model only the data before it and forecast the next horizon.
- Measure MASE (absolute error scaled by seasonal naive — below 1 means better than naive) for point forecasts, and WQL/CRPS for probabilistic forecasts.
- Always keep baselines: Seasonal Naive, ETS/AutoARIMA and, if you have one, your current LightGBM model. Without a baseline, any number "looks good".
- Check coverage: the share of actuals falling inside the P10–P90 band should be close to 80%.
8. Getting hands-on
8.1. Self-hosted Chronos-2 with covariates (Apache-2.0)
Install with pip install "chronos-forecasting>=2.0". The API works directly with "long" pandas DataFrames: each row is an (id, timestamp) pair with the target and covariate columns. It runs on both GPU and CPU.
import pandas as pd
from chronos import Chronos2Pipeline
pipeline = Chronos2Pipeline.from_pretrained("amazon/chronos-2", device_map="cuda")
# context_df: history -- id column (store_sku), timestamp, target (units_sold),
# plus covariates: is_promo, price, foot_traffic
# future_df : dates to forecast -- ONLY known-future covariates
# (is_promo, price), NO target column
context_df = pd.read_parquet("sales_history.parquet")
future_df = pd.read_parquet("promo_calendar_next_28d.parquet")
pred_df = pipeline.predict_df(
context_df,
future_df=future_df,
prediction_length=28,
quantile_levels=[0.1, 0.5, 0.6, 0.9],
id_column="store_sku",
timestamp_column="timestamp",
target="units_sold",
)
# pred_df has one column per quantile -> feed the 0.6 column to ordering (q* = 0.6)
Columns present in context_df but absent from future_df (like foot_traffic) are treated as past-only covariates; columns present in both are known-future covariates.
8.2. TimesFM-3 via BigQuery (commercially licensed path)
If your data already lives in BigQuery, the AI.FORECAST function lets you call TimesFM straight from SQL. Multivariate forecasting with covariates requires the 'TimesFM 3.0' model (currently in Preview):
SELECT *
FROM AI.FORECAST(
TABLE `retail.daily_sales_with_calendar`,
target_cols => ['units_sold'],
timestamp_col => 'sale_date',
model => 'TimesFM 3.0',
id_cols => ['store_id'],
horizon => 28,
past_covariate_cols => ['foot_traffic'],
future_covariate_cols => ['is_promo', 'is_holiday'],
confidence_level => 0.8
);
Notes on AI.FORECAST
- With
future_covariate_cols, the input data must span the full context window + horizon length — the table must already contain future rows with the promotion and holiday calendar. - For TimesFM 3.0: horizon up to 1,024 steps, context window a multiple of 32 (max 2,048). For TimesFM 2.5 (the default, univariate): horizon up to 10,000, context up to 15,360.
- Each series needs at least 3 data points; the
ai_forecast_statuscolumn reports errors per series instead of failing the whole query.
8.3. Writing your own perfect covariate test
def covariate_uplift(pipeline, ctx, fut, cov_cols, **kw):
"""Compare error with / without covariates on the same backtest window."""
with_cov = pipeline.predict_df(ctx, future_df=fut, **kw)
no_cov = pipeline.predict_df(ctx.drop(columns=cov_cols),
future_df=None, **kw)
return with_cov, no_cov
# If MASE(with_cov) ~= MASE(no_cov) on windows WITH promotions
# -> the model ignores covariates; don't trust it with your promo calendar.
9. When a foundation model is NOT the answer
TSFMs are a big step forward, but not a magic wand. Situations that call for caution:
- Intermittent demand: long-tail SKUs selling 0, 0, 1, 0, 0, 3... Sparse data is hard for every method; a 2025 study on sparse retail data with many gaps found tree models (XGBoost, LightGBM) still beat deep learning architectures such as N-BEATS, N-HiTS and TFT. Backtest this group carefully, and consider aggregating by week or product group.
- Structural breaks: pricing policy changes, new sales channels, pandemics. No model can learn what never happened in the history unless you feed it in as a covariate.
- Hierarchical coherence: store forecasts must add up to regional and national forecasts. TSFMs forecast each series independently; you still need a hierarchical reconciliation step such as MinT.
- Very long horizons: the Moirai 2.0 paper reports performance declining at longer horizons. Evaluate your 3-day and 6-month forecasts separately.
- Non-numeric covariates: product descriptions, images and news are not supported directly by Chronos-2 — you have to encode them as numbers or flags yourself.
- License traps: downloading TimesFM-3 weights to run in production violates its terms; TiRex requires large enterprises to pay for commercial use. Have legal read the model card before you write the first line of code.
10. Production architecture
flowchart TB DW["Data warehouse
sales, operations"] --> PREP["Data preparation
cleaning, gap filling"] CAL["Business calendar
promos, holidays, prices"] --> PREP PREP --> ROUTE{"Series segmentation"} ROUTE -->|"regular, long history"| TSFM["TSFM
Chronos-2 / TimesFM"] ROUTE -->|"sparse, intermittent"| TREE["Specialized models
LightGBM / Croston"] TSFM --> ENS["Combine + hierarchical reconciliation"] TREE --> ENS ENS --> STORE["Forecast quantile store"] STORE --> DEC["Decision systems
ordering, staffing, autoscaling"] STORE --> MON["Monitoring
MASE, coverage, drift"] MON -->|"degradation"| FB["Fallback
Seasonal Naive / ETS"] FB --> STORE style TSFM fill:#e94560,stroke:#fff,color:#fff style DEC fill:#2c3e50,stroke:#fff,color:#fff style ROUTE fill:#f8f9fa,stroke:#e94560,color:#2c3e50
Three design principles worth highlighting:
- Segment (route) before forecasting: not every series suits a TSFM. Classify by sparsity, history length and volatility, then pick the right model for each segment.
- Champion–challenger: run the new model (challenger) alongside the current one (champion) for a few weeks before switching; compare on the same series with the same metrics.
- Always have a fallback: if the GPU fails or metrics suddenly degrade, the system drops back to Seasonal Naive/ETS — worse, but never a blank forecast, because the ordering system downstream cannot wait.
Checklist for putting TSFMs into operations
- Define the decision being supported and its shortage/excess costs → derive the quantile to use.
- Filter models by license and deployment path (self-hosted vs cloud) before comparing accuracy.
- Run a rolling-origin backtest with at least 8 cutoffs, always including Seasonal Naive and your current model as baselines.
- Run the perfect covariate test for each important covariate (promotions, price, holidays).
- Separate sparse/intermittent series and evaluate them on their own.
- Measure quantile-band coverage weekly; alert when it drifts by more than 5–10 percentage points.
- Design the fallback path and hierarchical reconciliation from day one, not after an incident.
Conclusion
Time Series Foundation Models are doing for forecasting what LLMs did for language processing: turning a capability that used to require an expert team and months of work into a function call. Five things to remember:
- Zero-shot is real: pretrained on hundreds of billions to over a trillion points, leading TSFMs consistently beat classic baselines without any training.
- 2026 is the year of covariates: Chronos-2, TimesFM-3 and TiRex-2 learn in context from promotion calendars, prices and weather — but verify it with the perfect covariate test.
- Quantiles make the money: use the critical ratio to pick the right quantile for each decision instead of ordering from the median.
- License is the first filter: the benchmark leader is not necessarily what you are allowed to run in production.
- Backtest on your own data: leaderboards are for shortlisting; the final call belongs to rolling-origin backtests and baselines.
For most businesses, the question is no longer "should we use AI for forecasting?" but "when do we retire the 4-week moving-average spreadsheets?". It may well be the AI investment with the clearest, most measurable ROI in operations today.
References
- Google Research — TimesFM-3: A zero-shot foundation model for multivariate forecasting
- GitHub — google-research/timesfm
- Google Cloud — The AI.FORECAST function (BigQuery ML)
- arXiv — Chronos-2: From Univariate to Universal Forecasting
- Hugging Face — amazon/chronos-2 model card
- Amazon Science — Introducing Chronos-2
- arXiv — TiRex: Zero-Shot Forecasting Across Long and Short Horizons
- arXiv — Moirai 2.0: When Less Is More for Time Series Forecasting
- Datadog — Toto 2.0
- arXiv — It’s TIME: Towards the Next Generation of Time Series Forecasting Benchmarks
- arXiv — Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?
- AI Horizon Forecast — Can a Time-Series Foundation Model See a Promotion Coming?
- arXiv — Comparative Analysis of Modern Machine Learning Models for Retail Sales Forecasting
Disclaimer: The opinions expressed in this blog are solely my own and do not reflect the views or opinions of my employer or any affiliated organizations. The content provided is for informational and educational purposes only and should not be taken as professional advice. While I strive to provide accurate and up-to-date information, I make no warranties or guarantees about the completeness, reliability, or accuracy of the content. Readers are encouraged to verify the information and seek independent advice as needed. I disclaim any liability for decisions or actions taken based on the content of this blog.