The Art and Science of Forecasting

Chapter 15

The Foundation

A pretrained model forecasts series it has never seen; whether it helps is settled the old way, against a simple baseline on data it could not have seen.

Four demonstrations follow the chapter: a real zero-shot forecast that cannot anticipate a shift its context does not contain, a benchmark spoiled by test answers in the training material, how patches turn a series into tokens, and a pretrained model scored beside a seasonal naive at the same origins.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

A zero-shot forecast, and the shift it cannot know about

A pretrained model forecasts a series it has never seen. What can it get right, and what can it not?

The model reads only the 112 days before the origin, turns them into tokens and samples 64 future paths. The teal line is the median path and the band spans the middle 80 percent. On the stable series the weekly pattern carries forward; a jump that happens after the context ends is invisible to it.

Equation: mean absolute error equals the average of the absolute differences between y t and its forecast

Equation: coverage equals the share of the n actual values that lie between the 10th and 90th percentile forecasts

Scroll sideways for the whole equation

y is the actual value on day t and y-hat the forecast; n is 14 days. q 0.1 and q 0.9 are the 10th and 90th percentiles of 64 sampled paths, so the band between them is nominally 80 percent. The indicator 1(...) is 1 when the actual value lies inside the band and 0 otherwise.

Predict first. On the level-shift series, forecasting from day 144 (the level jumps at day 145), will the nominal 80 percent band hold most of the next 14 days?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A zero-shot forecast, and the shift it cannot know about. Level shift series around day 144: observed values, the Chronos median with its 10 to 90 percent band, and the seasonal naive line for the next 14 days. 1 of 14 actual values fall inside the band.
Synthetic series: Level shift at day 145, Forecast from day: 144
Constructed data: the chapter notebook's two seeded synthetic series (seed 20260933), forecast by the pinned pretrained Chronos-T5-Tiny checkpoint exactly as the notebook runs it; every coverage matches the notebook's printed output.

Calculated values

Chronos median MAE
12.47
Seasonal naive MAE
12.67
Chronos minus seasonal naive
-0.20
Days inside the 80% band
1 of 14
Band coverage
0.07

Level shift series, forecast from day 144: MAE difference 12.47 - 12.67 = -0.20, so Chronos has the lower error. Band coverage 1 / 14 = 0.07 against a nominal 0.80. The context ends at day 143 and the level jumps at day 145: nothing in the context announces it.

Worked steps

  1. Chronos median MAE over the 14 days: 12.47.
  2. Seasonal naive (repeat the last seven days twice): MAE 12.67.
  3. Difference: 12.47 - 12.67 = -0.20.
  4. Coverage: 1 / 14 = 0.07.

Use the idea

For a new series with no model of its own, a pretrained forecast is available at once. Score it against a seasonal naive at the same origins before relying on it, and expect it to miss what the history does not contain.

Where the conclusion applies

Both series are freshly generated from the notebook seed (20260933), so these exact values were not in the model's training data, though similar weekly shapes may have been. One small checkpoint, eight forecast windows and 64 samples per window: sampled percentiles carry Monte Carlo error and eight cases establish no ranking.

Check your understanding: A band covers 9 of 14 actual days. What is its coverage, and is that above or below the nominal 0.80?
9 / 14 = 0.64, below 0.80. One window cannot establish calibration either way.

Chapter 15 source: section "Section One: The Bet That Paid Off".

Demonstration 2 of 4

A benchmark the model has already seen

What happens to a benchmark score when the test answers were in the training material?

The method looks up the stored context closest to each test context and returns the future stored with it. Every value is independent noise, so with a disjoint library the lookup is no better than chance. Insert a test pair into the library and its nearest stored context is itself, at distance zero.

Equation: j star is the stored context j whose squared distance to the test context x is smallest

Equation: the forecast y hat is the stored future of the nearest context j star

Scroll sideways for the whole equation

x is a test context of 14 values and x j the j-th stored context; j-star is the stored context nearest to x. The forecast y-hat is the 7-value future stored with it. MAE averages the absolute error over the 7 values and the 30 test pairs.

Predict first. If all 30 test pairs are copied into the library, what MAE will the retrieval method score?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A benchmark the model has already seen. Bars of retrieval error for 30 test pairs; the first 0 are inserted in the library and score zero. Overall MAE 1.127 against 1.127 with a disjoint library.
Test pairs copied into the library: 0
Constructed data: the chapter notebook's seeded retrieval toy (100 library pairs, 30 test pairs, all standard normal noise); the notebook inserts none or all 30 pairs, the 10 and 20 cases apply its same retrieval rule to the same draws.

Calculated values

Test pairs inserted in the library
0 of 30
MAE, inserted pairs
undefined (no pair is inserted)
MAE, other pairs
1.127
MAE, all 30 test pairs
1.127
MAE with a disjoint library
1.127

With no test pair in the library, every pair is scored on a neighbour it has never seen, so the benchmark MAE is (0 x 0 + 30 x 1.127) / 30 = 1.127. This is the honest score: the library and the test share no pairs.

Worked steps

  1. Inserted pairs retrieve their own future: error 0 for 0 pairs.
  2. The other 30 pairs average 1.127.
  3. All 30: (0 x 0 + 30 x 1.127) / 30 = 1.127.

Use the idea

Before trusting a score, ask whether the evaluated observations could have been in the training data. Prefer observations generated after a documented training cutoff, or private series never published.

Where the conclusion applies

This is a separate retrieval toy, not Chronos: it shows what overlap does to a score and says nothing about any real model's training corpus. Real leakage is rarely total; the chapter also names indirect leakage through correlated series, which this toy does not model. Draws use the notebook seed (20260933).

Common wrong turn: A zero-shot score on a public benchmark measures generalization
The chapter warns that if the evaluation data were in the training corpus, the zero-shot score was not zero-shot at all: at least partly memorization dressed up as generalization.
Check your understanding: A benchmark has 40 test pairs. 10 are leaked and score 0; the other 30 average 1.20. What MAE does the benchmark report?
(10 x 0 + 30 x 1.20) / 40 = 36 / 40 = 0.90, a quarter below the honest 1.20.

Chapter 15 source: section "Section Two: The Leakage Problem".

Demonstration 3 of 4

Patches as tokens

How does grouping observations into patches change the sequence a transformer must read?

Single-point tokenization makes every value a token, so the sequence is as long as the data. Patching cuts the context into consecutive blocks; each block becomes one token, and the sequence is shorter by the patch length.

Equation: the number of tokens N equals the context length L divided by the patch length P

Scroll sideways for the whole equation

L is the number of values in the context, P the number of consecutive values grouped into one patch, and N the number of tokens the model reads.

Predict first. The notebook gives the model 112 days of context. With patches of 7 values, how many tokens is that?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Patches as tokens. The stable weekly series over the 112 days before day 144, divided into 16 patches of 7 values, alternate patches shaded.
Context length L: 112, Patch length P: 7
Constructed data: the chapter notebook's seeded stable weekly series (seed 20260933) and its 112-day Chronos context; the token counts are arithmetic on those lengths.

Calculated values

Context length L
112
Patch length P
7
Tokens N
16
Each token spans
exactly one weekly cycle

Grouping 112 values into patches of 7: N = 112 / 7 = 16 tokens, each spanning exactly one weekly cycle. The sequence the model attends over is shorter than the data, as the chapter describes.

Worked steps

  1. Context length: 112 values.
  2. Patch length: 7.
  3. Tokens: 112 / 7 = 16.

Use the idea

When comparing models that tokenize differently, count the tokens each one reads from the same context: a longer usable history at the same sequence length is one reason patching was adopted.

Where the conclusion applies

The arithmetic assumes the patch length divides the context length; otherwise a model pads or drops values. It counts tokens only: it does not show how a model embeds a patch or which patch size a model uses for which frequency.

Check your understanding: A daily context of 512 values is cut into patches of 32. How many tokens does the model read?
N = 512 / 32 = 16 tokens.

Chapter 15 source: section "Section Four: The Full Methodology".

Demonstration 4 of 4

The naive baseline is how you find out

Does the pretrained model beat a seasonal naive forecast at the same origins, on this series?

Each bar is one method's mean absolute error over the same 12 months. Three expanding origins choose between methods; the final 12 months are scored once and never used for choosing. The pretrained model is just another candidate: it earns its place by beating the seasonal naive at matched origins, or it does not.

Equation: mean absolute error equals the average of the absolute differences between y t and its forecast

Scroll sideways for the whole equation

y is the actual monthly value and y-hat a forecast; n is the 12 months after each origin. Seasonal naive repeats the value from 12 months earlier; naive repeats the last value; drift extends the line from first to last value; the equal ensemble averages the three.

Predict first. On the Nino 1+2 sea temperatures, does Chronos beat the seasonal naive on the final holdout?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The naive baseline is how you find out. Bars of 12-month MAE at the selection origin 1 on the constructed monthly example: Chronos 2.516, seasonal naive 2.209, and the equal ensemble, naive and drift baselines.
Series: Constructed monthly example, Forecast origin: Selection origin 1
Constructed data: the companion's seeded monthly example for this chapter (144 months, illustrative units). Real data: monthly average sea surface temperature in the Nino 1+2 region, 1950 to 2010, degrees Celsius, from the US National Oceanic and Atmospheric Administration (ERSST.V3B), a revised historical snapshot bundled with the companion. Chronos scores are the notebook's saved workshop results; the baselines are recomputed by the companion engine at the same origins.

Calculated values

Chronos median MAE
2.516
Seasonal naive MAE
2.209
Chronos minus seasonal naive
0.307
Months inside the 10 to 90 percent band
9 of 12
Chronos mean over 3 selection origins
2.481
Seasonal naive mean over 3 selection origins
2.908

On the constructed monthly example, selection origin 1 (12 months after 2017-12): 2.516 - 2.209 = 0.307, so the seasonal naive beats Chronos here. Over the three selection origins Chronos averages (2.516 + 3.498 + 1.428) / 3 = 2.481 against 2.908 for the seasonal naive. One window decides nothing; the comparison at matched origins does.

Worked steps

  1. Chronos median MAE: 2.516.
  2. Seasonal naive MAE at the same origin: 2.209.
  3. Difference: 2.516 - 2.209 = 0.307.
  4. Selection mean: (2.516 + 3.498 + 1.428) / 3 = 2.481.
  5. Band coverage: 9 / 12 = 0.75.

Use the idea

Before adopting a pretrained model, run it beside a seasonal naive at identical forecast origins on your own series, and keep the final holdout untouched until the choice is made.

Where the conclusion applies

Chronos runs zero-shot with no fine-tuning and no covariates; 64 sampled paths per origin. Three selection origins and one holdout are too few to rank methods in general. The sea temperature record has been public for decades, so its absence from the model's pretraining data is not certified.

What this does not settle

A public series may have been in the model's pretraining data, so a strong result on it deserves skepticism; the chapter recommends genuinely held-out or internal series.

Chapter 15 source: "you should be skeptical of results that look unusually strong".

Check your understanding: At three origins a model scores 2.10, 2.60 and 1.90 and the seasonal naive averages 2.30. Which wins on the selection average, and by how much?
(2.10 + 2.60 + 1.90) / 3 = 6.60 / 3 = 2.20, so the model wins by 2.30 - 2.20 = 0.10.

Chapter 15 source: section "Section Five: What This Means in Practice".