Demonstration 1 of 4
A zero-shot forecast, and the shift it cannot know about
A pretrained model forecasts a series it has never seen. What can it get right, and what can it not?
The model reads only the 112 days before the origin, turns them into tokens and samples 64 future paths. The teal line is the median path and the band spans the middle 80 percent. On the stable series the weekly pattern carries forward; a jump that happens after the context ends is invisible to it.
Scroll sideways for the whole equation
y is the actual value on day t and y-hat the forecast; n is 14 days. q 0.1 and q 0.9 are the 10th and 90th percentiles of 64 sampled paths, so the band between them is nominally 80 percent. The indicator 1(...) is 1 when the actual value lies inside the band and 0 otherwise.
Predict first. On the level-shift series, forecasting from day 144 (the level jumps at day 145), will the nominal 80 percent band hold most of the next 14 days?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's two seeded synthetic series (seed 20260933), forecast by the pinned pretrained Chronos-T5-Tiny checkpoint exactly as the notebook runs it; every coverage matches the notebook's printed output.
Calculated values
- Chronos median MAE
- 12.47
- Seasonal naive MAE
- 12.67
- Chronos minus seasonal naive
- -0.20
- Days inside the 80% band
- 1 of 14
- Band coverage
- 0.07
Level shift series, forecast from day 144: MAE difference 12.47 - 12.67 = -0.20, so Chronos has the lower error. Band coverage 1 / 14 = 0.07 against a nominal 0.80. The context ends at day 143 and the level jumps at day 145: nothing in the context announces it.
Worked steps
- Chronos median MAE over the 14 days: 12.47.
- Seasonal naive (repeat the last seven days twice): MAE 12.67.
- Difference: 12.47 - 12.67 = -0.20.
- Coverage: 1 / 14 = 0.07.
Use the idea
For a new series with no model of its own, a pretrained forecast is available at once. Score it against a seasonal naive at the same origins before relying on it, and expect it to miss what the history does not contain.
Where the conclusion applies
Both series are freshly generated from the notebook seed (20260933), so these exact values were not in the model's training data, though similar weekly shapes may have been. One small checkpoint, eight forecast windows and 64 samples per window: sampled percentiles carry Monte Carlo error and eight cases establish no ranking.
Check your understanding: A band covers 9 of 14 actual days. What is its coverage, and is that above or below the nominal 0.80?
Chapter 15 source: section "Section One: The Bet That Paid Off".
Demonstration 2 of 4
A benchmark the model has already seen
What happens to a benchmark score when the test answers were in the training material?
The method looks up the stored context closest to each test context and returns the future stored with it. Every value is independent noise, so with a disjoint library the lookup is no better than chance. Insert a test pair into the library and its nearest stored context is itself, at distance zero.
Scroll sideways for the whole equation
x is a test context of 14 values and x j the j-th stored context; j-star is the stored context nearest to x. The forecast y-hat is the 7-value future stored with it. MAE averages the absolute error over the 7 values and the 30 test pairs.
Predict first. If all 30 test pairs are copied into the library, what MAE will the retrieval method score?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded retrieval toy (100 library pairs, 30 test pairs, all standard normal noise); the notebook inserts none or all 30 pairs, the 10 and 20 cases apply its same retrieval rule to the same draws.
Calculated values
- Test pairs inserted in the library
- 0 of 30
- MAE, inserted pairs
- undefined (no pair is inserted)
- MAE, other pairs
- 1.127
- MAE, all 30 test pairs
- 1.127
- MAE with a disjoint library
- 1.127
With no test pair in the library, every pair is scored on a neighbour it has never seen, so the benchmark MAE is (0 x 0 + 30 x 1.127) / 30 = 1.127. This is the honest score: the library and the test share no pairs.
Worked steps
- Inserted pairs retrieve their own future: error 0 for 0 pairs.
- The other 30 pairs average 1.127.
- All 30: (0 x 0 + 30 x 1.127) / 30 = 1.127.
Use the idea
Before trusting a score, ask whether the evaluated observations could have been in the training data. Prefer observations generated after a documented training cutoff, or private series never published.
Where the conclusion applies
This is a separate retrieval toy, not Chronos: it shows what overlap does to a score and says nothing about any real model's training corpus. Real leakage is rarely total; the chapter also names indirect leakage through correlated series, which this toy does not model. Draws use the notebook seed (20260933).
Common wrong turn: A zero-shot score on a public benchmark measures generalization
Check your understanding: A benchmark has 40 test pairs. 10 are leaked and score 0; the other 30 average 1.20. What MAE does the benchmark report?
Chapter 15 source: section "Section Two: The Leakage Problem".
Demonstration 3 of 4
Patches as tokens
How does grouping observations into patches change the sequence a transformer must read?
Single-point tokenization makes every value a token, so the sequence is as long as the data. Patching cuts the context into consecutive blocks; each block becomes one token, and the sequence is shorter by the patch length.
Scroll sideways for the whole equation
L is the number of values in the context, P the number of consecutive values grouped into one patch, and N the number of tokens the model reads.
Predict first. The notebook gives the model 112 days of context. With patches of 7 values, how many tokens is that?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded stable weekly series (seed 20260933) and its 112-day Chronos context; the token counts are arithmetic on those lengths.
Calculated values
- Context length L
- 112
- Patch length P
- 7
- Tokens N
- 16
- Each token spans
- exactly one weekly cycle
Grouping 112 values into patches of 7: N = 112 / 7 = 16 tokens, each spanning exactly one weekly cycle. The sequence the model attends over is shorter than the data, as the chapter describes.
Worked steps
- Context length: 112 values.
- Patch length: 7.
- Tokens: 112 / 7 = 16.
Use the idea
When comparing models that tokenize differently, count the tokens each one reads from the same context: a longer usable history at the same sequence length is one reason patching was adopted.
Where the conclusion applies
The arithmetic assumes the patch length divides the context length; otherwise a model pads or drops values. It counts tokens only: it does not show how a model embeds a patch or which patch size a model uses for which frequency.
Check your understanding: A daily context of 512 values is cut into patches of 32. How many tokens does the model read?
Chapter 15 source: section "Section Four: The Full Methodology".
Demonstration 4 of 4
The naive baseline is how you find out
Does the pretrained model beat a seasonal naive forecast at the same origins, on this series?
Each bar is one method's mean absolute error over the same 12 months. Three expanding origins choose between methods; the final 12 months are scored once and never used for choosing. The pretrained model is just another candidate: it earns its place by beating the seasonal naive at matched origins, or it does not.
Scroll sideways for the whole equation
y is the actual monthly value and y-hat a forecast; n is the 12 months after each origin. Seasonal naive repeats the value from 12 months earlier; naive repeats the last value; drift extends the line from first to last value; the equal ensemble averages the three.
Predict first. On the Nino 1+2 sea temperatures, does Chronos beat the seasonal naive on the final holdout?
Choose an example
Scroll sideways for the whole figure
Constructed data: the companion's seeded monthly example for this chapter (144 months, illustrative units). Real data: monthly average sea surface temperature in the Nino 1+2 region, 1950 to 2010, degrees Celsius, from the US National Oceanic and Atmospheric Administration (ERSST.V3B), a revised historical snapshot bundled with the companion. Chronos scores are the notebook's saved workshop results; the baselines are recomputed by the companion engine at the same origins.
Calculated values
- Chronos median MAE
- 2.516
- Seasonal naive MAE
- 2.209
- Chronos minus seasonal naive
- 0.307
- Months inside the 10 to 90 percent band
- 9 of 12
- Chronos mean over 3 selection origins
- 2.481
- Seasonal naive mean over 3 selection origins
- 2.908
On the constructed monthly example, selection origin 1 (12 months after 2017-12): 2.516 - 2.209 = 0.307, so the seasonal naive beats Chronos here. Over the three selection origins Chronos averages (2.516 + 3.498 + 1.428) / 3 = 2.481 against 2.908 for the seasonal naive. One window decides nothing; the comparison at matched origins does.
Worked steps
- Chronos median MAE: 2.516.
- Seasonal naive MAE at the same origin: 2.209.
- Difference: 2.516 - 2.209 = 0.307.
- Selection mean: (2.516 + 3.498 + 1.428) / 3 = 2.481.
- Band coverage: 9 / 12 = 0.75.
Use the idea
Before adopting a pretrained model, run it beside a seasonal naive at identical forecast origins on your own series, and keep the final holdout untouched until the choice is made.
Where the conclusion applies
Chronos runs zero-shot with no fine-tuning and no covariates; 64 sampled paths per origin. Three selection origins and one holdout are too few to rank methods in general. The sea temperature record has been public for decades, so its absence from the model's pretraining data is not certified.
What this does not settle
A public series may have been in the model's pretraining data, so a strong result on it deserves skepticism; the chapter recommends genuinely held-out or internal series.
Chapter 15 source: "you should be skeptical of results that look unusually strong".
Check your understanding: At three origins a model scores 2.10, 2.60 and 1.90 and the seasonal naive averages 2.30. Which wins on the selection average, and by how much?
Chapter 15 source: section "Section Five: What This Means in Practice".