Demonstration 1 of 4
Every method faces the same future months
When four forecasts are scored on identical origins and horizons, does one method win at every horizon?
Each line is one method's average error at each horizon over the same thirty series and origins. The naive and drift lines bulge at mid horizons, where the seasonal cycle has moved furthest from the last observed month; they come back down at twelve months, where the last value is again in the same part of the cycle.
Scroll sideways for the whole equation
T is the last training month (the origin), h the horizon in months, y the series and y-hat a forecast. The naive forecast repeats the last value, the seasonal naive repeats the value twelve months earlier, drift extends the line from the first to the last training value, and the ensemble is their equal-weight average. MAE is the mean absolute error over thirty series at the chosen origins.
Predict first. Averaged over all five origins, the seasonal naive alone has the lowest MAE at horizon 4. Will it still be alone at horizon 12?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's thirty seeded synthetic monthly series (cells 1 and 3), five rolling origins and twelve horizons.
Calculated values
- Naive MAE
- 13.60
- Seasonal naive MAE
- 5.44
- Drift MAE
- 13.36
- Ensemble MAE
- 9.72
- Origins averaged
- all five origins
- Lowest MAE at this horizon
- Seasonal naive
At horizon 4, over all five origins and the same thirty series, ensemble MAE minus seasonal-naive MAE is 9.72 - 5.44 = 4.28, so the seasonal naive beats the ensemble here. The lowest MAE belongs to Seasonal naive. Every line is scored on the same future months, so the gaps between lines are differences between methods, not between test sets.
Worked steps
- Read each line at horizon 4: naive 13.60, seasonal naive 5.44, drift 13.36, ensemble 9.72.
- Ensemble minus seasonal naive: 9.72 - 5.44 = 4.28.
- Lowest at this horizon: Seasonal naive.
Use the idea
Before comparing a new method with an old one, fix the origins, horizons and target months once and score every method, baselines included, on exactly those cases.
Where the conclusion applies
Thirty synthetic seasonal series from the notebook (seed 20260930), not M-competition data. Results at one origin rest on thirty series and move from origin to origin, so a single origin is a thin basis for a ranking.
Check your understanding: If one method were scored at origins 72 to 120 and another only at origin 120, what could you conclude from their MAEs?
Chapter 12 source: section "Section Three: The Problem of Measuring Accuracy".
Demonstration 2 of 4
Three scores, three different questions
If a method has MASE close to 1, did it tie the seasonal naive on the same future months?
The bars show one score for all four methods over the same 1,800 forecasts. Switching the score changes the units and spacing of the bars; switching the method changes which bar and which calculation are spelled out.
Scroll sideways for the whole equation
y is the actual value and y-hat the forecast. MAE averages absolute errors, RMSE is the square root of the average squared error. s is the training scale: the mean absolute change between each training month and the same month a year earlier, recomputed for every series and origin from months 1 to T only. MASE divides held-out MAE by s.
Predict first. The seasonal naive is the benchmark that sets the MASE scale. Will its own MASE be exactly 1?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's thirty seeded synthetic series; scores replay notebook cell 7 (MAE 7.912, 5.198, 7.970, 6.301; seasonal MASE 1.575, 1.034, 1.585, 1.253).
Calculated values
- Method
- Seasonal naive
- MASE
- 1.034
- Seasonal naive, same metric
- 1.034
- Distance from 1
- 0.034
Seasonal naive MASE - 1 = 1.034 - 1 = 0.034. Its held-out error is above the training scale. Even the seasonal naive itself scores 1.034, not 1: its held-out error is compared with its error on the training months, so MASE 1 does not mean tied with the seasonal naive on the same future months.
Worked steps
- Seasonal naive: MAE 5.198, RMSE 6.474, MASE 1.034.
- Seasonal naive MASE - 1 = 1.034 - 1 = 0.034.
- Seasonal naive on the same metric: 1.034.
Use the idea
Report MASE as error relative to a training-sample scale, and when the claim is that a method beats a baseline, score the baseline on the same held-out months and compare the two directly.
Where the conclusion applies
The scale uses lag 12 differences of the training months, as the notebook does; the chapter's general definition uses the naive one-step error, with the seasonal naive for seasonal data. A constant training history gives s = 0, and then MASE is undefined, never zero.
Common wrong turn: MASE 1 means a tie with the naive forecast
Check your understanding: A method has held-out MAE 3 on a series whose training scale s is 2. What is its MASE, and what if the training history were constant?
Chapter 12 source: section "Section Three: The Problem of Measuring Accuracy".
Demonstration 3 of 4
How large is the gap to the benchmark?
If a method loses to the seasonal naive at every origin, is the size of that loss stable too?
Each origin's point divides one method's MAE by the seasonal naive's MAE on the same series and months. The dashed line at zero is the benchmark itself. Lines that stay on one side keep their rank, while their height shows how much the margin moves.
Scroll sideways for the whole equation
G is the gap in percent, MAE with subscript m is the method's mean absolute error at one origin and MAE with subscript sn is the seasonal naive's at the same origin, both over the same thirty series and the chosen horizons. Zero is a tie; negative is better than the benchmark.
Predict first. Over all twelve horizons, does the ensemble beat the seasonal naive at any of the five origins?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's thirty seeded synthetic series and five origins (cell 4); the horizon subsets reuse the same error array.
Calculated values
- Gap at origin 72
- 20.1 percent
- Gap at origin 84
- 24.0 percent
- Gap at origin 96
- 15.3 percent
- Gap at origin 108
- 22.8 percent
- Gap at origin 120
- 24.0 percent
- Origins where it beats the seasonal naive
- 0 of 5
- Largest minus smallest gap
- 8.7 points
Over horizons 1 to 12, at origin 84 the ensemble gap is 100 x (6.600 / 5.323 - 1) = 24.0 percent, and the gap moves by 24.0 - 15.3 = 8.7 points across origins. It loses to the seasonal naive at every origin: a ranking can hold while the size of the gap changes from one origin to the next.
Worked steps
- At origin 84: Ensemble MAE 6.600, seasonal naive MAE 5.323.
- Gap: 100 x (6.600 / 5.323 - 1) = 24.0 percent.
- Smallest gap, at origin 96: 15.3 percent.
- Spread: 24.0 - 15.3 = 8.7 points.
Use the idea
Report the gap to the seasonal naive at each origin, not only an average rank, so that a margin that shrinks or grows over time is visible.
Where the conclusion applies
Synthetic seasonal series in which the seasonal cycle dominates, so the seasonal naive is hard to beat by design. Five origins of one panel are a small sample; gaps on real series can behave differently.
Check your understanding: A method has MAE 8 at an origin where the seasonal naive has MAE 10. What is its gap?
Chapter 12 source: section "Section Four: A Systematic Account of What the Competitions Established".
Demonstration 4 of 4
Combination is a candidate, not a guarantee
Does an equal-weight average of forecasts beat each of the forecasts it averages?
Each column holds thirty dots, one per series, for one member or for the combination; the bar is the mean. A combination lands between its members, so whether it helps depends on whether the members fail on different series and months.
Scroll sideways for the whole equation
k is the number of member forecasts, y-hat with superscript i the forecast of member i, and y-hat with superscript c their equal-weight combination. The weights are fixed in advance; held-out errors are never used to choose them.
Predict first. The notebook's ensemble averages naive, seasonal naive and drift. On how many of the 30 series will it beat all three members by MAE?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's thirty seeded synthetic series, five origins and twelve horizons (cells 3 and 5); the two-member combinations average the same notebook forecasts.
Calculated values
- Naive MAE
- 7.912
- Seasonal naive MAE
- 5.198
- Drift MAE
- 7.970
- Combination MAE
- 6.301
- Series where it beats every member
- 4 of 30
- Series where it beats its worst member
- 30 of 30
Members naive, seasonal naive and drift: the average of their MAEs is (7.912 + 5.198 + 7.970) / 3 = 7.027, and the combination scores 6.301, so it loses to its best member on average (seasonal naive, 5.198). It beats every member on 4 of 30 series and its worst member on 30 of 30.
Worked steps
- Member MAEs: Naive 7.912, Seasonal naive 5.198, Drift 7.970.
- Average of members: (7.912 + 5.198 + 7.970) / 3 = 7.027.
- Combination: 6.301; best member Seasonal naive 5.198.
- Per series: beats every member 4 of 30, beats the worst 30 of 30.
Use the idea
Treat an equal-weight combination as one more candidate scored on the same held-out months as its members and the baselines; keep it when it wins there, not because it is an average.
Where the conclusion applies
Equal weights, fixed before scoring. In these series the seasonal naive's errors already carry most of what the other two know, which is the case the chapter names where the right weight on the other forecasts would be zero.
What this does not settle
The chapter adds that estimated optimal weights carry their own estimation error, so equal weights often do as well out of sample; this demonstration uses equal weights only and does not test weight estimation.
Chapter 12 source: "the forecast combination puzzle".
Check your understanding: Two forecasts have MAEs 7.912 and 7.970 on a panel. What does the average of these two numbers tell you about the MAE of their equal-weight combination?
Chapter 12 source: section "Section Four: A Systematic Account of What the Competitions Established".