The Art and Science of Forecasting

Chapter 12

The Competition

A forecasting method earns its place only by beating simple benchmarks on the same future observations.

Four demonstrations follow the chapter's benchmark lessons on one synthetic panel of thirty monthly series: scoring every method on the same origins and horizons, what MAE, RMSE and MASE each measure, how the gap to the seasonal naive moves across origins, and why an equal-weight combination is a candidate to test rather than a guaranteed improvement.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Every method faces the same future months

When four forecasts are scored on identical origins and horizons, does one method win at every horizon?

Each line is one method's average error at each horizon over the same thirty series and origins. The naive and drift lines bulge at mid horizons, where the seasonal cycle has moved furthest from the last observed month; they come back down at twelve months, where the last value is again in the same part of the cycle.

Equation: the naive forecast h months after origin T equals y at T

Equation: the seasonal naive forecast equals the value twelve months earlier

Equation: the drift forecast equals y at T plus h times the average change from the first to the last training value

Equation: mean absolute error equals the average of the absolute differences between actual and forecast

Scroll sideways for the whole equation

T is the last training month (the origin), h the horizon in months, y the series and y-hat a forecast. The naive forecast repeats the last value, the seasonal naive repeats the value twelve months earlier, drift extends the line from the first to the last training value, and the ensemble is their equal-weight average. MAE is the mean absolute error over thirty series at the chosen origins.

Predict first. Averaged over all five origins, the seasonal naive alone has the lowest MAE at horizon 4. Will it still be alone at horizon 12?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Every method faces the same future months. Mean absolute error against horizon 1 to 12 for four methods over all five origins, with horizon 4 marked: naive 13.60, seasonal naive 5.44, drift 13.36, ensemble 9.72.
Origins averaged: All five origins, Horizon marked: 4
Constructed data: the chapter notebook's thirty seeded synthetic monthly series (cells 1 and 3), five rolling origins and twelve horizons.

Calculated values

Naive MAE
13.60
Seasonal naive MAE
5.44
Drift MAE
13.36
Ensemble MAE
9.72
Origins averaged
all five origins
Lowest MAE at this horizon
Seasonal naive

At horizon 4, over all five origins and the same thirty series, ensemble MAE minus seasonal-naive MAE is 9.72 - 5.44 = 4.28, so the seasonal naive beats the ensemble here. The lowest MAE belongs to Seasonal naive. Every line is scored on the same future months, so the gaps between lines are differences between methods, not between test sets.

Worked steps

  1. Read each line at horizon 4: naive 13.60, seasonal naive 5.44, drift 13.36, ensemble 9.72.
  2. Ensemble minus seasonal naive: 9.72 - 5.44 = 4.28.
  3. Lowest at this horizon: Seasonal naive.

Use the idea

Before comparing a new method with an old one, fix the origins, horizons and target months once and score every method, baselines included, on exactly those cases.

Where the conclusion applies

Thirty synthetic seasonal series from the notebook (seed 20260930), not M-competition data. Results at one origin rest on thirty series and move from origin to origin, so a single origin is a thin basis for a ranking.

Check your understanding: If one method were scored at origins 72 to 120 and another only at origin 120, what could you conclude from their MAEs?
Nothing about the methods: the test sets differ. At origin 120 alone the seasonal naive's horizon 1 MAE differs from its five-origin value, so the gap could come from the months scored, not the method.

Chapter 12 source: section "Section Three: The Problem of Measuring Accuracy".

Demonstration 2 of 4

Three scores, three different questions

If a method has MASE close to 1, did it tie the seasonal naive on the same future months?

The bars show one score for all four methods over the same 1,800 forecasts. Switching the score changes the units and spacing of the bars; switching the method changes which bar and which calculation are spelled out.

Equation: mean absolute error equals the average of the absolute differences between actual and forecast

Equation: root mean squared error equals the square root of the average squared error

Equation: the training scale s is the mean absolute difference between each training value and the value twelve months earlier

Equation: MASE equals the held out MAE divided by the training scale s

Scroll sideways for the whole equation

y is the actual value and y-hat the forecast. MAE averages absolute errors, RMSE is the square root of the average squared error. s is the training scale: the mean absolute change between each training month and the same month a year earlier, recomputed for every series and origin from months 1 to T only. MASE divides held-out MAE by s.

Predict first. The seasonal naive is the benchmark that sets the MASE scale. Will its own MASE be exactly 1?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Three scores, three different questions. Horizontal bars of MASE for four methods, Seasonal naive highlighted at 1.034: Naive 1.575, Seasonal naive 1.034, Drift 1.585, Ensemble 1.253.
Score: MASE, Method spelled out: Seasonal naive
Constructed data: the chapter notebook's thirty seeded synthetic series; scores replay notebook cell 7 (MAE 7.912, 5.198, 7.970, 6.301; seasonal MASE 1.575, 1.034, 1.585, 1.253).

Calculated values

Method
Seasonal naive
MASE
1.034
Seasonal naive, same metric
1.034
Distance from 1
0.034

Seasonal naive MASE - 1 = 1.034 - 1 = 0.034. Its held-out error is above the training scale. Even the seasonal naive itself scores 1.034, not 1: its held-out error is compared with its error on the training months, so MASE 1 does not mean tied with the seasonal naive on the same future months.

Worked steps

  1. Seasonal naive: MAE 5.198, RMSE 6.474, MASE 1.034.
  2. Seasonal naive MASE - 1 = 1.034 - 1 = 0.034.
  3. Seasonal naive on the same metric: 1.034.

Use the idea

Report MASE as error relative to a training-sample scale, and when the claim is that a method beats a baseline, score the baseline on the same held-out months and compare the two directly.

Where the conclusion applies

The scale uses lag 12 differences of the training months, as the notebook does; the chapter's general definition uses the naive one-step error, with the seasonal naive for seasonal data. A constant training history gives s = 0, and then MASE is undefined, never zero.

Common wrong turn: MASE 1 means a tie with the naive forecast
The chapter says a MASE of 1.0 means held-out MAE equals the training-sample naive scaling error. Here the seasonal naive itself scores 1.034; a tie with a baseline has to be checked by scoring that baseline on the same held-out months.
Check your understanding: A method has held-out MAE 3 on a series whose training scale s is 2. What is its MASE, and what if the training history were constant?
MASE = 3 / 2 = 1.5. With a constant history s = 0, so MASE is undefined: report it as undefined and use an unscaled metric.

Chapter 12 source: section "Section Three: The Problem of Measuring Accuracy".

Demonstration 3 of 4

How large is the gap to the benchmark?

If a method loses to the seasonal naive at every origin, is the size of that loss stable too?

Each origin's point divides one method's MAE by the seasonal naive's MAE on the same series and months. The dashed line at zero is the benchmark itself. Lines that stay on one side keep their rank, while their height shows how much the margin moves.

Equation: the gap G equals 100 times the method MAE divided by the seasonal naive MAE, minus 1

Scroll sideways for the whole equation

G is the gap in percent, MAE with subscript m is the method's mean absolute error at one origin and MAE with subscript sn is the seasonal naive's at the same origin, both over the same thirty series and the chosen horizons. Zero is a tie; negative is better than the benchmark.

Predict first. Over all twelve horizons, does the ensemble beat the seasonal naive at any of the five origins?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How large is the gap to the benchmark?. Gap to the seasonal naive in percent at origins 72 to 120 for naive, drift and ensemble over horizons 1 to 12, Ensemble highlighted: 20.1, 24.0, 15.3, 22.8, 24.0 percent.
Method highlighted: Ensemble, Horizons averaged: Horizons 1 to 12
Constructed data: the chapter notebook's thirty seeded synthetic series and five origins (cell 4); the horizon subsets reuse the same error array.

Calculated values

Gap at origin 72
20.1 percent
Gap at origin 84
24.0 percent
Gap at origin 96
15.3 percent
Gap at origin 108
22.8 percent
Gap at origin 120
24.0 percent
Origins where it beats the seasonal naive
0 of 5
Largest minus smallest gap
8.7 points

Over horizons 1 to 12, at origin 84 the ensemble gap is 100 x (6.600 / 5.323 - 1) = 24.0 percent, and the gap moves by 24.0 - 15.3 = 8.7 points across origins. It loses to the seasonal naive at every origin: a ranking can hold while the size of the gap changes from one origin to the next.

Worked steps

  1. At origin 84: Ensemble MAE 6.600, seasonal naive MAE 5.323.
  2. Gap: 100 x (6.600 / 5.323 - 1) = 24.0 percent.
  3. Smallest gap, at origin 96: 15.3 percent.
  4. Spread: 24.0 - 15.3 = 8.7 points.

Use the idea

Report the gap to the seasonal naive at each origin, not only an average rank, so that a margin that shrinks or grows over time is visible.

Where the conclusion applies

Synthetic seasonal series in which the seasonal cycle dominates, so the seasonal naive is hard to beat by design. Five origins of one panel are a small sample; gaps on real series can behave differently.

Check your understanding: A method has MAE 8 at an origin where the seasonal naive has MAE 10. What is its gap?
100 x (8 / 10 - 1) = -20 percent: 20 percent lower MAE on those matched cases.

Chapter 12 source: section "Section Four: A Systematic Account of What the Competitions Established".

Demonstration 4 of 4

Combination is a candidate, not a guarantee

Does an equal-weight average of forecasts beat each of the forecasts it averages?

Each column holds thirty dots, one per series, for one member or for the combination; the bar is the mean. A combination lands between its members, so whether it helps depends on whether the members fail on different series and months.

Equation: the combined forecast equals the equal weight average of the k member forecasts

Scroll sideways for the whole equation

k is the number of member forecasts, y-hat with superscript i the forecast of member i, and y-hat with superscript c their equal-weight combination. The weights are fixed in advance; held-out errors are never used to choose them.

Predict first. The notebook's ensemble averages naive, seasonal naive and drift. On how many of the 30 series will it beat all three members by MAE?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Combination is a candidate, not a guarantee. Dots of per-series MAE for naive, seasonal naive and drift and their equal-weight combination, means Naive 7.912, Seasonal naive 5.198, Drift 7.970, Combination 6.301.
Forecasts combined: Naive, seasonal naive, drift, Score: MAE
Constructed data: the chapter notebook's thirty seeded synthetic series, five origins and twelve horizons (cells 3 and 5); the two-member combinations average the same notebook forecasts.

Calculated values

Naive MAE
7.912
Seasonal naive MAE
5.198
Drift MAE
7.970
Combination MAE
6.301
Series where it beats every member
4 of 30
Series where it beats its worst member
30 of 30

Members naive, seasonal naive and drift: the average of their MAEs is (7.912 + 5.198 + 7.970) / 3 = 7.027, and the combination scores 6.301, so it loses to its best member on average (seasonal naive, 5.198). It beats every member on 4 of 30 series and its worst member on 30 of 30.

Worked steps

  1. Member MAEs: Naive 7.912, Seasonal naive 5.198, Drift 7.970.
  2. Average of members: (7.912 + 5.198 + 7.970) / 3 = 7.027.
  3. Combination: 6.301; best member Seasonal naive 5.198.
  4. Per series: beats every member 4 of 30, beats the worst 30 of 30.

Use the idea

Treat an equal-weight combination as one more candidate scored on the same held-out months as its members and the baselines; keep it when it wins there, not because it is an average.

Where the conclusion applies

Equal weights, fixed before scoring. In these series the seasonal naive's errors already carry most of what the other two know, which is the case the chapter names where the right weight on the other forecasts would be zero.

What this does not settle

The chapter adds that estimated optimal weights carry their own estimation error, so equal weights often do as well out of sample; this demonstration uses equal weights only and does not test weight estimation.

Chapter 12 source: "the forecast combination puzzle".

Check your understanding: Two forecasts have MAEs 7.912 and 7.970 on a panel. What does the average of these two numbers tell you about the MAE of their equal-weight combination?
(7.912 + 7.970) / 2 = 7.941 is only the average of the two scores. The combination's MAE cannot exceed it, but how far below it falls depends on how often their errors offset: the naive and drift pair here scores 7.936, barely below. It has to be scored.

Chapter 12 source: section "Section Four: A Systematic Account of What the Competitions Established".