The Art and Science of Forecasting

Chapter 13

The Walmart War Room

A global gradient-boosted model wins when the features carry real, timely information, and only an honest backtest shows whether they do.

Four demonstrations work through the chapter's mechanics on a constructed retail panel: why a promotion flag helps on exactly the days a copy of last week fails, how to keep a rolling feature from leaking the answer, how an ablation across rolling origins shows which feature groups earn their place, and how a model's explanation adds up to its own prediction without measuring cause.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Promotion days are where the extra feature pays

A seasonal naive forecast copies last week. What happens to it on a day with a planned promotion?

The black line is the held-out demand, the teal line the chosen forecast, issued one day ahead each morning with yesterday's demand known. The seasonal naive cannot know a promotion is planned; a model given the promotion flag and price can.

Equation: the seasonal naive forecast of y at t equals y at t minus 7, the same weekday last week

Equation: mean absolute error equals the average of the absolute differences between y t and its forecast

Scroll sideways for the whole equation

y is demand of one item at one store on day t, the hat marks a forecast, and the seasonal naive copies the value seven days earlier. MAE is the average absolute miss over n days, in units.

Predict first. For item 3 in the last fold, which forecast has the largest error on promotion days?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Promotion days are where the extra feature pays. Item 3 demand over the last 28 synthetic days with LightGBM with promotion and price; MAE 2.46 overall and 3.09 on the 3 circled promotion days.
Item: 3, Forecast: LightGBM, plus promotion and price
Constructed data: the chapter notebook's synthetic retail panel, last rolling fold (origin day 272), with the notebook's seasonal naive and two LightGBM models refitted exactly as in the notebook.

Calculated values

Item
3
Forecast
With promotion
MAE, all 28 days
2.46
Promotion days
3
MAE on promotion days
3.09
MAE on other days
2.39

For item 3, LightGBM with promotion and price misses by 68.99 units in total over 28 days: 68.99 / 28 = 2.46. On the 3 promotion days the error is 9.27 / 3 = 3.09, larger than 2.39 on the other days.

Worked steps

  1. All days: 68.99 / 28 = 2.46.
  2. Promotion days: 9.27 / 3 = 3.09.
  3. Other days: 59.71 / 25 = 2.39.

Use the idea

Split your error by the days a known driver is active; a forecast that is fine on average can still fail exactly on the days the business cares about.

Where the conclusion applies

Promotions are scheduled in advance and their prices are known at the origin, as the notebook assumes; if a promotion is decided after the forecast is issued, the flag cannot be used. Constructed panel, not M5 data.

Check your understanding: On four promotion days a forecast misses by 12, 18, 9 and 15 units. What is its MAE on those days?
(12 + 18 + 9 + 15) / 4 = 13.5 units.

Chapter 13 source: section "Section Two: What the M5 Revealed".

Demonstration 2 of 4

A rolling feature must stop before the day it predicts

When you build a trailing seven day mean to forecast today, which seven days belong in it?

The dashed line is the forecast origin. The coloured points are the seven days averaged into the feature and the bar is their mean. Shifting by one day before averaging keeps the target out of its own feature.

Equation: the trailing seven day mean m at t equals one seventh of the sum of y over the seven days before t

Scroll sideways for the whole equation

m is the trailing seven day mean used as a feature for day t, and y is the demand of one item at one store on a given day. The sum runs over the seven days before t, never over day t itself.

Predict first. Day 15 is a promotion day with demand far above the days before it. If the window wrongly includes day 15, does the feature move toward the answer or away from it?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A rolling feature must stop before the day it predicts. Item 0 demand around day 15; the safe window covers days 8 to 14 and averages 30.83 units, while day 15 itself is 44.72.
Day forecast: 15, Window: Seven days before (safe)
Constructed data: item 0 of the chapter notebook's synthetic retail panel, with the notebook's shift-then-roll feature and the leaky alternative it warns against.

Calculated values

Day forecast
15
Demand on that day
44.72
Days in the window
8 to 14
Seven day mean used
30.83
Leakage-safe mean (notebook)
30.83
Mean used minus safe mean
0.00

For day 15 the seven values from day 8 to day 14 add to 215.81: 215.81 / 7 = 30.83. The window ends the day before, so the feature was known at the origin.

Worked steps

  1. Window: days 8 to 14.
  2. Sum of the seven values: 215.81.
  3. Mean: 215.81 / 7 = 30.83.
  4. Against the safe mean: 30.83 - 30.83 = 0.00.

Use the idea

Before trusting a validation score, recompute one row's features by hand and confirm that every value used was observed before the forecast origin.

Where the conclusion applies

Item 0 of the notebook's constructed panel (seed 20260931). Yesterday's demand is assumed reported by the morning of the forecast; if reporting lags, even the safe window must stop earlier.

Check your understanding: Demand on days 1 to 3 is 10, 12 and 14, and day 4 turns out to be 20. What is the leakage-safe three day mean for day 4, and what would a window that includes day 4 give?
Safe: (10 + 12 + 14) / 3 = 12. Leaky: (12 + 14 + 20) / 3 = 15.33, which already contains the day 4 answer.

Chapter 13 source: section "Section Three: The Feature Engineering Art".

Demonstration 3 of 4

Ablation: take a feature group away and score again

How do you find out whether a group of features earns its place?

Each bar is the same 28 held-out days scored by a different forecast. Promotion and price are removed together, because a promotion lowers the price and the price alone would reveal it.

Equation: mean absolute error equals the average of the absolute differences between y t and its forecast

Scroll sideways for the whole equation

MAE is the mean absolute error of one-day-ahead forecasts over the 28 days after the origin, for all 16 items, in units. Each fold trains only on days before the origin.

Predict first. Does removing promotion and price cost more accuracy than the lags and calendar add over the seasonal naive?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Ablation: take a feature group away and score again. Bar chart of one-step MAE at origin 272: seasonal naive 6.50, lags and calendar 4.85, plus promotion and price 2.40.
Forecast origin (day): 272
Constructed data: the chapter notebook's synthetic retail panel and its three rolling origins (days 216, 244 and 272); the MAE values reproduce the notebook's printed table.

Calculated values

Origin day
272
Training days
14 to 271
Seasonal naive MAE
6.50
Lags + calendar MAE
4.85
Plus promotion and price MAE
2.40
Lowest MAE
Plus promotion, price

Fitting through day 271 and scoring days 272 to 299: lags and calendar cut the seasonal naive's error by 6.50 - 4.85 = 1.65 units, and adding the planned promotion and price cuts it by a further 4.85 - 2.40 = 2.45. Lowest at this origin: Plus promotion, price. Grey ticks show the other two origins.

Worked steps

  1. Seasonal naive: 6.50.
  2. Lags and calendar: 6.50 - 4.85 = 1.65 better.
  3. Promotion and price: 4.85 - 2.40 = 2.45 better again.

Use the idea

Run your own ablation on held-out folds: a feature group that does not lower error on unseen days is complexity you pay for without return, on that panel.

Where the conclusion applies

Three chronological folds of one constructed panel with fixed LightGBM settings and no tuning. An ablation measures predictive usefulness here, not the causal effect of a promotion.

Check your understanding: A model scores MAE 3.10 with a weather feature and 3.05 without it on the same folds. Does the feature earn its place?
No: 3.05 - 3.10 = -0.05, so the model is slightly better without it on these folds.

Chapter 13 source: section "Section Three: The Feature Engineering Art".

Demonstration 4 of 4

The explanation adds up to the prediction

When a model says why it forecast a number, what exactly do the parts add up to?

Each bar is one feature's share of the gap between this prediction and the model's average output. Positive bars push the forecast up, negative bars down; with the base value they add exactly to the prediction.

Equation: the prediction equals the base value phi zero plus the sum of the feature contributions phi j

Scroll sideways for the whole equation

The hat marks the model's fitted prediction for one row, phi zero is the base value (the model's average output), and each phi j is the contribution of feature j to this prediction, in units of demand.

Predict first. On item 3's first promotion day in the last fold, is the base value plus all nine contributions equal to the prediction, or to the observed demand?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The explanation adds up to the prediction. Horizontal bars of nine feature contributions for item 3 on day 273; promotion flag 13.53, item (-7.44); with base 42.42 they sum to the prediction 46.13.
Promotion day: Day 273
Constructed data: item 3 of the chapter notebook's synthetic retail panel on its three promotion days in the last fold, explained with the notebook's native TreeSHAP calculation.

Calculated values

Day
273
Base value
42.42
Sum of contributions
3.71
Fitted prediction
46.13
Observed demand
46.99
Promotion flag contribution
13.53

Item 3 on promotion day 273: base 42.42 plus contributions totalling 3.71 gives 42.42 + 3.71 = 46.13 units, against 46.99 observed. The promotion flag accounts for 13.53 of the model's output. That is how the model divides its own prediction, not the lift a promotion would cause.

Worked steps

  1. Base value: 42.42.
  2. Promotion flag 13.53, item (-7.44), the other seven features (-2.38).
  3. Total: 13.53 + (-7.44) + (-2.38) = 3.71.
  4. Prediction: 42.42 + 3.71 = 46.13.

Use the idea

Use contributions to audit a model: a large share for a feature that should not matter is a sign of leakage or a data error worth investigating.

Where the conclusion applies

LightGBM's native TreeSHAP for the last model the notebook fits (origin 272, with promotion and price). Correlated features such as promotion and price can share credit; contributions describe the model, not the causal effect of running a promotion.

Check your understanding: A model's base value is 20 and three contributions are 3, -1 and 5. What is the model's output?
20 + 3 + (-1) + 5 = 27.

Chapter 13 source: section "Section Four: The Methods in Full".