The Art and Science of Forecasting

Chapter 14

The Globalizer

One model trained across many series can forecast a series with almost no history, and its forecast should be a distribution, checked like one.

Four demonstrations follow the chapter: a small network trained on 32 synthetic series forecasting eight it never saw, the chapter's safety stock arithmetic when the variability is wrong, how sampled paths turn a next-day distribution into a two-week band, and whether a nominal 80 percent band covers 80 percent of outcomes.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

A model trained on other series forecasts one it never saw

Can a network trained on 32 series forecast a 33rd it never trained on, from only 14 of its values?

Each pair of bars is one entity the network never trained on: its error on the left, the seasonal naive error on the right. Both see only the 14 values before each of three forecast origins. The network brings the weekly shape it learned from 32 other series; the naive forecast copies one noisy week.

Equation: mean absolute error equals the average of the absolute differences between y t and its forecast

Scroll sideways for the whole equation

y is the observed demand index on a forecast day and y-hat a forecast of it. MAE is the mean absolute error over the days scored, in index units. The global forecast is the median of the network's sampled paths; the local forecast repeats the entity's last observed week.

Predict first. Over all eight held-out entities and all 14 days, will the global network beat the local seasonal naive?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A model trained on other series forecasts one it never saw. Bars of mean absolute error for the eight held-out entities, global network against local seasonal naive, days 1 to 14, with all eight held-out entities highlighted: global 2.127, local 3.081.
Entities scored: All eight held out, Days ahead: Days 1 to 14
Constructed data: the chapter notebook's 40 seeded synthetic demand indices (seed 20260932), 32 for training and 8 held out, scored at origins 140, 154 and 166.

Calculated values

Forecasts scored
336
Global network MAE
2.127
Local seasonal naive MAE
3.081
Local minus global
0.954
Of all 8 entities, global wins on
8

Over all eight held-out entities, days 1 to 14, local MAE minus global MAE is 3.081 - 2.127 = 0.954, so the global model has the lower error here. Neither model saw these entities during training; each forecast starts from the same 14 recent values. The network's advantage is borrowed from the 32 training entities, which share the weekly shape by construction.

Worked steps

  1. Forecasts scored: 336 (8 entities x 3 origins x 14 days).
  2. Global network MAE: 2.127. Local seasonal naive MAE: 3.081.
  3. Difference: 3.081 - 2.127 = 0.954.
  4. The global model wins on 8 of the 8 entities over days 1 to 14.

Use the idea

When a new product has a few weeks of history, compare a model trained across related products with a simple local forecast on products held out of training, not on the products the model learned from.

Where the conclusion applies

The 40 series are synthetic and share one weekly pattern by design, which favours pooling. With unrelated series, or a held-out entity with a different cycle, the global model's advantage can vanish. This small network is not DeepAR: it has no recurrence and no item embeddings.

Check your understanding: A global model scores MAE 1.505 on an entity and a local forecast scores 1.949. By how much does the global model reduce the error, and does that prove pooling helps for every new product?
1.949 - 1.505 = 0.444 index units lower. No: it shows transfer to series that share the training series' structure; an unrelated product can reverse it.

Chapter 14 source: section "The Catalog and the Cold Start".

Demonstration 2 of 4

The same point forecast, a different safety stock

If two products both average 100 units a week, can they share one safety stock?

The curve is the true weekly demand; the solid line is the stock, set at the mean plus 1.645 of the planned standard deviations. The shaded tail beyond the line is the share of weeks that run out. When the planned and true standard deviations differ, the stock and the target come apart.

Equation: the stock S equals the mean demand mu plus z standard deviations sigma

Scroll sideways for the whole equation

S is the stock held, mu the mean weekly demand (100 units), sigma the standard deviation of weekly demand in units and z the number of standard deviations the target needs: 1.645 for a 95 percent service level.

Predict first. Stock is set for a standard deviation of 10, but demand really varies with standard deviation 30. Roughly what service level is reached?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The same point forecast, a different safety stock. Normal demand curve with mean 100 and standard deviation 30; stock line at 116.45 with the stockout tail shaded; the stock reaches 70.8 percent of weeks.
Standard deviation used to set the stock: 10, True standard deviation of demand: 30
Constructed data: the chapter's own inventory example (mean 100, standard deviation 10 or 30, 95 percent service level), with standard deviation 20 added.

Calculated values

Safety stock set
16.45
Stock held
116.45
Standard deviations covered
0.548
Service level reached
70.8 percent
Stock needed for 95 percent
149.35

Safety stock 1.645 x 10 = 16.45, so the stock is 116.45. Against true variability 30 it covers (116.45 - 100) / 30 = 0.548 standard deviations, which the normal curve turns into 70.8 percent: the stock falls short of the 95 percent target. The shaded tail is the share of weeks that run out.

Worked steps

  1. Safety stock: 1.645 x 10 = 16.45.
  2. Stock held: 100 + 16.45 = 116.45.
  3. Standard deviations covered: (116.45 - 100) / 30 = 0.548.
  4. Normal curve below 0.548 standard deviations: 70.8 percent of weeks without a stockout.

Use the idea

Before reusing one safety stock buffer across products, check that each product's own demand variability is the one the buffer was computed for; the point forecast alone cannot tell you.

Where the conclusion applies

Demand is treated as normal with a known mean, as in the chapter's calculation. Count demand that is skewed or intermittent breaks the normal approximation, which is why the chapter turns to the negative binomial.

Common wrong turn: A fixed buffer is safe for every product
The chapter warns that applying the same safety stock buffer to every product ignores differences in variability: with standard deviation 30, the 116.45 units set for standard deviation 10 reach about 71 percent of weeks, not 95.
Check your understanding: Mean weekly demand is 100 with standard deviation 20. What stock gives a 95 percent service level?
1.645 x 20 = 32.9 units of safety stock, so 100 + 32.9 = 132.9 units.

Chapter 14 source: section "Section Three: The Probabilistic Imperative".

Demonstration 3 of 4

A distribution, sampled one day at a time

How does a network that predicts only the next day produce a band for the next fourteen?

The network sees only the last 14 values, divided by their mean absolute level, and predicts a mean and a spread for the next day. Drawing from that distribution, appending the draw and repeating gives one path; 160 paths give the median and the band. No value after the origin enters any path.

Equation: the sampled value y t equals the context scale s t times 1 plus the network mean m t plus the network spread d t times a standard normal draw epsilon t

Scroll sideways for the whole equation

s is the mean absolute level of the 14 most recent values, m and d the network's mean and spread for the day on that relative scale, and epsilon a standard normal draw. Each drawn y is added to the history before the next day is drawn.

Predict first. For entity 32 at origin 154, will the 80 percent band be wider on day 14 than on day 1?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A distribution, sampled one day at a time. Entity 32: 14 observed values, then 14 forecast days from origin 154 with ten sampled paths, the median and an 80 percent band from 7.31 wide on day 1 to 8.29 on day 14; the actual values lie inside on 13 days.
Held-out entity: 32, Forecast origin (day): 154
Constructed data: the chapter notebook's seeded synthetic entities, never used for weight updates, with the notebook's 160 recursive sample paths per origin (seed 20260932).

Calculated values

Context scale s
60.31
Network mean m, day 1
-0.003
Network spread d, day 1
0.050
Day 1 mean
60.13
Day 1 standard deviation
3.02
80 percent band width, day 1
7.31
80 percent band width, day 14
8.29
Days inside the band
13 of 14

Entity 32 at origin 154: the 14 context values have mean absolute level s = 60.31. The network returns m = -0.003 and d = 0.050 for day 1, so the mean is 60.31 x (1 + (-0.003)) = 60.13 and the standard deviation 60.31 x 0.050 = 3.02. Each sampled value is fed back as a lag for the next day. Here the band on day 14 is wider than on day 1: 8.29 against 7.31. The actual values fell inside it on 13 of 14 days.

Worked steps

  1. Scale from the context only: s = 60.31.
  2. Day 1 mean: 60.31 x (1 + (-0.003)) = 60.13.
  3. Day 1 standard deviation: 60.31 x 0.050 = 3.02.
  4. Draw 160 values, feed each back as a lag, and repeat for days 2 to 14.
  5. Band width: 7.31 on day 1, 8.29 on day 14.

Use the idea

To read a probabilistic forecast, ask how its band was built: sampled paths carry uncertainty forward, and a band from 10th to 90th percentiles of each day is a statement about each day, not about a whole path.

Where the conclusion applies

The model is a small Gaussian network, not DeepAR's recurrent network with a negative binomial output; a Gaussian can put probability on negative demand. Each band uses 160 sampled paths (the book's figure uses 400 for entity 32).

Check your understanding: The 14 context values have mean absolute level 50, and the network predicts m = 0.1 and d = 0.2. What are the mean and standard deviation of the next value?
Mean 50 x (1 + 0.1) = 55; standard deviation 50 x 0.2 = 10.

Chapter 14 source: section "Section Four: The Neural Forecasting Landscape".

Demonstration 4 of 4

Audit the spread, not only the center

Does a band labelled 80 percent contain the outcome 80 percent of the time?

Each point is the share of cases at that horizon whose actual value fell inside the band; the dashed line is the share the band promises. A well calibrated band sits near its line on average, while a band can reach any coverage by growing wide enough, which is why width belongs beside coverage.

Equation: coverage equals the fraction of the n cases whose actual value y i lies between the band ends L i and U i

Scroll sideways for the whole equation

L and U are the lower and upper ends of the band for case i, y the value that happened, and the indicator 1 counts a case when y falls inside. A case is one entity, origin and horizon; n is the number of cases.

Predict first. Over all eight held-out entities, what fraction of actual values fall inside the nominal 80 percent band?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Audit the spread, not only the center. Observed coverage by horizon for the 80 percent band over all eight held-out entities, against the dashed nominal line; overall 0.804, horizons ranging 0.667 to 0.958.
Nominal band: 80 percent, Cases scored: All eight held-out entities
Constructed data: the chapter notebook's 160 sample paths per entity and origin (seed 20260932); the 50 and 98 percent bands are read from the same paths.

Calculated values

Cases scored
336
Inside the band
270
Observed coverage
0.804
Nominal coverage
0.80
Observed minus nominal
0.004
Mean band width (index units)
6.88
Lowest and highest horizon
0.667 and 0.958

For the 80 percent band over all eight held-out entities, coverage = 270 / 336 = 0.804, which is above the nominal level (0.804 - 0.80 = 0.004). Each horizon rests on only 24 cases from overlapping origins, so the points swing from 0.667 to 0.958. The mean width is 6.88: a wider band covers more simply by being wider.

Worked steps

  1. Cases: 8 entities x 3 origins x 14 days = 336.
  2. Actual values inside the 80 percent band: 270.
  3. Coverage: 270 / 336 = 0.804.
  4. Against nominal: 0.804 - 0.80 = 0.004, with mean width 6.88.

Use the idea

When someone hands you an interval forecast, ask for its coverage on cases it was not fitted to, the number of cases behind it and the band's width, before using the interval in a decision.

Where the conclusion applies

Coverage here is marginal, day by day, on related synthetic series. Overlapping origins make cases dependent, so 336 cases are worth fewer independent checks; the result is a diagnostic, not a calibration guarantee.

Check your understanding: A forecaster reports 90 percent bands, and 61 of 80 held-out outcomes fall inside. What is the coverage, and is it near the nominal level?
61 / 80 = 0.7625, about 0.76, well below 0.90 (0.7625 - 0.90 = -0.1375), so the bands are too narrow for these cases.

Chapter 14 source: section "Section Five: The Distribution and the Decision".