Demonstration 1 of 4
A model trained on other series forecasts one it never saw
Can a network trained on 32 series forecast a 33rd it never trained on, from only 14 of its values?
Each pair of bars is one entity the network never trained on: its error on the left, the seasonal naive error on the right. Both see only the 14 values before each of three forecast origins. The network brings the weekly shape it learned from 32 other series; the naive forecast copies one noisy week.
Scroll sideways for the whole equation
y is the observed demand index on a forecast day and y-hat a forecast of it. MAE is the mean absolute error over the days scored, in index units. The global forecast is the median of the network's sampled paths; the local forecast repeats the entity's last observed week.
Predict first. Over all eight held-out entities and all 14 days, will the global network beat the local seasonal naive?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's 40 seeded synthetic demand indices (seed 20260932), 32 for training and 8 held out, scored at origins 140, 154 and 166.
Calculated values
- Forecasts scored
- 336
- Global network MAE
- 2.127
- Local seasonal naive MAE
- 3.081
- Local minus global
- 0.954
- Of all 8 entities, global wins on
- 8
Over all eight held-out entities, days 1 to 14, local MAE minus global MAE is 3.081 - 2.127 = 0.954, so the global model has the lower error here. Neither model saw these entities during training; each forecast starts from the same 14 recent values. The network's advantage is borrowed from the 32 training entities, which share the weekly shape by construction.
Worked steps
- Forecasts scored: 336 (8 entities x 3 origins x 14 days).
- Global network MAE: 2.127. Local seasonal naive MAE: 3.081.
- Difference: 3.081 - 2.127 = 0.954.
- The global model wins on 8 of the 8 entities over days 1 to 14.
Use the idea
When a new product has a few weeks of history, compare a model trained across related products with a simple local forecast on products held out of training, not on the products the model learned from.
Where the conclusion applies
The 40 series are synthetic and share one weekly pattern by design, which favours pooling. With unrelated series, or a held-out entity with a different cycle, the global model's advantage can vanish. This small network is not DeepAR: it has no recurrence and no item embeddings.
Check your understanding: A global model scores MAE 1.505 on an entity and a local forecast scores 1.949. By how much does the global model reduce the error, and does that prove pooling helps for every new product?
Chapter 14 source: section "The Catalog and the Cold Start".
Demonstration 2 of 4
The same point forecast, a different safety stock
If two products both average 100 units a week, can they share one safety stock?
The curve is the true weekly demand; the solid line is the stock, set at the mean plus 1.645 of the planned standard deviations. The shaded tail beyond the line is the share of weeks that run out. When the planned and true standard deviations differ, the stock and the target come apart.
Scroll sideways for the whole equation
S is the stock held, mu the mean weekly demand (100 units), sigma the standard deviation of weekly demand in units and z the number of standard deviations the target needs: 1.645 for a 95 percent service level.
Predict first. Stock is set for a standard deviation of 10, but demand really varies with standard deviation 30. Roughly what service level is reached?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter's own inventory example (mean 100, standard deviation 10 or 30, 95 percent service level), with standard deviation 20 added.
Calculated values
- Safety stock set
- 16.45
- Stock held
- 116.45
- Standard deviations covered
- 0.548
- Service level reached
- 70.8 percent
- Stock needed for 95 percent
- 149.35
Safety stock 1.645 x 10 = 16.45, so the stock is 116.45. Against true variability 30 it covers (116.45 - 100) / 30 = 0.548 standard deviations, which the normal curve turns into 70.8 percent: the stock falls short of the 95 percent target. The shaded tail is the share of weeks that run out.
Worked steps
- Safety stock: 1.645 x 10 = 16.45.
- Stock held: 100 + 16.45 = 116.45.
- Standard deviations covered: (116.45 - 100) / 30 = 0.548.
- Normal curve below 0.548 standard deviations: 70.8 percent of weeks without a stockout.
Use the idea
Before reusing one safety stock buffer across products, check that each product's own demand variability is the one the buffer was computed for; the point forecast alone cannot tell you.
Where the conclusion applies
Demand is treated as normal with a known mean, as in the chapter's calculation. Count demand that is skewed or intermittent breaks the normal approximation, which is why the chapter turns to the negative binomial.
Common wrong turn: A fixed buffer is safe for every product
Check your understanding: Mean weekly demand is 100 with standard deviation 20. What stock gives a 95 percent service level?
Chapter 14 source: section "Section Three: The Probabilistic Imperative".
Demonstration 3 of 4
A distribution, sampled one day at a time
How does a network that predicts only the next day produce a band for the next fourteen?
The network sees only the last 14 values, divided by their mean absolute level, and predicts a mean and a spread for the next day. Drawing from that distribution, appending the draw and repeating gives one path; 160 paths give the median and the band. No value after the origin enters any path.
Scroll sideways for the whole equation
s is the mean absolute level of the 14 most recent values, m and d the network's mean and spread for the day on that relative scale, and epsilon a standard normal draw. Each drawn y is added to the history before the next day is drawn.
Predict first. For entity 32 at origin 154, will the 80 percent band be wider on day 14 than on day 1?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded synthetic entities, never used for weight updates, with the notebook's 160 recursive sample paths per origin (seed 20260932).
Calculated values
- Context scale s
- 60.31
- Network mean m, day 1
- -0.003
- Network spread d, day 1
- 0.050
- Day 1 mean
- 60.13
- Day 1 standard deviation
- 3.02
- 80 percent band width, day 1
- 7.31
- 80 percent band width, day 14
- 8.29
- Days inside the band
- 13 of 14
Entity 32 at origin 154: the 14 context values have mean absolute level s = 60.31. The network returns m = -0.003 and d = 0.050 for day 1, so the mean is 60.31 x (1 + (-0.003)) = 60.13 and the standard deviation 60.31 x 0.050 = 3.02. Each sampled value is fed back as a lag for the next day. Here the band on day 14 is wider than on day 1: 8.29 against 7.31. The actual values fell inside it on 13 of 14 days.
Worked steps
- Scale from the context only: s = 60.31.
- Day 1 mean: 60.31 x (1 + (-0.003)) = 60.13.
- Day 1 standard deviation: 60.31 x 0.050 = 3.02.
- Draw 160 values, feed each back as a lag, and repeat for days 2 to 14.
- Band width: 7.31 on day 1, 8.29 on day 14.
Use the idea
To read a probabilistic forecast, ask how its band was built: sampled paths carry uncertainty forward, and a band from 10th to 90th percentiles of each day is a statement about each day, not about a whole path.
Where the conclusion applies
The model is a small Gaussian network, not DeepAR's recurrent network with a negative binomial output; a Gaussian can put probability on negative demand. Each band uses 160 sampled paths (the book's figure uses 400 for entity 32).
Check your understanding: The 14 context values have mean absolute level 50, and the network predicts m = 0.1 and d = 0.2. What are the mean and standard deviation of the next value?
Chapter 14 source: section "Section Four: The Neural Forecasting Landscape".
Demonstration 4 of 4
Audit the spread, not only the center
Does a band labelled 80 percent contain the outcome 80 percent of the time?
Each point is the share of cases at that horizon whose actual value fell inside the band; the dashed line is the share the band promises. A well calibrated band sits near its line on average, while a band can reach any coverage by growing wide enough, which is why width belongs beside coverage.
Scroll sideways for the whole equation
L and U are the lower and upper ends of the band for case i, y the value that happened, and the indicator 1 counts a case when y falls inside. A case is one entity, origin and horizon; n is the number of cases.
Predict first. Over all eight held-out entities, what fraction of actual values fall inside the nominal 80 percent band?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's 160 sample paths per entity and origin (seed 20260932); the 50 and 98 percent bands are read from the same paths.
Calculated values
- Cases scored
- 336
- Inside the band
- 270
- Observed coverage
- 0.804
- Nominal coverage
- 0.80
- Observed minus nominal
- 0.004
- Mean band width (index units)
- 6.88
- Lowest and highest horizon
- 0.667 and 0.958
For the 80 percent band over all eight held-out entities, coverage = 270 / 336 = 0.804, which is above the nominal level (0.804 - 0.80 = 0.004). Each horizon rests on only 24 cases from overlapping origins, so the points swing from 0.667 to 0.958. The mean width is 6.88: a wider band covers more simply by being wider.
Worked steps
- Cases: 8 entities x 3 origins x 14 days = 336.
- Actual values inside the 80 percent band: 270.
- Coverage: 270 / 336 = 0.804.
- Against nominal: 0.804 - 0.80 = 0.004, with mean width 6.88.
Use the idea
When someone hands you an interval forecast, ask for its coverage on cases it was not fitted to, the number of cases behind it and the band's width, before using the interval in a decision.
Where the conclusion applies
Coverage here is marginal, day by day, on related synthetic series. Overlapping origins make cases dependent, so 336 cases are worth fewer independent checks; the result is a diagnostic, not a calibration guarantee.
Check your understanding: A forecaster reports 90 percent bands, and 61 of 80 held-out outcomes fall inside. What is the coverage, and is it near the nominal level?
Chapter 14 source: section "Section Five: The Distribution and the Decision".