The Art and Science of Forecasting

Chapter 21

The Bullwhip

A supply chain can manufacture its own volatility, and the cure starts with forecasting the right signal for the right decision.

Four demonstrations follow the chapter: how a sensible restocking rule amplifies steady demand stage by stage, how much of that comes from forecasting a neighbour's orders instead of consumer sales, why the cheapest stock level is a cost-weighted quantile, and which intermittent-demand forecast notices when an item stops selling.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

A sensible restocking rule turns steady demand into swings

If customers buy at a steady pace, why do the orders reaching the factory swing so hard?

Each stage keeps a stock: arrivals flow in, the orders it fills flow out. It forecasts what it is asked for and orders enough to lift stock plus pipeline to a target of lead time plus one weeks of forecast plus the buffer. A small rise in demand both lifts the forecast and opens an inventory gap, so the order overshoots; the next stage up treats that overshoot as demand and repeats it.

Equation: stock at t plus 1 equals stock at t plus inflow at t minus outflow at t

Equation: the forecast d hat t equals alpha times d t plus 1 minus alpha times the previous forecast

Equation: the order o t is the larger of zero and L plus 1 times the forecast, plus 30, minus net inventory I t, minus pipeline P t

Scroll sideways for the whole equation

d is the demand a stage receives in week t (consumer purchases for the retailer, the downstream stage's orders for everyone else) and d-hat its smoothed forecast with weight alpha = 0.5. o is the order placed, L the lead time in weeks, I the net inventory (negative when backlogged), P the units already ordered but not yet arrived, and 30 a fixed safety buffer.

Predict first. With a three week lead time, will the factory's orders vary more than 100 times as much as consumer demand?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A sensible restocking rule turns steady demand into swings. Weekly consumer demand stays near 50 while factory orders with a 3 week lead time swing between 0 and 462 over weeks 150 to 199.
Stage shown: Factory, Lead time (weeks): 3
Constructed data: the chapter notebook's seeded consumer demand and its three stage inventory-position chain (cell 3), with the lead time varied.

Calculated values

Stage
Factory
Lead time (weeks)
3
Consumer demand variance
22.04
Factory order variance
7472.8
Variance ratio
339.1
Week 160 order
0.00

In week 160 the factory forecasts 24.83 a week and orders 4 x 24.83 + 30 - 172.80 - 174.05 = -217.53, below zero, so the order is 0. After the 50 warm-up weeks its orders vary 7472.8 / 22.04 = 339.1 times as much as consumer demand, though every stage follows the same sensible rule with a 3 week delay.

Worked steps

  1. Target position: 4 x 24.83 + 30 = 129.32.
  2. Subtract net inventory 172.80 and pipeline 174.05: order 0.00.
  3. Variance ratio after warm-up: 7472.8 / 22.04 = 339.1.

Use the idea

Before blaming customers for volatile orders, compare the variance of orders at your stage with the variance of end-consumer sales over the same weeks.

Where the conclusion applies

Consumer demand is seeded normal noise around 50 (seed 20260939); every stage starts with 200 units and a pipeline of 50 a week and orders weekly. Lead times 1 and 5 vary the notebook's three week rule. Real chains add batching, price promotions and shortage gaming, which this rule leaves out.

Check your understanding: A stage forecasts 52 a week, has a two week lead time, holds 120 units and has 95 on order. With the buffer of 30, what does it order?
3 x 52 + 30 - 120 - 95 = 156 + 30 - 215 = -29, below zero, so it orders 0 this week.

Chapter 21 source: section "Section One: Forrester at the Plant".

Demonstration 2 of 4

Forecast the consumer, not the neighbour's orders

How much of the upstream swing comes from forecasting orders as if they were demand?

When a stage forecasts its neighbour's orders, it reads the neighbour's inventory corrections as changes in demand and adds its own on top. Forecasting the consumer signal removes that echo; lead times, buffers and ordering rules stay exactly the same, so whatever amplification remains comes from the stock-correction rule itself.

Equation: B equals the variance of orders divided by the variance of consumer demand

Equation: the forecast d hat t equals alpha times d t plus 1 minus alpha times the previous forecast

Scroll sideways for the whole equation

B is the variance of a stage's weekly orders divided by the variance of consumer demand, both measured after 50 warm-up weeks. alpha is the forecast weight on the newest observation.

Predict first. At alpha 0.5, does forecasting consumer sales at every stage cut the factory's ratio by more than half?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Forecast the consumer, not the neighbour's orders. Bars of variance ratio by stage on a log scale for alpha 0.5: factory 339.0 with local forecasts against 52.7 with shared consumer sales.
Forecast weight alpha: 0.5
Constructed data: the chapter notebook's seeded chain with local and shared consumer information (cells 3 and 10), with the forecast weight varied.

Calculated values

Forecast weight alpha
0.5
Retailer ratio (either)
10.1
Factory ratio, own orders
339.0
Factory ratio, consumer sales
52.7
Reduction factor at the factory
6.4

With alpha 0.5 the factory's variance ratio is 339.0 when each stage forecasts the orders it receives and 52.7 when it forecasts consumer sales: 339.0 / 52.7 = 6.4, so sharing lowers the factory ratio here. The retailer already sees consumer sales, so its ratio (10.1) is the same either way; the stock-correction part of the rule still amplifies, which is why sharing shrinks the whip without removing it.

Worked steps

  1. Factory ratio forecasting neighbour's orders: 339.0.
  2. Factory ratio forecasting consumer sales: 52.7.
  3. Reduction: 339.0 / 52.7 = 6.4.

Use the idea

Measure amplification by comparing order variance with consumer variance in the same units and weeks, then decide whether to change the forecasting input, the ordering policy, or both. A large forecast error on its own is a reason to inspect the model, not proof of bullwhip.

Where the conclusion applies

Same seeded consumer demand and chain as the notebook (cell 10); only the information each stage forecasts from changes. Alpha 0.2 and 0.8 vary the notebook's 0.5. Sharing a signal does not remove physical delays or capacity limits.

Check your understanding: A distributor's order variance is 400 and consumer demand variance is 25. After sharing sales data its order variance is 150. What are the two ratios?
400 / 25 = 16 before and 150 / 25 = 6 after, so sharing cut the ratio by 16 / 6 = 2.7 times.

Chapter 21 source: section "Section Three: When the Model Sees Noise as Signal".

Demonstration 3 of 4

The cheapest stock level is a quantile, not the mean

If running short costs three times as much as a leftover, how much more than average demand should you stock?

The curve is the average cost over 30,000 simulated demands for each stock level. Stock one more unit and you save c u with probability 1 - F(q) and lose c o with probability F(q); the two balance where F(q) equals c u / (c u + c o). The condition is on the cumulative probability that demand does not exceed q, not on the chance of selling one more unit.

Equation: the cumulative demand probability at the best quantity q star equals c u over c u plus c o

Scroll sideways for the whole equation

q is the quantity stocked and q-star the cheapest one. F is the cumulative probability that demand is at most q. c u is the underage cost per unit short, c o the overage cost per unit left over (3 here). Demand is normal with mean 100 and standard deviation 20.

Predict first. With underage cost 9 and overage cost 3, will the cheapest stock be above 110 units?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The cheapest stock level is a quantile, not the mean. Expected cost against units stocked for underage cost 9 and overage cost 3; the minimum is near 113.5, marked against the mean demand of 100.
Underage cost per unit short: 9
Constructed data: the chapter notebook's 30,000 seeded normal demands and stocking cost curve (cell 5), with the underage cost varied around its 9.

Calculated values

Underage cost
9
Overage cost
3
Critical ratio
0.750
Critical-fractile quantity
113.5
Cheapest whole quantity on the curve
113
Expected cost at 100
95.3
Expected cost at 113
75.5

Critical ratio = 9 / (9 + 3) = 0.750. The demand quantile at that probability is 100 + 20 x 0.674 = 113.5 units. Stocking the mean costs 95.3 against 75.5 at 113, because a shortage costs more than a leftover. The simulated cost curve bottoms out at 113, agreeing with the formula to within one unit.

Worked steps

  1. Critical ratio: 9 / (9 + 3) = 0.750.
  2. Standard normal quantile at 0.750: 0.674.
  3. Stock: 100 + 20 x 0.674 = 113.5.
  4. Simulated cheapest whole quantity: 113.

Use the idea

When a shortage and a leftover cost different amounts, choose the stock level as the quantile of the demand forecast at the critical ratio rather than the mean forecast.

Where the conclusion applies

One selling period, demand normal with known mean and spread, costs linear per unit. The simulated demand uses the notebook's seeded draws (seed 20260939). With a misjudged spread the quantile is wrong too.

Common wrong turn: Stock the expected demand
The chapter says the right inventory level is not the expected demand level: it moves up or down with the asymmetry of the costs. Only equal costs give the median, here 100.
Check your understanding: Underage costs 4 and overage costs 1. With demand normal, mean 200 and standard deviation 30, and the 0.8 quantile of a standard normal 0.842, what should you stock?
4 / (4 + 1) = 0.8, so stock 200 + 30 x 0.842 = 225.3, about 225 units.

Chapter 21 source: section "Section Four: The Full Methodology".

Demonstration 4 of 4

When an item stops selling, which forecast notices?

After an item's last sale, which intermittent-demand forecast lets go of it?

Croston splits the series into how much is bought when something is bought and how long between purchases, and divides one by the other. SBA multiplies by 1 - alpha / 2 to correct Croston's upward bias. TSB replaces the interval with an occurrence probability that is smoothed every period, so a run of zeros pulls it down.

Equation: the Croston forecast equals smoothed size z hat divided by smoothed interval p hat

Equation: the SBA forecast equals 1 minus alpha over 2, times z hat over p hat

Equation: the TSB forecast equals the smoothed occurrence probability pi hat times z hat

Scroll sideways for the whole equation

z-hat is the smoothed demand size when demand occurs, p-hat the smoothed number of periods between demands, and pi-hat the smoothed probability that a period has demand. alpha is the smoothing weight for all three. Forecasts are made before each period's demand is seen.

Predict first. Forty periods after the last sale (period 119, alpha 0.15), will Croston's forecast be lower than it was at period 80?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: When an item stops selling, which forecast notices?. Intermittent demand with mostly zero periods, stopping after period 80, with Croston, SBA and TSB forecasts at smoothing 0.15; at period 119 Croston is 0.866 and TSB 0.0006.
Period inspected: 119, Smoothing weight alpha: 0.15
Constructed data: the chapter notebook's seeded intermittent series that stops after period 80 (cell 7). The bundled example file for this chapter has no zero-demand periods, so it cannot show this behaviour.

Calculated values

Zero-demand periods in 0 to 79
70 of 80
Smoothed size
5.51
Smoothed interval
6.36
Smoothed occurrence probability
0.0001
Croston
0.866
SBA
0.801
TSB
0.0006

Croston = 5.51 / 6.36 = 0.866 units a period; SBA = 0.925 x 0.866 = 0.801; TSB = 0.0001 x 5.51 = 0.0006. No demand occurs from period 80 on, so at period 119 Croston and SBA still hold the values set at the last sale, because they update only when demand occurs, while TSB has shrunk its occurrence probability in every zero period.

Worked steps

  1. Croston: size 5.51 / interval 6.36 = 0.866.
  2. SBA: (1 - 0.15 / 2) x 0.866 = 0.801.
  3. TSB: probability 0.0001 x size 5.51 = 0.0006.

Use the idea

For spare parts and end-of-life items, check whether your intermittent forecast can fall during a long run of zeros before trusting it to stop replenishment.

Where the conclusion applies

The series is the notebook's constructed one: each period has demand with probability 0.2, sized 1 plus a Poisson count with mean 5, and all demand stops after period 80 (seed 20260939). Alpha 0.3 varies the notebook's 0.15. Classifying a series by demand interval and size variability suggests candidate methods, but only validation shows which works.

Check your understanding: A Croston forecast has smoothed size 5 and smoothed interval 4. With alpha 0.2, what are the Croston and SBA forecasts?
Croston: 5 / 4 = 1.25 units a period. SBA: (1 - 0.2 / 2) x 1.25 = 0.9 x 1.25 = 1.125.

Chapter 21 source: section "Section Four: The Full Methodology".