The Art and Science of Forecasting

Chapter 18

The Hierarchy

A total and its parts cannot both be right if they do not add up, and the fix should use every level's information.

Four demonstrations follow the chapter: what forcing every level to one official number keeps and loses, why summing the stores can help or hurt a national forecast, how bottom-up, OLS and MinT reconciliation move each level, and why accuracy after reconciliation has to be measured on held-out cases.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

One number by decree

When the total and its parts disagree, what does scaling everything to the official total keep and what does it lose?

Coherence means the bars for the parts add up to the bar for the total. Scaling restores it by multiplying every part by one factor, which keeps the relative proportions and forces the total. Summing the parts restores it the other way, by discarding the total.

Equation: y, the vector of values at every level, equals the summing matrix S times b, the vector of bottom level values

Equation: the scale factor k equals the official total forecast divided by the sum of the forecasts being scaled

Scroll sideways for the whole equation

y is the vector of values at every level, b the bottom-level values and S the summing matrix whose rows say which bottom values add up to each node. k is the scale factor: the official total divided by the sum of the parts being scaled.

Predict first. In the two-store example the national forecast is 210 and the stores say 90 and 105. If the stores are scaled to 210, does the ratio of store A to store B change?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: One number by decree. Bars for the national forecast, store A, store B and the sum of stores: 210.0, 96.9, 113.1, 210.0 units.
Example: Two stores (notebook), How the conflict is resolved: Scale the parts to the national number
Constructed data: the chapter notebook's illustrative base forecasts (national 210, stores 90 and 105, cell 3) and the chapter's composite retail figures of 4.2, 4.4 and 3.9 billion dollars.

Calculated values

Scale factor k
1.0769
Store A scaled
96.9
Store B scaled
113.1
Store A over store B
0.857
Coherent
yes

The national number wins: k = 210 / 195 = 1.0769, so store A becomes 90 x 1.0769 = 96.9 and store B 105 x 1.0769 = 113.1, and 96.9 + 113.1 = 210.0. Store A over store B: 90 / 105 = 0.857 before and 0.857 after: the relative store proportions survive, but a store whose own error ran the other way is pushed in the wrong direction.

Worked steps

  1. k = 210 / 195 = 1.0769.
  2. Store A: 90 x 1.0769 = 96.9.
  3. Store B: 105 x 1.0769 = 113.1.
  4. Check: 96.9 + 113.1 = 210.0.

Use the idea

Before a planning meeting declares one number, write down what each level's forecast knew that the others did not, and ask which of that knowledge the chosen resolution throws away.

Where the conclusion applies

The two-store figures are the chapter notebook's illustrative base forecasts; the 4.2, 4.4 and 3.9 billion dollar figures are the chapter's composite, not a real company. Scaling fixes the total, but it cannot correct an error that belongs to one location, and it says nothing about whether the official total was right.

Check your understanding: A national forecast is 500 and two stores forecast 200 and 250. With a scale factor rounded to 1.111, what does proportional scaling give each store?
k = 500 / 450 = 1.111. Store one: 200 x 1.111 = 222.2. Store two: 250 x 1.111 = 277.8. Check: 222.2 + 277.8 = 500.0.

Chapter 18 source: section "Section One: The Forecasts That Didn't Add Up".

Demonstration 2 of 4

Does adding the stores cancel their errors?

When is the sum of store forecasts a better national forecast than forecasting the nation directly?

Store errors that move together add up in the total; errors that offset cancel. The covariance bar is added to the two store variances to give the bottom-up total, which is compared with the direct total's variance.

Equation: the variance of the sum of the two store errors equals the variance of e A plus the variance of e B plus twice their covariance

Scroll sideways for the whole equation

e A and e B are the forecast errors of store A and store B. Var is a variance and Cov a covariance, in units squared. The direct total is a forecast made at the national level with its own error variance.

Predict first. With the notebook's error covariances (store variances 36 and 25, covariance 4, direct total 64), is the bottom-up total more accurate than the direct total?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Does adding the stores cancel their errors?. Bars of error variance: store A 36, store B 25, twice the covariance 8, bottom-up total 69 and direct total 64.
Covariance of the two store errors: 4, Error variance of the direct total: 64 (notebook)
Constructed data: the chapter notebook's error covariance (cell 5) and its 400 seeded test cases (cell 7, seed 20260936), with the store covariance and the direct variance varied for teaching.

Calculated values

Covariance of store errors
4
Bottom-up total variance
69
Bottom-up total standard deviation
8.31
Direct total standard deviation
8.00
Notebook test RMSE, bottom-up and direct
8.63 and 8.57

Bottom-up total variance = 36 + 25 + 2 x 4 = 69, a standard deviation of sqrt(69) = 8.31, against sqrt(64) = 8.00 for the direct total, so the direct total forecast has the smaller error here. On the notebook's 400 test cases the measured total RMSE agrees: 8.63 for bottom-up against 8.57 for the direct forecast. Shared errors add, offsetting errors cancel: which level wins is a held-out question, not a rule.

Worked steps

  1. Store variances: 36 + 25 = 61.
  2. Covariance term: 2 x 4 = 8.
  3. Bottom-up total: 61 + 8 = 69.
  4. Standard deviations: sqrt(69) = 8.31 against sqrt(64) = 8.00.

Use the idea

Before choosing bottom-up for a national number, estimate how much of the store errors is shared (weather, the economy, a national promotion) and compare both routes on held-out periods.

Where the conclusion applies

The variances 36, 25 and 64 and the covariance 4 are the notebook's constructed error process; the other covariances and the direct variance of 81 are teaching variations. Real errors need estimating from past forecasts, and an estimate from a short history can be far from the truth.

Check your understanding: Two store errors have variances 9 and 16 and covariance -6. What is the variance of the bottom-up total?
9 + 16 + 2 x (-6) = 13, smaller than either 9 + 16 = 25 because the errors partly offset.

Chapter 18 source: section "Section Two: Why Naive Reconciliation Fails".

Demonstration 3 of 4

Reconciliation moves every level

How do bottom-up, equal-weight OLS and covariance-weighted MinT turn conflicting forecasts into ones that add up?

Each method finds coherent forecasts from the base ones. Bottom-up keeps the stores, OLS moves every level by the same amount, and MinT moves most the levels whose estimated errors are largest. Every reconciled set adds up.

Equation: the reconciled vector y tilde equals S times G times the base forecast vector y hat

Equation: G equals the inverse of S transpose times W inverse times S, multiplied by S transpose times W inverse

Scroll sideways for the whole equation

y-hat is the vector of base forecasts (total, store A, store B) and y-tilde the reconciled vector. S is the summing matrix, G the matrix that maps base forecasts to bottom-level values, and W the error covariance matrix: the identity for OLS, an estimate from 180 earlier errors for MinT.

Predict first. With base forecasts 210, 90 and 105, will MinT's reconciled total be closer to 210 or to 195?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Reconciliation moves every level. Paired bars of base and MinT forecasts for total, store A and store B: base 210, 90, 105; reconciled 202.0, 93.7, 108.3.
Base forecast of the total: 210, Reconciliation: MinT (estimated covariance)
Constructed data: the chapter notebook's base forecasts and its 180 seeded past errors (cell 5, seed 20260936), reconciled by the companion's own reconciliation tool; base totals 195 and 225 are teaching variations.

Calculated values

Base gap, total minus stores
15
MinT total
201.96
MinT store A
93.68
MinT store B
108.28
Total moved by
8.04

Base total 210 against stores 90 + 105 = 195, a gap of 15. MinT moves the total down by 210 - 201.96 = 8.04, more than OLS would, because the estimated covariance says the total's errors are the largest. After reconciliation 93.68 + 108.28 = 201.96, so the forecasts are coherent. Coherence is guaranteed by the projection; whether accuracy improved needs outcomes.

Worked steps

  1. Gap: 210 - (90 + 105) = 15.
  2. MinT moves the total down by 210 - 201.96 = 8.04, more than OLS would, because the estimated covariance says the total's errors are the largest.
  3. Coherence check: 93.68 + 108.28 = 201.96.

Use the idea

When levels of a plan disagree, reconcile with the error history of each level rather than by rank, and keep the base forecasts beside the reconciled ones so the moves can be explained.

Where the conclusion applies

The covariance is estimated from the notebook's 180 seeded past errors and shrunk 20 percent toward its diagonal. With unbiased base forecasts and the true covariance, MinT cannot do worse than the base forecast at any node; an estimated covariance carries no such guarantee, and nothing here keeps forecasts nonnegative.

Check your understanding: Base forecasts are 120 for the total and 40 and 50 for the stores. What does equal-weight OLS reconciliation give?
Gap: 120 - (40 + 50) = 30, and 30 / 3 = 10. Stores 50 and 60, total 110. Check: 50 + 60 = 110.

Chapter 18 source: section "Section Three: The MinT Solution".

Demonstration 4 of 4

Coherence is guaranteed, accuracy is measured

Does reconciling with an estimated covariance make every level more accurate on cases it never saw?

Each group of bars is one node; each bar a method's root mean squared error over the chosen test cases. MinT used a covariance estimated before these cases were drawn, so this is an honest held-out comparison.

Equation: root mean squared error equals the square root of the average squared difference between the forecast and the outcome

Scroll sideways for the whole equation

RMSE is the root mean squared error of a node's forecasts over n test cases, y-tilde the forecast and y the outcome, in units. Base forecasts are the outcome plus an error drawn from the notebook's error process.

Predict first. Over all 400 test cases, is the bottom-up total more accurate than the base total forecast?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Coherence is guaranteed, accuracy is measured. Grouped bars of test RMSE for base, bottom-up and MinT at total, store A and store B over 400 cases; total: 8.57, 8.63 and 6.93.
Test cases scored: All 400
Constructed data: the chapter notebook's 400 seeded test cases (cell 7, seed 20260936), scored with the notebook's MinT matrix and bottom-up sums.

Calculated values

Test cases
400
Total RMSE, base
8.57
Total RMSE, bottom-up
8.63
Total RMSE, MinT
6.93
Nodes where MinT beats base
3 of 3

Over the first 400 test cases, base minus MinT total RMSE is 8.57 - 6.93 = 1.64, so MinT is lower at the total, and MinT is lower at 3 of 3 nodes. Bottom-up minus base at the total is 8.63 - 8.57 = 0.06: summing the stores inherits their errors. Coherence holds in every case; accuracy is what these held-out errors measure, under this one constructed error process.

Worked steps

  1. Base total RMSE: 8.57; MinT: 6.93.
  2. Difference: 8.57 - 6.93 = 1.64.
  3. Bottom-up minus base: 8.63 - 8.57 = 0.06.
  4. MinT beats base at 3 of 3 nodes.

Use the idea

Report a reconciliation with a held-out comparison against the base forecasts, node by node, and treat a coherent forecast that lost accuracy as a finding.

Where the conclusion applies

The test truths and errors are the notebook's seeded draws from one known error process, independent of the 180 errors used to estimate the covariance. The guarantee of no worse mean squared error at any node needs unbiased base forecasts and the true covariance; with an estimate, and with few cases, a node can lose.

Check your understanding: On a holdout the base total RMSE is 8.57 and the MinT total RMSE is 6.93. By what fraction is MinT lower?
(8.57 - 6.93) / 8.57 = 0.191, about 19 percent lower for this constructed error process.

Chapter 18 source: section "Section Four: The Full Methodology".