The Art and Science of Forecasting

Chapter 27

Giving the Oracle Direction

Model what the data can support, estimate what it cannot, and let a second route test the number.

Four demonstrations follow the chapter's procedure: the router that decides between modelling and estimating, one shared scale calibrated on established products and tested on held-out ones, the build-up that decides how much of a launch lands in year one, and two forecasting engines whose gap points at a wrong input.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Model or estimate: the router reads the history first

How much history does a seasonal series need before it should be modelled at all?

The teal stretch is the history the router is given; the dashed lines show how far back the two thresholds reach. When the teal stretch does not reach the router line, the tool returns a request for evidence and no number, rather than fitting a curve to too little data.

Equation: the number of observations N is at least two times the season length s

Equation: the router minimum equals the larger of 36 and four times the horizon h

Scroll sideways for the whole equation

N is the number of monthly observations available, s the season length (12 months), and h the forecast horizon in months. N min is the companion router's minimum before it compares models.

Predict first. With 47 months of history and a 12-month horizon, will the router model the series or ask for evidence?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Model or estimate: the router reads the history first. Monthly sea surface temperature for the 120 months to December 2010, with the last 24 months highlighted and dashed lines at the router minimum of 48 months and at two cycles, 24 months.
Months of history given: 24, Forecast horizon (months): 12
Real data: monthly Nino 1+2 sea surface temperature in degrees Celsius, 1950 to 2010, from the US National Oceanic and Atmospheric Administration (ERSST v3b), public domain, bundled with the companion. The decision is the chapter's applied routing check; the full model comparison is not run on this page.

Calculated values

Months of history given
24
Forecast horizon (months)
12
Router minimum, max(36, 4 times horizon)
48
Two full seasonal cycles
24
Meets the two-cycle floor
yes
Router decision
needs evidence: no number is produced

With a 12-month horizon the router needs 4 x 12 = 48 months, or 36 if that is larger, so 48. Two seasonal cycles are 2 x 12 = 24 months. Given 24 months, the decision is: needs evidence: no number is produced. This history clears the two-cycle floor (24 of 24 months) but not the router's 48: the companion asks for more than the bare minimum before it compares models. Fewer points than that would hand back sampling noise as a seasonal pattern, so the honest output is a request for evidence, not a curve.

Worked steps

  1. Four horizons: 4 x 12 = 48 months.
  2. Router minimum: the larger of 36 and 48, so 48.
  3. Two seasonal cycles: 2 x 12 = 24 months.
  4. History given: 24 months, so needs evidence: no number is produced.

Use the idea

Before asking for a forecast, count the clean periods you actually have and compare them with the season and the horizon. If they fall short, switch to the estimate path: analogues, defended proxies and a range.

Where the conclusion applies

History length is only one input to the routing decision: the chapter also requires a stable regime and an identifiable relationship, which no count of points can check. The companion's minimum of the larger of 36 and four horizons is stricter than the chapter's two-cycle floor; both are working conventions, not laws.

Check your understanding: A quarterly series (season length 4) is to be forecast 6 quarters ahead. How many quarters do two seasonal cycles need, and how many does the router rule ask for?
Two cycles: 2 x 4 = 8 quarters. Router: the larger of 36 and 4 x 6 = 24, so 36 observations.

Chapter 27 source: section "Section Three: The Router: Model or Estimate?".

Demonstration 2 of 4

One shared scale, tested on products it never saw

Can one conversion number, fitted on established products, explain products it was not fitted to?

Every product's drivers multiply to an exposure; the shared scale turns exposure into units. Teal dots fitted the scale, terracotta squares were held out. Points on the dashed line are explained exactly. Because only one number is fitted, the model cannot bend itself to each product, so held-out agreement means something.

Equation: exposure X i equals eligible buyers times awareness times availability times interest times units per buyer

Equation: the shared scale c is the sum of X i times V i over the sum of X i squared, and modelled units V hat i equal c times X i

Scroll sideways for the whole equation

For product i, B is eligible buyers, a awareness among them, d availability among the aware, r purchase interest and u units per buyer over 24 months; X is their product. V is observed 24-month units, c the one shared scale, and V-hat the modelled units.

Predict first. If the scale is fitted on only the first 6 products instead of 24, will the held-out error stay small?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: One shared scale, tested on products it never saw. Scatter of observed against modelled 24-month units for 30 synthetic products; 24 fit the shared scale 0.8538 and six held-out squares sit close to the equal line, mean absolute error 1,843 units.
Products used to fit the scale: 24
Constructed data: the chapter notebook's 30 seeded synthetic reference products (seed 20260945), 24 for calibration and 6 held out; the scale is the companion's own shared calibration.

Calculated values

Products used to fit
24
Shared scale
0.8538
Held-out MAE (units)
1,843
Held-out mean absolute percentage error
3.0 percent
Scale planted by the generator
0.85

Fitting one scale to 24 products gives c = 0.8538. For held-out product 25 the drivers multiply to 97,916, so the model says 97,916 x 0.8538 = 83,601 units against 88,181 observed. Across the six held-out products the mean absolute error is 1,843 units (3.0 percent). One number, fitted elsewhere, explains products it never saw; that supports the relationship across this portfolio, not its transfer to a new category or its split into trial and repeat.

Worked steps

  1. Exposure of product 25: eligible buyers x awareness x availability x interest x units per buyer = 97,916.
  2. Shared scale from 24 products: 0.8538.
  3. Modelled units: 97,916 x 0.8538 = 83,601.
  4. Held-out MAE over six products: 1,843 units.

Use the idea

Calibrate a launch model on as many comparable mature products as you have, hold some out before fitting, and report the held-out error. Do not tune each product until it fits.

Where the conclusion applies

The 30 products are synthetic, built with a planted scale of 0.85 and 6 percent noise, so recovering about 0.85 shows the mechanics, not that real markets obey one scale. Using the scale for trial conversion on a new product is a further assumption the holdout does not test.

Check your understanding: A held-out product has exposure 50,000 and observed 24-month sales of 41,000. With a shared scale of 0.85, what is the model's error?
50,000 x 0.85 = 42,500 modelled units, so the error is 42,500 - 41,000 = 1,500 units.

Chapter 27 source: section "Section Four: The Build-Up Problem".

Demonstration 3 of 4

The build-up: when trial happens decides year one

If two launches win the same number of triers over two years, will they sell the same in year one?

Teal bars are first purchases, spread over 24 months by the gamma timing curve; grey bars add repeat purchases from every cohort that has had time to buy again. A slower peak moves trials into year two, and those late triers also have less time to repeat inside the horizon.

Equation: the share of trials in month m is F of m minus F of m minus 1, divided by F of 24

Equation: shape k minus 1, times scale theta, equals 4, so the trial rate peaks in month four

Equation: units in month m equal trials in month m plus 0.15 times the sum of trials in all earlier months

Scroll sideways for the whole equation

F is the gamma cumulative distribution of trial timing with shape k and scale theta (months); s is the share of the declared trial total in month m; T is trials in month m and U all units that month: each trier buys one unit at trial and 0.15 units a month after.

Predict first. Slip the peak from month 4 to month 5, with repeat at 0.15. Do year-one units fall by more or less than 10 percent?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The build-up: when trial happens decides year one. Monthly bars over 24 months: trial (teal) peaks in month 4; grey bars add repeat purchases on top. Year-one trials 10,442, year-one units 19,398.
Month trial peaks: 4, standard, What the trial total counts: All trials within 24 months, Repeat units per trier per month: 0.15, as in the notebook
Constructed data: the chapter notebook's synthetic launch (120,000 eligible buyers, awareness 0.65, availability 0.70, interest 0.28, seed 20260945), its gamma timing and its repeat assumption, computed with the companion's tools.

Calculated values

Trial total declared
13,052
Year-one trials
10,442
Year-one share of the declared total
0.800
Trials after month 24
0 (all allocated within 24 months)
Year-one units, trial plus repeat
19,398
24-month units
43,726

Peak month 4: year-one trials are 10,442 of 13,052, and 10,442 / 13,052 = 0.800; all 13,052 declared trials fall inside 24 months. Repeat adds 19,398 - 10,442 = 8,956 units in year one, more when trial comes early, because early triers have longer to buy again.

Worked steps

  1. Reach: 120,000 x 0.65 x 0.70 x 0.28 = 15,288 buyers; times the shared scale 0.8538, assumed to transfer, gives a declared 13,052 triers.
  2. Gamma timing with peak month 4: year-one trials 10,442.
  3. Year-one share: 10,442 / 13,052 = 0.800.
  4. Year-one units including repeat: 19,398; 24 months: 43,726.

Use the idea

State the denominator of any launch timing figure: 80 percent of a two-year total is not 80 percent of an eventual total. Keep trial and repeat apart, and let a slower build lower year one.

Where the conclusion applies

Peak months 3, 4 and 5 and the 80 percent year-one share are the author's practitioner conventions, not measured laws; the repeat rate of 0.15 units a month is an illustrative assumption. The timing curve is a trial-timing curve for the first two years, not the Bass adoption model of Chapter 19.

Check your understanding: A declared two-year trial total is 4,000 under the standard timing (80 percent in year one). How many trials fall in each year?
4,000 x 0.8 = 3,200 in year one and 4,000 - 3,200 = 800 in year two.

Chapter 27 source: section "Section Four: The Build-Up Problem".

Demonstration 4 of 4

Two engines: what agreement proves and what a gap points at

When a bottom-up and a top-down launch forecast agree, what has that agreement shown?

The bottom-up engine uses penetration; the top-down engine uses category volume. Everything else (awareness, trial, distribution) they share. When the two category inputs describe different categories, the bars split. When both are moved together, the bars agree again, at a different level.

Equation: the gap equals the absolute difference between bottom up B and top down T, divided by their average

Scroll sideways for the whole equation

B is the bottom-up forecast, built from the population through category penetration, trial and repeat; T is the top-down forecast, built from the category's unit volume through a share of choice. Both are year-one units in millions.

Predict first. Type category penetration of 40 percent while leaving category volume at 415 million. Will the engines still agree within 10 percent?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Two engines: what agreement proves and what a gap points at. Two bars: bottom-up 3.545 and top-down 3.525 million year-one units, a gap of 0.6 percent, with penetration 53 percent and category volume 415 million units.
Category penetration typed: 53 percent, as in the example, Category volume typed (million units a year): 415, as in the example
Constructed data: the chapter's illustrative launch inputs (118 million households, 53 percent penetration, 415 million units a year, 63 percent distribution) run through the companion's top-down bottom-up triangulation model, with penetration and category volume varied.

Calculated values

Category penetration typed
53 percent
Category volume typed (million units)
415
Bottom-up units (million)
3.545
Top-down units (million)
3.525
Gap
0.6 percent
Penetration the top-down engine implies
52.7 percent

Gap = (3.545 - 3.525) / ((3.545 + 3.525) / 2) = 0.0057, or 0.6 percent, so the engines agree. They share awareness, trial and distribution, so this says the inputs describe one coherent category, not that the launch will sell this much.

Worked steps

  1. Bottom-up builds from penetration 53 percent of 118 million households: 3.545 million units.
  2. Top-down starts from category volume 415 million units: 3.525 million units.
  3. Gap: (3.545 - 3.525) / 3.5350 = 0.0057.
  4. Under 10 percent: consistent inputs. Over: find the wrong input; do not average.

Use the idea

Run any launch forecast two ways and treat a gap above your threshold as a pointer to a wrong input, not as a range to average. Treat agreement as a consistency check only.

Where the conclusion applies

The model's weights are illustrative starting values, not measured elasticities, and the 10 percent threshold is a working convention, not a statistical test. Because the routes share most inputs, agreement cannot validate the forecast; only scoring against the actual launch can.

Common wrong turn: Two routes that agree have validated the number
The chapter says the two routes share awareness, trial and distribution, so their agreement is an internal consistency check, not independent validation. With penetration 70 percent and volume 550 million they agree to 0.2 percent at 4.68 million units, a different forecast.
Check your understanding: One engine says 4.0 million units and the other 3.0 million. What is the gap, and what should you do?
(4.0 - 3.0) / ((4.0 + 3.0) / 2) = 1.0 / 3.5 = 0.286, a gap of 28.6 percent: above 10 percent, so find the input that is wrong instead of averaging.

Chapter 27 source: section "Section Four: The Build-Up Problem".