The Art and Science of Forecasting

Chapter 16

The Prophet

A forecast built from named pieces lets the analyst add what she knows, and obliges her to test it.

Four demonstrations follow the chapter on a synthetic activity series with a scheduled campaign: how trend, weekly cycle and campaign add up to a forecast, how trend flexibility is chosen on earlier origins, how often a nominal 80 percent band actually covers the outcome, and whether the analyst's calendar improves a held-out forecast, on schedule and when it goes stale.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Trend, weekly cycle and campaign add up to the forecast

What does it buy an analyst that the forecast is a sum of named pieces?

The model is fitted once, on days 0 to 374. On any later day it reports the trend, the weekly effect for that weekday and the campaign effect for that date; added together they give exactly the point forecast. The waterfall on the right builds the forecast one component at a time.

Equation: y of t equals the trend g of t plus the seasonality s of t plus the holiday effect h of t plus the error epsilon t

Equation: the point forecast y hat of t equals the fitted trend plus the fitted seasonality plus the fitted holiday effect

Scroll sideways for the whole equation

y(t) is activity on day t; g(t) is the trend, s(t) the weekly seasonal effect, h(t) the effect of the campaign calendar (the chapter's holiday component), and epsilon the error. y-hat(t) is the point forecast, the sum of the three fitted components, all in units of the activity index.

Predict first. Day 385 is the first day of a scheduled campaign. Will the campaign component be close to the 22 points built into the data?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Trend, weekly cycle and campaign add up to the forecast. Left: the final test period with the forecast line and day 385 circled. Right: trend 132.73, weekly -0.23 and campaign 21.62 stacked into the forecast 154.12, beside the actual 151.67.
Day of the final test period: Day 385 (campaign, first day)
Constructed data: the chapter notebook's seeded synthetic activity series (seed 20260934) and its scheduled campaign calendar, fitted as in the notebook's component audit.

Calculated values

Day
385 (Monday)
Trend
132.73
Weekly
-0.23
Campaign
21.62
Forecast, the sum
154.12
Actual
151.67
Actual minus forecast
-2.45

Day 385 is a campaign day, so the campaign component is 21.62. Forecast = 132.73 + (-0.23) + 21.62 = 154.12; the actual value was 151.67, an error of 151.67 - 154.12 = -2.45. Each piece can be checked on its own: a campaign effect on the wrong day or a weekly swing of the wrong size would show here before anyone trusts the total.

Worked steps

  1. Trend on day 385: 132.73.
  2. Add the weekly effect for a Monday: 132.73 + (-0.23) = 132.50.
  3. Add the campaign effect: 132.50 + 21.62 = 154.12.
  4. Error: 151.67 - 154.12 = -2.45.

Use the idea

When a forecast looks wrong, read its components before arguing with the total: a trend that kept climbing, a weekly swing of the wrong size or a campaign on the wrong day each point to a different fix.

Where the conclusion applies

The data were generated with the same additive structure the model assumes (a bent trend, a weekly sine wave, a 22 point two-day campaign every 45 days, noise with standard deviation 2.4), so the components are recovered well. On a real series the components are fitted explanations, not observed causes, and a perfect sum checks the arithmetic, not whether the calendar was right.

Check your understanding: In additive mode a trend of 80, a weekly effect of +4 and a holiday effect of -10 give what forecast?
80 + 4 + (-10) = 74.

Chapter 16 source: section "Section Two: What Prophet Got Right".

Demonstration 2 of 4

Choose trend flexibility on earlier origins

How should the changepoint prior scale be set without peeking at the final test?

Each state refits the model on the days before an origin and forecasts the next 45. The stiffest trend cannot follow the faster growth that starts at day 200; the most flexible one bends more than it needs to. The prior scale with the lowest average over both origins is frozen before the final test.

Equation: mean absolute error equals the average of the absolute differences between y t and its forecast

Scroll sideways for the whole equation

y is the activity on day t and y-hat its forecast; MAE is the mean absolute error over the n = 45 days after a forecast origin. The changepoint prior scale sets how freely the trend may change slope: a small scale is strong regularization, a large one weak regularization.

Predict first. Which prior scale will have the lowest average validation MAE over origins 285 and 330?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Choose trend flexibility on earlier origins. Activity from day 200 with the forecast fitted to day 284 at prior scale 0.05; the shaded 45-day validation block has MAE 2.041.
Changepoint prior scale: 0.05, Validation origin: Day 285
Constructed data: the chapter notebook's seeded synthetic activity series, with the model refitted at the notebook's two validation origins and three prior scales.

Calculated values

Changepoint prior scale
0.05
Validation MAE, origin 285
2.041
Validation MAE, origin 330
1.868
Average of the two origins
1.9545
Selected prior scale
0.05

Fitted on days 0 to 284 with prior scale 0.05, the forecast of the next 45 days has MAE 2.041. Over both origins the average is (2.041 + 1.868) / 2 = 1.9545, the lowest average of the three, so validation selects it. The trend bends at the change in growth without chasing noise.

Worked steps

  1. MAE at origin 285: 2.041; at origin 330: 1.868.
  2. Average: (2.041 + 1.868) / 2 = 1.9545.
  3. Averages for 0.001, 0.05 and 0.5: 11.2545, 1.9545, 2.1475.
  4. Lowest average: prior scale 0.05, fixed before the final 45 days are scored.

Use the idea

Tune any flexibility setting with rolling origins that end before the final test, and record the choice before the test is scored, as the chapter's section on hyperparameter optimization asks.

Where the conclusion applies

Two origins and three candidate scales are the notebook's choices; with more origins the averages would be steadier. Validation can mislead when the validation blocks are calmer than the future: a break after day 375 would not be seen by any of these scores.

Check your understanding: If a fourth prior scale scored 1.900 at origin 285 and 2.030 at origin 330, would it beat 0.05?
(1.900 + 2.030) / 2 = 1.965, which is above 1.9545, so 0.05 still has the lowest average.

Chapter 16 source: section "Section Four: The Full Methodology".

Demonstration 3 of 4

An 80 percent band is a claim to check

Does a band labelled 80 percent contain the outcome 80 percent of the time?

The model's interval comes from its maximum a posteriori fit plus simulated future trend changes and observation noise. Whether that band covers 80 percent of outcomes is an empirical question, answered by counting outcomes inside it on data the model never saw.

Equation: coverage equals the share of the n days on which the outcome y t lies between the lower band L t and the upper band U t

Scroll sideways for the whole equation

L and U are the lower and upper ends of the nominal 80 percent band on day t, y the outcome, and the indicator 1(...) is 1 when the outcome lies inside the band and 0 otherwise. Coverage is the share of the n scored days inside.

Predict first. Over all 45 final test days, will measured coverage come out exactly 0.800?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: An 80 percent band is a claim to check. Final 45 test days with the forecast, its shaded 80 percent band and the outcomes; all 45 test days highlighted, 37 of 45 inside.
Days scored: All 45 test days
Constructed data: the chapter notebook's seeded synthetic activity series and the final fit with 300 interval samples (global seed 16), scored on the untouched last 45 days.

Calculated values

Days scored
45
Inside the band
37
Measured coverage
0.822
Nominal coverage
0.800
Average band width
6.43

Over all 45 test days, 37 of 45 outcomes fall inside the nominal 80 percent band: coverage = 37 / 45 = 0.822, above the nominal 0.800. One holdout of 45 days is a single measurement, not a calibration certificate.

Worked steps

  1. Outcomes inside the band: 37 of 45.
  2. Coverage: 37 / 45 = 0.822.
  3. Compare with the nominal 0.800: 0.822 - 0.800 = 0.022.

Use the idea

Report the measured coverage of a held-out period next to any interval, and treat a band whose coverage has never been measured as a modelling statement, not a guarantee.

Where the conclusion applies

The notebook draws 300 interval samples with seed 16; another seed gives slightly different bands. This fit treats the seasonal pattern as known, and the residuals are assumed independent, so correlated errors or an unannounced break can defeat the band, as the chapter warns.

What this does not settle

The chapter notes that full Bayesian sampling is needed for uncertainty in the seasonal components; this band treats them as known. One 45-day holdout on synthetic data that share the model's own structure cannot certify calibration on real series.

Chapter 16 source: "required to obtain uncertainty estimates in the seasonality components".

Check your understanding: A 30-day holdout has 21 outcomes inside a nominal 80 percent band. What is the measured coverage, and how far is it from nominal?
21 / 30 = 0.700, which is 0.800 - 0.700 = 0.100 below nominal.

Chapter 16 source: section "Section Four: The Full Methodology".

Demonstration 4 of 4

Test whether the analyst's calendar helps

Does adding the analyst's known event improve later forecasts, and what happens when that knowledge goes stale?

Both models were tuned on the same earlier origins and scored once on the untouched last 45 days. The calendar's whole contribution sits on the campaign days; on the other days the two forecasts are close. Moving the last campaign a week later, without updating the calendar, turns that contribution into two wrong days.

Equation: mean absolute error equals the average of the absolute differences between y t and its forecast

Scroll sideways for the whole equation

y is the outcome on day t and y-hat the forecast; MAE is the mean absolute error over the n days scored. Each model chose its own changepoint prior scale on the earlier origins (both chose 0.05).

Predict first. With the campaign on schedule, over all 45 test days, which model has the lower MAE?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Test whether the analyst's calendar helps. Final 45 test days: outcomes with the campaign as scheduled, and the forecast with the calendar; MAE 1.753 over all 45 test days.
Last campaign: As scheduled (days 385 and 386), Model: With the campaign calendar, Days scored: All 45 test days
Constructed data: the chapter notebook's seeded synthetic activity series and its calendar ablation; the a-week-late campaign follows the notebook's exercise, with the same seeded noise.

Calculated values

Campaign timing
as scheduled
Model
with calendar
Days scored
45
Sum of absolute errors
78.88
MAE
1.753
MAE of the other model, same days
2.637

Scored on all 45 test days, the model with the calendar has MAE = 78.88 / 45 = 1.753, lower than the 2.637 of the model without it. The calendar told the model the right days, and the model sized the effect from earlier campaigns.

Worked steps

  1. Sum of absolute errors over all 45 test days: 78.88.
  2. MAE: 78.88 / 45 = 1.753.
  3. Other model on the same days: 2.637; difference 1.753 - 2.637 = -0.884.

Use the idea

Keep an ablation beside any forecast that uses business knowledge: score it with and without the calendar on later data, record the calendar's date of issue, and check scheduled events against what actually happened.

Where the conclusion applies

The late campaign is the notebook's own exercise, applied after the fact: training data end at day 375, before that campaign, so both fitted models are the notebook's and only the outcomes on days 385 to 393 move. The data share the model's additive structure, so the calendar's gain here is not evidence about real series.

Check your understanding: Over 4 campaign days a model's absolute errors sum to 86.40. What is its MAE on those days?
86.40 / 4 = 21.600.

Chapter 16 source: section "Section Five: The Distribution of Responsibility".