Demonstration 1 of 4
Trend, weekly cycle and campaign add up to the forecast
What does it buy an analyst that the forecast is a sum of named pieces?
The model is fitted once, on days 0 to 374. On any later day it reports the trend, the weekly effect for that weekday and the campaign effect for that date; added together they give exactly the point forecast. The waterfall on the right builds the forecast one component at a time.
Scroll sideways for the whole equation
y(t) is activity on day t; g(t) is the trend, s(t) the weekly seasonal effect, h(t) the effect of the campaign calendar (the chapter's holiday component), and epsilon the error. y-hat(t) is the point forecast, the sum of the three fitted components, all in units of the activity index.
Predict first. Day 385 is the first day of a scheduled campaign. Will the campaign component be close to the 22 points built into the data?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded synthetic activity series (seed 20260934) and its scheduled campaign calendar, fitted as in the notebook's component audit.
Calculated values
- Day
- 385 (Monday)
- Trend
- 132.73
- Weekly
- -0.23
- Campaign
- 21.62
- Forecast, the sum
- 154.12
- Actual
- 151.67
- Actual minus forecast
- -2.45
Day 385 is a campaign day, so the campaign component is 21.62. Forecast = 132.73 + (-0.23) + 21.62 = 154.12; the actual value was 151.67, an error of 151.67 - 154.12 = -2.45. Each piece can be checked on its own: a campaign effect on the wrong day or a weekly swing of the wrong size would show here before anyone trusts the total.
Worked steps
- Trend on day 385: 132.73.
- Add the weekly effect for a Monday: 132.73 + (-0.23) = 132.50.
- Add the campaign effect: 132.50 + 21.62 = 154.12.
- Error: 151.67 - 154.12 = -2.45.
Use the idea
When a forecast looks wrong, read its components before arguing with the total: a trend that kept climbing, a weekly swing of the wrong size or a campaign on the wrong day each point to a different fix.
Where the conclusion applies
The data were generated with the same additive structure the model assumes (a bent trend, a weekly sine wave, a 22 point two-day campaign every 45 days, noise with standard deviation 2.4), so the components are recovered well. On a real series the components are fitted explanations, not observed causes, and a perfect sum checks the arithmetic, not whether the calendar was right.
Check your understanding: In additive mode a trend of 80, a weekly effect of +4 and a holiday effect of -10 give what forecast?
Chapter 16 source: section "Section Two: What Prophet Got Right".
Demonstration 2 of 4
Choose trend flexibility on earlier origins
How should the changepoint prior scale be set without peeking at the final test?
Each state refits the model on the days before an origin and forecasts the next 45. The stiffest trend cannot follow the faster growth that starts at day 200; the most flexible one bends more than it needs to. The prior scale with the lowest average over both origins is frozen before the final test.
Scroll sideways for the whole equation
y is the activity on day t and y-hat its forecast; MAE is the mean absolute error over the n = 45 days after a forecast origin. The changepoint prior scale sets how freely the trend may change slope: a small scale is strong regularization, a large one weak regularization.
Predict first. Which prior scale will have the lowest average validation MAE over origins 285 and 330?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded synthetic activity series, with the model refitted at the notebook's two validation origins and three prior scales.
Calculated values
- Changepoint prior scale
- 0.05
- Validation MAE, origin 285
- 2.041
- Validation MAE, origin 330
- 1.868
- Average of the two origins
- 1.9545
- Selected prior scale
- 0.05
Fitted on days 0 to 284 with prior scale 0.05, the forecast of the next 45 days has MAE 2.041. Over both origins the average is (2.041 + 1.868) / 2 = 1.9545, the lowest average of the three, so validation selects it. The trend bends at the change in growth without chasing noise.
Worked steps
- MAE at origin 285: 2.041; at origin 330: 1.868.
- Average: (2.041 + 1.868) / 2 = 1.9545.
- Averages for 0.001, 0.05 and 0.5: 11.2545, 1.9545, 2.1475.
- Lowest average: prior scale 0.05, fixed before the final 45 days are scored.
Use the idea
Tune any flexibility setting with rolling origins that end before the final test, and record the choice before the test is scored, as the chapter's section on hyperparameter optimization asks.
Where the conclusion applies
Two origins and three candidate scales are the notebook's choices; with more origins the averages would be steadier. Validation can mislead when the validation blocks are calmer than the future: a break after day 375 would not be seen by any of these scores.
Check your understanding: If a fourth prior scale scored 1.900 at origin 285 and 2.030 at origin 330, would it beat 0.05?
Chapter 16 source: section "Section Four: The Full Methodology".
Demonstration 3 of 4
An 80 percent band is a claim to check
Does a band labelled 80 percent contain the outcome 80 percent of the time?
The model's interval comes from its maximum a posteriori fit plus simulated future trend changes and observation noise. Whether that band covers 80 percent of outcomes is an empirical question, answered by counting outcomes inside it on data the model never saw.
Scroll sideways for the whole equation
L and U are the lower and upper ends of the nominal 80 percent band on day t, y the outcome, and the indicator 1(...) is 1 when the outcome lies inside the band and 0 otherwise. Coverage is the share of the n scored days inside.
Predict first. Over all 45 final test days, will measured coverage come out exactly 0.800?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded synthetic activity series and the final fit with 300 interval samples (global seed 16), scored on the untouched last 45 days.
Calculated values
- Days scored
- 45
- Inside the band
- 37
- Measured coverage
- 0.822
- Nominal coverage
- 0.800
- Average band width
- 6.43
Over all 45 test days, 37 of 45 outcomes fall inside the nominal 80 percent band: coverage = 37 / 45 = 0.822, above the nominal 0.800. One holdout of 45 days is a single measurement, not a calibration certificate.
Worked steps
- Outcomes inside the band: 37 of 45.
- Coverage: 37 / 45 = 0.822.
- Compare with the nominal 0.800: 0.822 - 0.800 = 0.022.
Use the idea
Report the measured coverage of a held-out period next to any interval, and treat a band whose coverage has never been measured as a modelling statement, not a guarantee.
Where the conclusion applies
The notebook draws 300 interval samples with seed 16; another seed gives slightly different bands. This fit treats the seasonal pattern as known, and the residuals are assumed independent, so correlated errors or an unannounced break can defeat the band, as the chapter warns.
What this does not settle
The chapter notes that full Bayesian sampling is needed for uncertainty in the seasonal components; this band treats them as known. One 45-day holdout on synthetic data that share the model's own structure cannot certify calibration on real series.
Chapter 16 source: "required to obtain uncertainty estimates in the seasonality components".
Check your understanding: A 30-day holdout has 21 outcomes inside a nominal 80 percent band. What is the measured coverage, and how far is it from nominal?
Chapter 16 source: section "Section Four: The Full Methodology".
Demonstration 4 of 4
Test whether the analyst's calendar helps
Does adding the analyst's known event improve later forecasts, and what happens when that knowledge goes stale?
Both models were tuned on the same earlier origins and scored once on the untouched last 45 days. The calendar's whole contribution sits on the campaign days; on the other days the two forecasts are close. Moving the last campaign a week later, without updating the calendar, turns that contribution into two wrong days.
Scroll sideways for the whole equation
y is the outcome on day t and y-hat the forecast; MAE is the mean absolute error over the n days scored. Each model chose its own changepoint prior scale on the earlier origins (both chose 0.05).
Predict first. With the campaign on schedule, over all 45 test days, which model has the lower MAE?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded synthetic activity series and its calendar ablation; the a-week-late campaign follows the notebook's exercise, with the same seeded noise.
Calculated values
- Campaign timing
- as scheduled
- Model
- with calendar
- Days scored
- 45
- Sum of absolute errors
- 78.88
- MAE
- 1.753
- MAE of the other model, same days
- 2.637
Scored on all 45 test days, the model with the calendar has MAE = 78.88 / 45 = 1.753, lower than the 2.637 of the model without it. The calendar told the model the right days, and the model sized the effect from earlier campaigns.
Worked steps
- Sum of absolute errors over all 45 test days: 78.88.
- MAE: 78.88 / 45 = 1.753.
- Other model on the same days: 2.637; difference 1.753 - 2.637 = -0.884.
Use the idea
Keep an ablation beside any forecast that uses business knowledge: score it with and without the calendar on later data, record the calendar's date of issue, and check scheduled events against what actually happened.
Where the conclusion applies
The late campaign is the notebook's own exercise, applied after the fact: training data end at day 375, before that campaign, so both fitted models are the notebook's and only the outcomes on days 385 to 393 move. The data share the model's additive structure, so the calendar's gain here is not evidence about real series.
Check your understanding: Over 4 campaign days a model's absolute errors sum to 86.40. What is its MAE on those days?
Chapter 16 source: section "Section Five: The Distribution of Responsibility".