Demonstration 1 of 4
The old model keeps predicting the old world
What happens to a model fitted before a structural break once the level moves?
The dashed line is the frozen forecast, the mean of the shaded fitting window. Before period 100 the series scatters around it; after period 100 the whole series sits above it, so every error has the same sign and the average error is roughly the size of the shift.
Scroll sideways for the whole equation
y is the observed value in period t and y-hat the forecast. The frozen forecast is y-bar, the mean of the first fifty periods, never updated. MAE is the mean absolute error over the periods named.
Predict first. With the notebook's level shift of 15, will the frozen model's error after period 100 be more or less than five times its error before?
Choose an example
Scroll sideways for the whole figure
Constructed data: notebook 24 cell 1, seeded draws of 100 periods around 50 and 100 around 65 (seed 20260942); smaller shifts move the second block down by a fixed amount.
Calculated values
- Level shift
- 15
- Shift in calibration standard deviations
- 9.3
- Frozen forecast
- 49.98
- MAE, periods 50 to 99
- 1.69
- MAE, periods 100 to 199
- 14.79
- Error ratio, after over before
- 8.75
The frozen model forecasts 49.98 for ever. Before the break its MAE is 1.69; after a shift of 15 it is 14.79, a ratio of 14.79 / 1.69 = 8.75, so the dashboard would turn red at once. The model is not noisier than before: it is answering a question about a world that ended at period 100.
Worked steps
- Frozen forecast: mean of periods 0 to 49 = 49.98.
- Mean absolute error over periods 50 to 99: 1.69.
- Mean absolute error over periods 100 to 199: 14.79.
- Ratio: 14.79 / 1.69 = 8.75.
Use the idea
When a model's errors start to share one sign, compare recent error with the error the model showed when it was validated, as the chapter's dashboard did, before trusting its next forecast.
Where the conclusion applies
The series is constructed: independent noise around 50 and then around 50 plus the shift, from the notebook's seeded draws (seed 20260942). Real breaks can arrive with changing variance, gradual drift or seasonality, which make the error jump less clean than here.
Check your understanding: A frozen forecast of 100 had an MAE of 4 before a break; afterwards demand settles at 130 with the same noise. Roughly what MAE should you expect after the break?
Chapter 24 source: section "Section One: The Dashboard Turns Red".
Demonstration 2 of 4
An alarm that has to accumulate evidence
How long does a CUSUM detector take to notice a level shift, and what decides it?
Before the shift the residuals average about zero, the 0.5 allowance pulls the statistic back to zero, and it stays low. After a shift of d standard deviations each period adds about d - 0.5, so the statistic climbs like a ramp; a small d makes a slow ramp, a large d a jump. A lower threshold answers sooner but risks false alarms.
Scroll sideways for the whole equation
z is the forecast error in period t divided by s, the standard deviation of the first fifty periods (1.61), after subtracting their mean y-bar. S is the CUSUM statistic, k = 0.5 the allowance subtracted each period, and h the threshold: when S passes h an alarm is raised and S restarts at zero.
Predict first. With the notebook's large shift of 15 and threshold 8, how many periods after the shift is the first alarm?
Choose an example
Scroll sideways for the whole figure
Constructed data: notebook 24 cell 3, the standardised one-sided CUSUM with k 0.5 and threshold 8 on the seeded series (seed 20260942); smaller shifts and threshold 4 are added variations.
Calculated values
- Shift in calibration standard deviations
- 0.93
- Alarms before the shift (false)
- 0
- Alarms after the shift
- 5
- Periods from shift to first alarm
- 8
- Standardised residual at period 100
- 3.04
Each post-shift period adds on average 0.93 - 0.5 = 0.43 to the statistic, against a threshold of 8: the evidence has to pile up for 8 periods before the first alarm. 0 false alarms before period 100 and 5 after. Alarms after a reset repeat one finding, that the old mean no longer fits; they are not separate discoveries.
Worked steps
- Calibration: mean 49.98, standard deviation 1.61 from periods 0 to 49.
- Shift in standard deviations: 1.5 / 1.61 = 0.93.
- Average gain per period after the shift: 0.93 - 0.5 = 0.43.
- First alarm: period 108, 8 periods after the shift.
- False alarms before the shift: 0.
Use the idea
Set the threshold on stable data before monitoring starts, then read the first alarm after a change as the moment the old explanation must earn trust again, not as the decision itself.
Where the conclusion applies
The threshold values are illustrative, as in the notebook, not calibrated to a false-alarm rate; with threshold 4 on independent noise a false alarm is more likely. The detector watches only upward shifts, and the noise is the notebook's seeded draws (seed 20260942).
Check your understanding: Errors after a shift average 1.5 standard deviations, k is 0.5 and h is 8. Roughly how many periods until the first alarm?
Chapter 24 source: section "Section Three: Building for Structural Breaks".
Demonstration 3 of 4
Adaptation policies have different costs
Should a forecast forget the past quickly or remember all of it when a break might come?
A rolling mean drops pre-break values after w periods, so its error falls back within about w periods of the shift. The expanding mean carries the whole old regime and drifts upward slowly. The frozen mean never moves. In the stable stretch the opposite holds a little: fewer values make the rolling mean noisier.
Scroll sideways for the whole equation
Each forecast for period t uses only earlier values. Frozen keeps the mean of the first fifty periods; rolling averages the last w periods; expanding averages every earlier period. The curves are the mean absolute error over the ten most recent periods.
Predict first. With the notebook's rolling window of 20 and shift of 15, which policy has the lowest MAE after the shift?
Choose an example
Scroll sideways for the whole figure
Constructed data: notebook 24 cell 4, one-step frozen, rolling 20 and expanding forecasts on the seeded series (seed 20260942); windows 10 and 40 and the smaller shift are added variations.
Calculated values
- Frozen MAE, stable periods 50 to 99
- 1.69
- Rolling 20 MAE, stable periods 50 to 99
- 1.79
- Expanding MAE, stable periods 50 to 99
- 1.70
- Frozen MAE, periods 100 to 199
- 14.79
- Rolling 20 MAE, periods 100 to 199
- 2.93
- Expanding MAE, periods 100 to 199
- 10.07
After the shift the rolling 20 window gains 10.07 - 2.93 = 7.14 in MAE over the expanding mean; in the stable periods it pays 1.79 - 1.70 = 0.09. Lowest MAE after the shift: rolling 20; in the stable stretch: frozen. Forgetting fast helps after a break and costs a little when nothing has changed.
Worked steps
- Stable periods, rolling 20: MAE 1.79; expanding: 1.70.
- Cost of forgetting in the stable stretch: 1.79 - 1.70 = 0.09.
- After the shift, rolling 20: 2.93; expanding: 10.07; frozen: 14.79.
- Gain from forgetting after the shift: 10.07 - 2.93 = 7.14.
Use the idea
Choose how fast a production model forgets by comparing its error in stable stretches with its error after a known change, as the chapter's monitoring framework recommends, rather than by its error in one period.
Where the conclusion applies
One upward level shift on independent noise from the notebook's seeded draws (seed 20260942). A temporary outlier or slow drift would rank the policies differently, as the notebook's exercise suggests.
Check your understanding: A rolling mean of 10 periods meets a level shift of 20. Roughly how many periods after the shift is the rolling forecast fully at the new level?
Chapter 24 source: section "Section Four: The Full Methodology".
Demonstration 4 of 4
A real break, and a calibration window that must be stable
On a real series with a known shift, what do the detector and the adaptation policies do, and what if the stable window is not stable?
Binary segmentation dates a single fall in the mean at 1899. A detector calibrated on 1871 to 1890 sees the low flows after the break as large negative errors and alarms within a few years, then again after each reset because the frozen reference never moves. Calibrating on 40 years puts the break inside the reference, which is then no longer a single regime.
Scroll sideways for the whole equation
The detector is a two-sided version of S: one statistic for upward and one for downward shifts, standardised by the calibration years' mean and standard deviation, with its threshold set by simulation for a 5 percent false-alarm chance over the monitored years. MAE is scored on the years after the calibration window.
Predict first. With the notebook's 20 calibration years and rolling window 20, which policy has the lowest MAE on the Nile?
Choose an example
Scroll sideways for the whole figure
Real data: annual flow of the Nile at Aswan, 1871 to 1970, in hundreds of millions of cubic metres, a public-domain series first analysed by Cobb (1978) and bundled with the companion, annual-nile series with timestamps; all numbers come from the companion's break tool run as in the notebook's observed-data cell, with the calibration and window varied.
Calculated values
- Dated changepoint
- 1899
- CUSUM alarms
- 11
- Years from break to first alarm
- 4
- Frozen MAE
- 213.7
- Rolling 20 MAE
- 113.2
- Expanding MAE
- 144.7
- Reset after alarm MAE
- 118.3
Frozen MAE minus rolling 20 MAE is 213.7 - 113.2 = 100.5; the lowest MAE here is rolling 20. The calibration window ends in 1890, before the break, and the first alarm comes 4 years after it. One river is one case: it shows the mechanism on real data, not that a policy works in general.
Worked steps
- Calibration: 1871 to 1890 (20 years).
- Dated changepoint: 1899.
- Frozen MAE 213.7, rolling 20 113.2, expanding 144.7, reset after alarm 118.3.
- Frozen minus rolling: 213.7 - 113.2 = 100.5.
Use the idea
Before monitoring a series, choose a calibration window you believe is one regime, check that no dated break falls inside it, and ask whether the dated break matches an event the business or the record knows about.
Where the conclusion applies
The changepoint search looks for shifts in the mean, not the variance; the threshold assumes independent normal errors. The dates are candidates for explanation, not explanations.
Check your understanding: A frozen model scores MAE 214 and a rolling model 113 on the same years. By how much does forgetting reduce the error, as a share of the frozen error?
Chapter 24 source: section "Section Four: The Full Methodology".