Demonstration 1 of 4
Difference the changing level
Why model the changes of a wandering series instead of its level?
The level adds up every past change, so it drifts and its mean depends on which stretch you look at. Differencing undoes that accumulation: the changes have a stable mean, and their remaining dependence (phi) is what the AR part of ARIMA describes.
Scroll sideways for the whole equation
The first difference is the value in period t minus the value in the period before. The changes follow an autoregression: each change is phi times the previous change plus a fresh random shock. phi is the AR coefficient.
Predict first. Switch from the level to the first difference. Will the two half means of the series move closer together?
Choose an example
Scroll sideways for the whole figure
Constructed data: notebook 06 cell 1, an integrated autoregression with phi 0.65; phi 0 and 0.9 reuse the same shocks.
Calculated values
- Level: mean of first half
- 90.3
- Level: mean of second half
- 99.8
- Change: mean of first half
- -0.11
- Change: mean of second half
- 0.03
- Lag 1 autocorrelation of level
- 0.98
- Lag 1 autocorrelation of change
- 0.63
The level's half means differ by 99.8 - 90.3 = 9.5: the mean wanders, so the level is not stationary. The changes have half means -0.11 and 0.03, both near zero. With phi 0.65, the changes keep a lag 1 autocorrelation of 0.63, which is the structure an AR term models after one difference.
Worked steps
- First difference: each change is the value minus the value before, for example 100.90 - 100.00 = 0.90.
- Level half means: 90.3 and 99.8, a gap of 9.5.
- Change half means: -0.11 and 0.03.
Use the idea
Before fitting an ARIMA model, plot the level and its first difference, and difference only when the level's mean is clearly not stable.
Where the conclusion applies
The series is the notebook's constructed integrated autoregression (seed 20260924), rebuilt from the same shocks for each phi. Differencing a series that is already stationary adds structure instead of removing it, so not every trend should be differenced.
Check your understanding: A series reads 100, 103, 101, 106. What are its first differences?
Chapter 6 source: section "Part Three: What ARIMA Actually Is".
Demonstration 2 of 4
A signature in the lag structure
How do the ACF and PACF tell an autoregression apart from noise?
For an AR(1), the PACF has one spike at lag 1 and then nothing, while the ACF decays gradually because each change echoes through the next. With phi 0 the changes are white noise, yet lag 11 still crosses the band: a chance spike.
Scroll sideways for the whole equation
The ACF at lag k is the correlation of the changes with themselves k periods back; the PACF removes the effect of shorter lags first. n is the number of differences (189). Bars inside the band are not distinguishable from zero at the 5 percent level.
Predict first. At phi 0.65, which PACF lags stand far outside the band?
Choose an example
Scroll sideways for the whole figure
Constructed data: notebook 06 cells 1 and 3, the ACF and PACF of the training differences; phi 0 and 0.9 reuse the same shocks.
Calculated values
- Differences used
- 189
- Band
- plus or minus 0.143
- PACF at lag 1
- 0.626
- PACF at lag 2
- -0.060
- Lags outside the band
- 1, 6, 10
The band is 1.96 / sqrt(189) = 1.96 / 13.75 = 0.143. The PACF at lag 1 is 0.626 and at lag 2 is -0.060; lags outside the band: 1, 6, 10. An AR(1) process should show a PACF that cuts off after lag 1 and an ACF that tails off. A lone spike at a high lag can be sampling noise: one in twenty lags crosses a 5 percent band by chance.
Worked steps
- Band: 1.96 / 13.75 = 0.143, since sqrt(189) is 13.75.
- PACF at lag 1: 0.626; at lag 2: -0.060.
- Lags beyond the band: 1, 6, 10.
Use the idea
Read the ACF and PACF of the stationary (differenced) series to propose candidate orders, then let the residual check and holdout errors decide.
Where the conclusion applies
Only training differences (189 values) are used. The band assumes white noise and is approximate; it screens candidates and does not select an order with certainty.
Common wrong turn: Every bar that crosses the band is real structure
Check your understanding: With 100 observations, how wide is the 5 percent band on an ACF plot?
Chapter 6 source: section "Part Four: The Full Methodology".
Demonstration 3 of 4
Seasonal ARIMA against seasonal naive, origin by origin
Does the seasonal model beat the simplest seasonal benchmark when scored honestly over time?
At each origin all three models see only earlier data and forecast the next twelve points. Seasonal naive wins at origins 120 and 144, SARIMA at 168. The ranking depends on the origin, so only the average over origins counts.
Scroll sideways for the whole equation
The series depends on its value twelve periods earlier, with seasonal AR coefficient capital phi 0.85 in the generator. SARIMA(0,0,0)(1,0,0) with period 12 models that; ARIMA(1,0,0) sees only lag 1; seasonal naive repeats the last twelve values.
Predict first. Averaged over three origins, will SARIMA beat seasonal naive by a wide margin?
Choose an example
Scroll sideways for the whole figure
Constructed data: notebook 06 cell 7, a series with lag-12 dependence (coefficient 0.85, seed 20260924) scored at origins 120, 144 and 168.
Calculated values
- Mean MAE, SARIMA
- 1.85
- Mean MAE, ARIMA
- 3.42
- Mean MAE, Seasonal naive
- 1.86
Averaged over the three origins, SARIMA's MAE is (1.30 + 1.82 + 2.42) / 3 = 1.85, against 1.86 for seasonal naive and 3.42 for nonseasonal ARIMA. SARIMA is lowest, but SARIMA and seasonal naive are within 0.01 of each other; the nonseasonal model, which cannot see lag 12, trails both.
Worked steps
- SARIMA: (1.30 + 1.82 + 2.42) / 3 = 1.85.
- Seasonal naive mean: 1.86.
- Nonseasonal ARIMA mean: 3.42.
Use the idea
Compare any seasonal model with seasonal naive over several chronological origins before believing it adds anything.
Where the conclusion applies
Orders are fixed from the generator, not selected on the scored outcomes; real order selection must happen inside each training window. Three origins are few, and a margin of 0.01 is not a meaningful difference.
Check your understanding: A model's twelve-step MAEs at three origins are 1.30, 1.82 and 2.42. What is its mean?
Chapter 6 source: section "Part Four: The Full Methodology".
Demonstration 4 of 4
Listen to the residuals
How do you know an ARIMA model has said everything the data allows?
The random walk forecast is flat and its residuals are the raw changes, which are autocorrelated. One AR term absorbs that memory: residuals turn white and the forecast bends toward the drift. A second AR term adds nothing the check or the holdout rewards.
Scroll sideways for the whole equation
ARIMA(p,d,q): p autoregressive lags, d differences, q moving-average lags. The residual check is a Ljung-Box test on 12 lags of the residuals: a small p-value means autocorrelation remains. MAE is mean absolute error on the held-out 30 points.
Predict first. ARIMA(0,1,0) is the random walk, with no AR term. Will its residuals pass the white noise check?
Choose an example
Scroll sideways for the whole figure
Constructed data: notebook 06 cell 5, ARIMA(1,1,0) on the constructed series (MAE 6.13 against persistence 6.82); orders (0,1,0) and (2,1,0) are fitted the same way.
Calculated values
- Model
- ARIMA(1,1,0)
- Residual check p-value (12 lags)
- 0.279
- Residuals look like white noise
- yes
- Holdout MAE, first 30 steps
- 6.13
- Persistence MAE, first 30 steps
- 6.82
ARIMA(1,1,0): the residual check gives p 0.279, so the residuals show no autocorrelation the test can detect. Over the first 30 held-out steps, persistence MAE minus model MAE is 6.82 - 6.13 = 0.69 in the model's favour.
Worked steps
- Residual check (Ljung-Box, 12 lags): p 0.279, pass.
- Model MAE over 30 steps: 6.13.
- Persistence MAE: 6.82.
- Difference: 6.82 - 6.13 = 0.69.
Use the idea
Fit, then test the residuals before reading the forecast; if the check fails, return to identification as Box and Jenkins prescribe.
Where the conclusion applies
The order (1,1,0) matches how the data were generated, which real data never tells you. White residuals do not prove all predictability (for example nonlinear) has been captured, and one 30-point path is not a calibration test of the band.
Check your understanding: Over 30 steps, a model's MAE is 6.13 and persistence's is 6.82. By how much does the model improve on persistence?
Chapter 6 source: section "Part Five: Listening to the Residuals".