The Art and Science of Forecasting

Chapter 6

The Unexpected Route

Box and Jenkins turned forecasting into a cycle: identify, estimate, check the residuals, and look again.

Four demonstrations follow the Box-Jenkins cycle: differencing a wandering level, reading the ACF and PACF for an order, scoring a seasonal model against seasonal naive over several origins, and checking whether the residuals are white noise.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Difference the changing level

Why model the changes of a wandering series instead of its level?

The level adds up every past change, so it drifts and its mean depends on which stretch you look at. Differencing undoes that accumulation: the changes have a stable mean, and their remaining dependence (phi) is what the AR part of ARIMA describes.

Equation: the first difference of y at t equals y t minus y t minus 1

Equation: the change at t equals phi times the previous change plus a shock

Scroll sideways for the whole equation

The first difference is the value in period t minus the value in the period before. The changes follow an autoregression: each change is phi times the previous change plus a fresh random shock. phi is the AR coefficient.

Predict first. Switch from the level to the first difference. Will the two half means of the series move closer together?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Difference the changing level. The level of a constructed series with phi 0.65; level half means 90.3 and 99.8, change half means -0.11 and 0.03.
AR coefficient of the changes, phi: 0.65, What is plotted: The level
Constructed data: notebook 06 cell 1, an integrated autoregression with phi 0.65; phi 0 and 0.9 reuse the same shocks.

Calculated values

Level: mean of first half
90.3
Level: mean of second half
99.8
Change: mean of first half
-0.11
Change: mean of second half
0.03
Lag 1 autocorrelation of level
0.98
Lag 1 autocorrelation of change
0.63

The level's half means differ by 99.8 - 90.3 = 9.5: the mean wanders, so the level is not stationary. The changes have half means -0.11 and 0.03, both near zero. With phi 0.65, the changes keep a lag 1 autocorrelation of 0.63, which is the structure an AR term models after one difference.

Worked steps

  1. First difference: each change is the value minus the value before, for example 100.90 - 100.00 = 0.90.
  2. Level half means: 90.3 and 99.8, a gap of 9.5.
  3. Change half means: -0.11 and 0.03.

Use the idea

Before fitting an ARIMA model, plot the level and its first difference, and difference only when the level's mean is clearly not stable.

Where the conclusion applies

The series is the notebook's constructed integrated autoregression (seed 20260924), rebuilt from the same shocks for each phi. Differencing a series that is already stationary adds structure instead of removing it, so not every trend should be differenced.

Check your understanding: A series reads 100, 103, 101, 106. What are its first differences?
103 - 100 = 3, 101 - 103 = -2, 106 - 101 = 5, so the differences are 3, -2, 5.

Chapter 6 source: section "Part Three: What ARIMA Actually Is".

Demonstration 2 of 4

A signature in the lag structure

How do the ACF and PACF tell an autoregression apart from noise?

For an AR(1), the PACF has one spike at lag 1 and then nothing, while the ACF decays gradually because each change echoes through the next. With phi 0 the changes are white noise, yet lag 11 still crosses the band: a chance spike.

Equation: plus or minus 1.96 over the square root of n

Equation: the change at t equals phi times the previous change plus a shock

Scroll sideways for the whole equation

The ACF at lag k is the correlation of the changes with themselves k periods back; the PACF removes the effect of shorter lags first. n is the number of differences (189). Bars inside the band are not distinguishable from zero at the 5 percent level.

Predict first. At phi 0.65, which PACF lags stand far outside the band?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A signature in the lag structure. PACF of the differences for phi 0.65 at lags 1 to 12 with a band of plus or minus 0.143; outside: 1, 6, 10.
AR coefficient of the changes, phi: 0.65, Function: PACF (partial)
Constructed data: notebook 06 cells 1 and 3, the ACF and PACF of the training differences; phi 0 and 0.9 reuse the same shocks.

Calculated values

Differences used
189
Band
plus or minus 0.143
PACF at lag 1
0.626
PACF at lag 2
-0.060
Lags outside the band
1, 6, 10

The band is 1.96 / sqrt(189) = 1.96 / 13.75 = 0.143. The PACF at lag 1 is 0.626 and at lag 2 is -0.060; lags outside the band: 1, 6, 10. An AR(1) process should show a PACF that cuts off after lag 1 and an ACF that tails off. A lone spike at a high lag can be sampling noise: one in twenty lags crosses a 5 percent band by chance.

Worked steps

  1. Band: 1.96 / 13.75 = 0.143, since sqrt(189) is 13.75.
  2. PACF at lag 1: 0.626; at lag 2: -0.060.
  3. Lags beyond the band: 1, 6, 10.

Use the idea

Read the ACF and PACF of the stationary (differenced) series to propose candidate orders, then let the residual check and holdout errors decide.

Where the conclusion applies

Only training differences (189 values) are used. The band assumes white noise and is approximate; it screens candidates and does not select an order with certainty.

Common wrong turn: Every bar that crosses the band is real structure
The chapter warns that misreading the ACF and PACF is the most common source of wrong orders: part of the skill is telling meaningful deviations from sampling noise.
Check your understanding: With 100 observations, how wide is the 5 percent band on an ACF plot?
1.96 / sqrt(100) = 1.96 / 10 = 0.196, so plus or minus 0.196.

Chapter 6 source: section "Part Four: The Full Methodology".

Demonstration 3 of 4

Seasonal ARIMA against seasonal naive, origin by origin

Does the seasonal model beat the simplest seasonal benchmark when scored honestly over time?

At each origin all three models see only earlier data and forecast the next twelve points. Seasonal naive wins at origins 120 and 144, SARIMA at 168. The ranking depends on the origin, so only the average over origins counts.

Equation: x at t equals capital phi times x twelve periods earlier plus a shock

Scroll sideways for the whole equation

The series depends on its value twelve periods earlier, with seasonal AR coefficient capital phi 0.85 in the generator. SARIMA(0,0,0)(1,0,0) with period 12 models that; ARIMA(1,0,0) sees only lag 1; seasonal naive repeats the last twelve values.

Predict first. Averaged over three origins, will SARIMA beat seasonal naive by a wide margin?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Seasonal ARIMA against seasonal naive, origin by origin. Bars of twelve-step MAE for SARIMA, ARIMA and seasonal naive at origins 120, 144 and 168.
Forecast origin: Average of all three
Constructed data: notebook 06 cell 7, a series with lag-12 dependence (coefficient 0.85, seed 20260924) scored at origins 120, 144 and 168.

Calculated values

Mean MAE, SARIMA
1.85
Mean MAE, ARIMA
3.42
Mean MAE, Seasonal naive
1.86

Averaged over the three origins, SARIMA's MAE is (1.30 + 1.82 + 2.42) / 3 = 1.85, against 1.86 for seasonal naive and 3.42 for nonseasonal ARIMA. SARIMA is lowest, but SARIMA and seasonal naive are within 0.01 of each other; the nonseasonal model, which cannot see lag 12, trails both.

Worked steps

  1. SARIMA: (1.30 + 1.82 + 2.42) / 3 = 1.85.
  2. Seasonal naive mean: 1.86.
  3. Nonseasonal ARIMA mean: 3.42.

Use the idea

Compare any seasonal model with seasonal naive over several chronological origins before believing it adds anything.

Where the conclusion applies

Orders are fixed from the generator, not selected on the scored outcomes; real order selection must happen inside each training window. Three origins are few, and a margin of 0.01 is not a meaningful difference.

Check your understanding: A model's twelve-step MAEs at three origins are 1.30, 1.82 and 2.42. What is its mean?
(1.30 + 1.82 + 2.42) / 3 = 5.54 / 3 = 1.85.

Chapter 6 source: section "Part Four: The Full Methodology".

Demonstration 4 of 4

Listen to the residuals

How do you know an ARIMA model has said everything the data allows?

The random walk forecast is flat and its residuals are the raw changes, which are autocorrelated. One AR term absorbs that memory: residuals turn white and the forecast bends toward the drift. A second AR term adds nothing the check or the holdout rewards.

Equation: the change at t equals phi times the previous change plus a shock

Scroll sideways for the whole equation

ARIMA(p,d,q): p autoregressive lags, d differences, q moving-average lags. The residual check is a Ljung-Box test on 12 lags of the residuals: a small p-value means autocorrelation remains. MAE is mean absolute error on the held-out 30 points.

Predict first. ARIMA(0,1,0) is the random walk, with no AR term. Will its residuals pass the white noise check?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Listen to the residuals. ARIMA(1,1,0) forecast with a 95 percent band over 30 held-out steps against the actual path; model MAE 6.13, persistence 6.82, residual p 0.279.
Model order: ARIMA(1,1,0), Held-out steps scored: 30
Constructed data: notebook 06 cell 5, ARIMA(1,1,0) on the constructed series (MAE 6.13 against persistence 6.82); orders (0,1,0) and (2,1,0) are fitted the same way.

Calculated values

Model
ARIMA(1,1,0)
Residual check p-value (12 lags)
0.279
Residuals look like white noise
yes
Holdout MAE, first 30 steps
6.13
Persistence MAE, first 30 steps
6.82

ARIMA(1,1,0): the residual check gives p 0.279, so the residuals show no autocorrelation the test can detect. Over the first 30 held-out steps, persistence MAE minus model MAE is 6.82 - 6.13 = 0.69 in the model's favour.

Worked steps

  1. Residual check (Ljung-Box, 12 lags): p 0.279, pass.
  2. Model MAE over 30 steps: 6.13.
  3. Persistence MAE: 6.82.
  4. Difference: 6.82 - 6.13 = 0.69.

Use the idea

Fit, then test the residuals before reading the forecast; if the check fails, return to identification as Box and Jenkins prescribe.

Where the conclusion applies

The order (1,1,0) matches how the data were generated, which real data never tells you. White residuals do not prove all predictability (for example nonlinear) has been captured, and one 30-point path is not a calibration test of the band.

Check your understanding: Over 30 steps, a model's MAE is 6.13 and persistence's is 6.82. By how much does the model improve on persistence?
6.82 - 6.13 = 0.69 units of the series per step.

Chapter 6 source: section "Part Five: Listening to the Residuals".