The Art and Science of Forecasting

Chapter 5

The Filter

How much should new information change your mind? In proportion to how uncertain your prediction was and how noisy the measurement is.

Four demonstrations follow the chapter: one predict and update step of the Kalman filter and the gain it settles to, the difference between filtering in real time and smoothing in hindsight, the real Nile series with years removed, and what happens when the filter is told the wrong measurement noise.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Predict, then update: one step of the filter

How much should one new measurement move the estimate of something you cannot see?

Each step the filter first predicts (same position, more uncertainty), then weighs the measurement by the gain. Uncertainty grows by Q in the prediction and shrinks by the factor 1 - K in the update, so it settles where the two balance, and the gain settles with it.

Equation: the predicted variance equals the previous variance plus the process variance Q

Equation: the Kalman gain equals the predicted variance divided by the predicted variance plus the measurement variance R

Equation: the new estimate equals the previous estimate plus the gain times the measurement minus the previous estimate

Equation: the updated variance equals 1 minus the gain, times the predicted variance

Scroll sideways for the whole equation

The hidden position drifts with variance Q = 0.2 per step; each measurement adds noise with variance R = 2. The prediction carries the last estimate forward with its variance grown by Q; the gain K is predicted variance over predicted plus measurement variance; the innovation is the measurement minus the prediction.

Predict first. Start very unsure (variance 5) or fairly sure (variance 0.5). By step 30, will the two gains still differ?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Predict, then update: one step of the filter. Noisy measurements of a hidden random walk over 120 steps with the filtered estimate; step 1 marked, where the gain is 0.7222 and the estimate becomes -0.087.
Time step: 1, Starting variance: 5, very unsure (notebook)
Constructed data: the chapter notebook's simulated tracking example (hidden random walk, Q = 0.2, R = 2, seed 20260923), filtered as in its cell 3.

Calculated values

Predicted variance
5.2000
Kalman gain
0.7222
Measurement
-0.1206
Innovation (measurement minus prediction)
-0.1206
Updated estimate
-0.087
Updated variance
1.444

Step 1: predicted variance 5.0000 + 0.2 = 5.2000, so the gain is 5.2000 / (5.2000 + 2) = 0.7222. The estimate moves that share of the way to the measurement: 0.0000 + 0.72222 x (-0.1206) = -0.087. Early on the filter is unsure of itself, so it trusts the measurement heavily.

Worked steps

  1. Predict: 5.0000 + 0.2 = 5.2000.
  2. Gain: 5.2000 / (5.2000 + 2) = 0.7222.
  3. Innovation: -0.1206 - 0.0000 = -0.121.
  4. Update: 0.0000 + 0.72222 x (-0.1206) = -0.087.
  5. New variance: (1 - 0.72222) x 5.2000 = 1.444.

Use the idea

When a new report arrives about something you track, ask how uncertain your current estimate is and how noisy the report is; the share you move is the first over the sum of both.

Where the conclusion applies

Constructed data (seed 20260923) from exactly the random walk and noise the filter assumes, so Q and R are known. With fixed Q and R the gain ignores how large a surprise is; a large innovation alone does not make this filter adapt.

Check your understanding: The predicted variance is 3.0 and the measurement variance is 1.0. What is the gain, and where does an estimate of 10 go after a measurement of 14?
Gain = 3.0 / (3.0 + 1.0) = 0.75. New estimate = 10 + 0.75 x 4 = 13.

Chapter 5 source: section "The State-Space Framework".

Demonstration 2 of 4

Filtering and smoothing answer different questions

If you may use later measurements to estimate an earlier state, how much better can you do?

The filter can only react after a measurement surprises it. The smoother revisits each step knowing what came next, so a move the filter caught late is moved back to where it began.

Equation: the smoothed value equals the filtered value plus P over P plus Q, times the next smoothed value minus the filtered value

Scroll sideways for the whole equation

The filtered estimate uses measurements up to each step; the smoothed estimate starts from the last filtered value and runs backward, pulling each filtered value toward the next smoothed one by the factor P over P plus Q. RMSE is the root mean squared distance from the hidden state, known here only because the data are simulated.

Predict first. Over all 120 steps, will the smoother's error be less than 80 percent of the filter's?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Filtering and smoothing answer different questions. Measurements and the hidden state for steps 1 to 120, with the smoothed estimate.
Estimate shown: Smoother, looking back, Steps shown: All 120 steps
Constructed data: the chapter notebook's simulated tracking example (seed 20260923), filtered and smoothed as in its cells 3 and 4.

Calculated values

Filter RMSE, all steps
0.767
Smoother RMSE, all steps
0.587
Filter RMSE, steps 1 to 120
0.767
Smoother RMSE, steps 1 to 120
0.587

Over steps 1 to 120 the smoother's error against the hidden state is 0.587 against 0.767 for the filter, a ratio of 0.587 / 0.767 = 0.765. The smoother runs backward with later measurements, so it is the better record of the past and useless for navigating in real time.

Worked steps

  1. Filter RMSE over steps 1 to 120: 0.767.
  2. Smoother RMSE over the same steps: 0.587.
  3. Ratio: 0.587 / 0.767 = 0.765.

Use the idea

Use smoothed states to reconstruct history (a trend decomposition, a past demand level) and filtered states for anything decided in real time; never score a real-time method with smoothed values.

Where the conclusion applies

Constructed data (seed 20260923) from the model the filter assumes, with the hidden state available for scoring. On real data the hidden state is never seen, so these errors cannot be computed.

Check your understanding: A filtered value at step 50 is 2.0 with variance 0.6, Q is 0.2, and the smoothed value at step 51 is 3.0. What is the smoothed value at step 50?
2.0 + 0.6 / (0.6 + 0.2) x (3.0 - 2.0) = 2.0 + 0.75 x 1.0 = 2.75.

Chapter 5 source: section "The Methods".

Demonstration 3 of 4

The Nile with years missing

What does a state-space model do when measurements simply do not arrive?

In a missing year the filter skips the update and keeps only the prediction, so the level stays flat and its variance grows by the level variance each year. The smoother, looking from both ends, bridges the gap.

Equation: the observation equals the hidden level plus noise

Equation: the next hidden level equals the current level plus a random disturbance

Scroll sideways for the whole equation

The observed flow in year t is a hidden level plus irregular noise; the level drifts as a random walk. Both variances are estimated by maximum likelihood. The band is 1.2816 standard errors either side, an 80 percent interval.

Predict first. Remove 1900 to 1909. Will the filtered band for 1909 be more or less than 1.5 times as wide as for 1900 with nothing missing?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The Nile with years missing. Annual Nile flow 1871 to 1970 with the filtered local level and its 80 percent band; 1900 to 1909 missing; 1909 marked at 1050.3 plus or minus 151.8.
Years removed: 1900 to 1909, Estimate shown: Filtered, real time
Real data: annual flow of the Nile at Aswan, 1871 to 1970, in hundreds of millions of cubic metres, a public-domain series first analysed by Cobb (1978) and bundled with the companion. Fitted with the companion's chapter 5 local-level tool; the missing years are removed for this demonstration.

Calculated values

Year inspected
1909
Filtered level
1050.3
80 percent half-width
151.8
Level variance
1049
Irregular variance
15485
Signal-to-noise ratio
0.068

1900 to 1909 missing. The filtered level for 1909 is 1050.3, with an 80 percent band 1050.3 - 151.8 = 898.5 up to 1202.1. With no measurement in the gap the filter carries the last level forward and its band widens every missing year. Estimated signal-to-noise ratio: 1049 / 15485 = 0.068.

Worked steps

  1. Estimate for 1909: 1050.3; 80 percent half-width 151.8.
  2. Band: 1050.3 - 151.8 = 898.5 to 1050.3 + 151.8 = 1202.1.
  3. Signal-to-noise: 1049 / 15485 = 0.068.

Use the idea

When reports arrive late or not at all, keep the estimate and widen its uncertainty for each missing period instead of inventing values; the next real measurement updates it normally.

Where the conclusion applies

A local level with Gaussian noise; the missing years are removed here on purpose and are missing at random. The model does not know about the sharp drop around 1899, so the level adapts to it gradually; that drop is a known feature of this series, not a causal finding.

What this does not settle

The chapter's filter is optimal only under one assumption: that you knew the right model. A local level on the Nile is a choice, and a wrong choice gives confident wrong bands.

Chapter 5 source: "Kalman's filter was optimal under one assumption".

Check your understanding: A filtered level has variance 4,000 and the level variance is 1,500 per year. With three years missing, what is the predicted variance at the end of the gap?
4,000 + 3 x 1,500 = 8,500.

Chapter 5 source: section "The Methods".

Demonstration 4 of 4

Telling the filter the wrong noise

What happens to the estimate and its stated uncertainty when the filter's measurement noise is wrong?

The gain comes from the assumed variances. Too small an R makes the gain large, so the estimate copies noisy measurements while claiming to be precise. Too large an R makes the gain small, so the estimate lags and the band balloons.

Equation: the Kalman gain equals the predicted variance divided by the predicted variance plus the measurement variance R

Equation: root mean squared error equals the square root of the average squared gap between the hidden state and the estimate

Scroll sideways for the whole equation

R is the measurement variance the filter is told; the data were generated with 2. Coverage is the share of steps where the hidden state falls inside the filter's 95 percent band; RMSE is the root mean squared distance from it.

Predict first. Tell the filter R = 0.2, ten times too small. Will its 95 percent band still hold the hidden state at least 80 percent of the time?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Telling the filter the wrong noise. Hidden state and the filtered estimate with its 95 percent band under assumed measurement variance 0.2; coverage 45.8 percent.
Measurement variance told to the filter: 0.2, ten times too small
Constructed data: the chapter notebook's simulated tracking example refiltered under three assumed measurement variances (cell 7).

Calculated values

Settled gain
0.6180
State RMSE
1.040
Nominal 95 percent coverage
45.8 percent
Mean band width
1.382

With assumed R 0.2 the gain settles at 0.3236 / (0.3236 + 0.2) = 0.6180. The state error is 1.040 and the 95 percent band holds the hidden state in 55 of 120 steps, 45.8 percent. The filter believes the measurements too much: it chases noise and its band is far too narrow.

Worked steps

  1. Settled gain: 0.3236 / (0.3236 + 0.2) = 0.6180.
  2. Coverage: 55 / 120 = 0.458.
  3. Mean band width 1.382; state RMSE 1.040.

Use the idea

Check a filter's innovations and interval coverage before trusting its stated uncertainty: honest weighting depends on both the prediction's uncertainty and the measurement's, neither one alone.

Where the conclusion applies

Constructed data (seed 20260923); the true R is known only because this is a controlled experiment. Coverage is measured on one dependent path of 120 steps, not an operational calibration guarantee.

Check your understanding: A settled predicted variance is 0.74 and the filter is told R = 0.74. What is its gain?
0.74 / (0.74 + 0.74) = 0.5.

Chapter 5 source: section "What You Cannot See But Must Estimate".