The Art and Science of Forecasting

Chapter 4

The Smoother

One number, alpha, decides how fast a forecast forgets, and the family built on it survives because it is simple and hard to beat.

Four demonstrations follow the chapter: the geometric weights behind Brown's formula, the trade between reacting to a shift and chasing noise, how damping bends a trend at long horizons, and a holdout test in which seasonality has to earn its place.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

How the past fades: geometric weights

How much does a month from a year ago count in an exponentially smoothed forecast?

Unrolling the recursion gives the newest month weight alpha, the one before alpha x (1 - alpha), and so on, a geometric decline. The moving average instead weights twelve months equally and then forgets month thirteen entirely.

Equation: the new forecast equals alpha times the most recent observation plus 1 minus alpha times the previous forecast

Equation: the weight of an observation k periods old is alpha times 1 minus alpha to the power k

Scroll sideways for the whole equation

alpha is the smoothing parameter between 0 and 1; the weight of an observation k months old is alpha times (1 - alpha) to the power k. The bars are those weights; the red step is a twelve-month moving average.

Predict first. Brown's example keeps 0.8 of the weight at each step back. Does the last year then carry more or less than 90 percent of the forecast?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How the past fades: geometric weights. Bars of exponential weights falling from 0.200 for the newest month, against a flat moving-average weight of 1/12 that stops after twelve months.
Smoothing parameter alpha: 0.2
Constructed data: the chapter's own weighting example (each step back keeps 0.8 of the weight, alpha 0.2), with alpha varied.

Calculated values

Weight on the newest month
0.200
Weight one month back
0.160
Weight two months back
0.128
Share of weight in the last 12 months
0.9313

With alpha 0.2 each step back keeps 1 - 0.2 = 0.8 of the weight before it: 0.2 x 0.8^2 = 0.128 for the month two back. The last twelve months carry 1 - 0.8^12 = 0.9313 of the total; the moving average gives each of them 1/12 = 0.083 and then drops month thirteen to nothing.

Worked steps

  1. Newest month: alpha = 0.2.
  2. One back: 0.2 x 0.8 = 0.160.
  3. Two back: 0.2 x 0.8^2 = 0.128.
  4. Last twelve months together: 1 - 0.8^12 = 0.9313.

Use the idea

When someone asks how far back a smoothed forecast looks, answer with the weights: at alpha 0.5 the last four months already carry 1 - 0.5^4 = 0.9375 of the forecast.

Where the conclusion applies

Weights shown for an unbounded history; a finite history also gives the initial level the leftover weight. The weights describe the method, not whether old data are still relevant to your series.

Check your understanding: With alpha 0.3, what weight does the observation three months back receive?
0.3 x 0.7^3 = 0.3 x 0.343 = 0.1029.

Chapter 4 source: section "Chapter 4: The Smoother".

Demonstration 2 of 4

How fast should a forecast forget?

When demand jumps to a new level, how many months does exponential smoothing need to follow it?

Each month the level moves alpha of the way toward the new observation. A large alpha closes the jump in a month or two but copies each month's noise; a small alpha ignores noise and also ignores the jump for a long time.

Equation: the new forecast equals alpha times the most recent observation plus 1 minus alpha times the previous forecast

Scroll sideways for the whole equation

The updated level is alpha times this month's observation plus (1 - alpha) times the previous level. The constructed series sits near 20 for 30 months and near 35 after, with noise of standard deviation 2.

Predict first. With alpha 0.1, will the level reach within 2 units of 35 before month 40?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How fast should a forecast forget?. Noisy constructed series jumping from about 20 to about 35 at month 31, with the smoothed level for alpha 0.5; month 31 highlighted at 28.30.
Smoothing parameter alpha: 0.5, Month inspected: 31
Constructed data: the chapter notebook's synthetic level shift (cell 3) and its three alpha values.

Calculated values

Observation this month
35.205
Previous level
21.397
Updated level
28.30
First month within 2 of 35
33
Level wobble, months 41 to 60 (sd)
0.75

Month 31: level = 0.5 x 35.205 + 0.5 x 21.397 = 28.30. With alpha 0.5 the level first comes within 2 units of the new 35 in month 33, and once settled it wobbles with standard deviation 0.75. A middle setting reacts within a few months without following every observation.

Worked steps

  1. Observation: 35.205; previous level: 21.397.
  2. New level: 0.5 x 35.205 + 0.5 x 21.397 = 28.30.

Use the idea

When you choose a smoothing parameter, ask how fast the real level changes compared with the month-to-month noise, as the chapter's new employee must, and check both the lag after a shift and the wobble after it.

Where the conclusion applies

Constructed data with seed 20260922; the updated levels are computed after seeing each month, as in the notebook, so they are filters, not forecasts made beforehand. One shift and one noise level; other series reward other settings.

Check your understanding: The level is 30.00, the new observation is 36.00 and alpha is 0.3. What is the updated level?
0.3 x 36.00 + 0.7 x 30.00 = 10.80 + 21.00 = 31.80.

Chapter 4 source: section "The Methods".

Demonstration 3 of 4

A trend need not continue forever

How much does damping change a trend forecast, and at which horizons?

Both models are fitted to the same 50 months. The undamped forecast adds one full trend step per month, a straight line. The damped forecast multiplies each further step by phi, so the line bends; with phi estimated near 1 the bend is gentle and shows mostly at long horizons.

Equation: the forecast h periods ahead is the current level plus h times the current trend

Equation: the damped forecast is the level plus phi plus phi squared up to phi to the h, times the trend

Scroll sideways for the whole equation

The forecast h months ahead starts from the final level and adds the final trend per month times a multiplier: h for Holt, phi + phi^2 + ... + phi^h for the damped trend, with damping coefficient phi between 0 and 1.

Predict first. At 30 months ahead, will the damped forecast sit more or less than 2 units below the undamped one?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A trend need not continue forever. Constructed trending history for 50 months with Holt and damped Holt forecasts for 30 months; Damped Holt at horizon 30 marked at 56.12.
Model highlighted: Damped Holt, Months ahead: 30
Constructed data: the chapter notebook's synthetic trend (cell 5) with its Holt and damped Holt fits.

Calculated values

Final level
43.9765
Final trend per month
0.4370
Trend multiplier
27.784
Damped Holt forecast at h = 30
56.12
Other model at the same horizon
59.28

Damped Holt at 30 months ahead: 43.9765 + 27.784 x 0.4370 = 56.12, where the multiplier adds phi + phi^2 + ... + phi^30 with phi 0.995 instead of 30. The undamped model gives 59.28 at the same horizon, a gap of 3.16.

Worked steps

  1. Final level 43.9765, final trend 0.4370 per month.
  2. Trend multiplier at h = 30: 27.784 (damped).
  3. Forecast: 43.9765 + 27.784 x 0.4370 = 56.12.

Use the idea

For long-horizon plans built on a trend, compare the undamped and damped forecasts and ask whether the extra growth in the straight line is something you would bet on.

Where the conclusion applies

Parameters are estimated by the fitting routine on constructed data (seed 20260922); here phi reaches 0.995, the top of its allowed range, so damping is mild. The chapter's point is skepticism about long trends, not that damping always wins.

Check your understanding: A damped model has final level 100, trend 2 per month and phi 0.9. What is its forecast two months ahead?
100 + (0.9 + 0.81) x 2 = 100 + 1.71 x 2 = 103.42.

Chapter 4 source: section "The Methods".

Demonstration 4 of 4

Seasonality earns its place on the holdout

Which member of the smoothing family forecasts two unseen years of seasonal demand best?

Each model is fitted on the first 120 months and forecasts the last 24 in one go. The figure shows the chosen model's forecast against the held-out months; the MAE scores the gap.

Equation: mean absolute error equals the average of the absolute differences between the observations and the forecasts

Scroll sideways for the whole equation

MAE is the mean absolute error over the 24 held-out months, in units. Seasonal naive repeats the last observed year; SES tracks a level only; Holt-Winters adds trend and additive seasonality; Theta is the competition-winning benchmark.

Predict first. Will simple exponential smoothing beat the seasonal naive reference on this seasonal series?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Seasonality earns its place on the holdout. Constructed seasonal series of 144 months with the Holt-Winters forecast over the last 24 held-out months, MAE 1.64.
Model shown: Holt-Winters additive
Constructed data: the companion's seeded seasonal series as fitted in the chapter notebook (cell 7), last 24 months held out.

Calculated values

Seasonal naive MAE
2.88
SES MAE
7.16
Holt-Winters MAE
1.64
Theta MAE
1.82
Ratio to seasonal naive
0.57

Holt-Winters holdout MAE 1.64 against seasonal naive 2.88: 1.64 / 2.88 = 0.57, so it beats the seasonal naive reference. Seasonality earns its place only because it wins on months the models never saw.

Worked steps

  1. Holdout: the last 24 months, never used for fitting or choosing.
  2. Holt-Winters MAE over those months: 1.64.
  3. Ratio to seasonal naive: 1.64 / 2.88 = 0.57.

Use the idea

Score any candidate against a seasonal naive reference on months it never saw, and report the result even when the more complex method fails to win.

Where the conclusion applies

One constructed series (seed 20260922) with additive seasonality and a gentle trend, which suits Holt-Winters. A single holdout is one draw; the chapter's competitions scored thousands of series.

Check your understanding: A model's holdout MAE is 2.10 and the seasonal naive MAE is 2.80. What is the ratio, and does the model beat the reference?
2.10 / 2.80 = 0.75, below 1, so it beats the reference by a quarter of its error.

Chapter 4 source: section "The Methods".