The Art and Science of Forecasting

Chapter 2

The Mathematics of Belief

Belief is a quantity that evidence should move, by the right amount, and a forecaster should be scored on it.

Four demonstrations follow the chapter: Bayes's update of a belief by seven successes in ten, Galton's regression toward the mean, Wald's missing aircraft, and Brier's score that rewards honest probabilities.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Seven heads in ten: updating a belief

After seeing seven successes in ten trials, what should you believe about the underlying rate?

The dashed curve is what you believed before the trials, the solid curve after. Multiplying the prior by the likelihood of the observed successes moves the curve toward the observed rate and narrows it; more trials move it further and narrow it more.

Equation: the posterior is proportional to the likelihood times the prior

Equation: theta given the data follows a Beta distribution with parameters a plus s and b plus f

Scroll sideways for the whole equation

theta is the unknown success probability. The prior is a Beta distribution with parameters a and b (here 2 times the prior strength each, mean 0.5). After s successes and f failures the posterior is Beta with parameters a plus s and b plus f.

Predict first. With seven successes in ten and the Beta(2, 2) prior, will the posterior mean be above, at, or below 0.7?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Seven heads in ten: updating a belief. Prior Beta(2.0, 2.0) as a dashed curve and posterior Beta(9.0, 5.0) peaking near 0.64, with the region from 0.6 to 0.8 shaded.
Evidence observed: 7 successes in 10 trials, Prior strength (Beta(2, 2) scaled): 1: Beta(2, 2)
Constructed data: the chapter's seven successes in ten trials, and the companion's three-batch example (21 in 32); the posterior comes from the companion's Beta-binomial tool.

Calculated values

Successes, failures
7, 3
Prior
Beta(2.0, 2.0)
Posterior mean
0.643
95 percent interval
0.386 to 0.861
Probability between 0.6 and 0.8
0.548

Posterior mean = (2.0 + 7) / (4.0 + 10) = 9.0 / 14.0 = 0.643, between the prior mean 0.500 and the observed rate 0.700. The posterior gives probability 0.548 that the true rate lies between 0.6 and 0.8. A stronger prior holds the estimate nearer 0.5; more trials pull it toward the data.

Worked steps

  1. Prior: Beta(2.0, 2.0), mean 0.500.
  2. Add 7 successes and 3 failures: Beta(2.0 + 7, 2.0 + 3) = Beta(9.0, 5.0).
  3. Posterior mean: 9.0 / 14.0 = 0.643.
  4. Probability the rate is between 0.6 and 0.8: 0.548.

Use the idea

When a new product's first results arrive, write down the prior you held, update it by the counts, and report the posterior interval rather than the raw percentage.

Where the conclusion applies

Trials are exchangeable with one common success probability, and every trial is recorded. The prior's strength is a choice: if the three strengths give materially different answers, the data are not yet decisive.

Check your understanding: With a Beta(1, 1) prior and 7 successes in 10 trials, what is the posterior mean?
Beta(1 + 7, 1 + 3) = Beta(8, 4), mean 8 / 12 = 0.667.

Chapter 2 source: section "Part One: The Essay in the Drawer".

Demonstration 2 of 4

Taller parents, less tall children: regression toward the mean

If you pick the people who scored highest once, how high will they score next time?

Every point is one person measured twice. The fitted line is flatter than the equality line because an extreme first score is partly luck, and luck does not persist. The diamond marks the selected group's two averages.

Equation: the predicted y equals a plus b times x

Scroll sideways for the whole equation

x is the first measurement and y-hat the predicted second measurement; b is the slope and a the intercept. Each measurement is a lasting ability plus fresh noise, both standardised; the noise standard deviation is 0.5 or 1.

Predict first. With noise as large as ability (standard deviation 1), will the top 10 percent keep more or less than two thirds of their lead?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Taller parents, less tall children: regression toward the mean. Scatter of 600 first against second measurements with noise 1.0; the fitted line has slope 0.53, flatter than the equality line, and the top 10 percent fall from 2.48 to 1.15.
Group selected on the first measurement: Top 10 percent, Noise standard deviation: 1
Constructed data: the chapter notebook's 600 seeded ability and measurement draws, with the noise rescaled for the second setting.

Calculated values

Fitted slope
0.53
Expected slope
0.50
Top 10 percent, first mean
2.48
Top 10 percent, second mean
1.15
Share of the extreme kept
0.46

The top 10 percent on the first measurement averaged 2.48; the same people averaged 1.15 the second time, 1.15 / 2.48 = 0.46 of their first extreme. With ability and noise variances 1 and 1.00, the expected slope is 1 / (1 + 1.00) = 0.50; the fitted slope is 0.53. Nobody changed: the luck in the first score did not repeat.

Worked steps

  1. Select the top 10 percent by first measurement: mean 2.48.
  2. Their second measurement mean: 1.15.
  3. Kept share: 1.15 / 2.48 = 0.46.
  4. Expected slope: 1 / (1 + 1.00) = 0.50; fitted 0.53.

Use the idea

When a product has its best month ever, forecast next month from a line that slopes less than one through the history, not from the record month itself.

Where the conclusion applies

Ability and noise are independent and normal, and the second measurement has the same noise as the first (seed 20260920, the notebook's draws). The spread of the population does not shrink: regression is about selected extremes, not a force toward mediocrity.

Common wrong turn: Regression means the group is getting more ordinary
The chapter calls Galton's phrase 'regression to mediocrity' an error: the distribution of heights stayed stable. Only the selected extremes move back, because part of their extremity was chance.
Check your understanding: If ability and noise each have variance 1, and someone scores 2.0 on the first test, what second score do you expect?
Slope 1 / (1 + 1) = 0.5, so 0.5 x 2.0 = 1.0.

Chapter 2 source: section "Part Three: Taller Parents, Less Tall Children".

Demonstration 3 of 4

The planes that did not come back

If you study only the aircraft that returned, where will the damage seem to be?

Grey bars count every hit; teal bars count hits on returning aircraft only. Where survival is low, the teal bar shrinks: missing aircraft take their damage with them.

Equation: the probability a hit was in the engine given the aircraft returned is proportional to the chance of returning after an engine hit times the chance of an engine hit

Scroll sideways for the whole equation

Each of 10,000 aircraft is hit once, equally often in the wing, fuselage or engine. The chance of returning depends on where it was hit: 0.9 for the wing, and the chosen values for the fuselage and engine.

Predict first. If engine-hit aircraft return only 2 times in 10 and fuselage-hit aircraft 8 times in 10, what share of hits on returning aircraft will be in the engine?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The planes that did not come back. Paired bars: engine hits are 0.331 of all hits but 0.108 among returning aircraft when engine survival is 0.2 and fuselage survival 0.8.
Chance of returning after an engine hit: 0.2, Chance of returning after a fuselage hit: 0.8
Constructed data: the chapter notebook's simulated damage and survival process, not wartime records.

Calculated values

Aircraft hit
10000
Returned
6373
Engine hits, all aircraft
0.331
Engine hits, returned only
0.108
Engine hits that returned
689 of 3307

Engine hits are 0.331 of all hits but 689 / 6373 = 0.108 of the hits on returning aircraft, because only 0.2 of engine-hit aircraft return. Counting only survivors makes the engine look rarely hit, when it is the place hits are most deadly.

Worked steps

  1. Of 10000 aircraft hit, 3307 were hit in the engine.
  2. 689 of those returned; 6373 aircraft returned in all.
  3. Engine share among survivors: 689 / 6373 = 0.108.
  4. Expected: 0.2 / (0.9 + 0.8 + 0.2) = 0.105 with equal exposure.

Use the idea

Before reading a pattern in any dataset, ask which cases could not enter it: closed stores, churned customers, failed products, funds that shut down.

Where the conclusion applies

Hit locations are equally likely and survival depends only on location (seed 20260920, the notebook's draws). Blank areas alone do not prove vulnerability; the inference needs comparable exposure, as the chapter's notebook warns.

Check your understanding: If wing, fuselage and engine hits are equally common and return with probabilities 0.9, 0.8 and 0.2, what share of returning aircraft were hit in the engine?
0.2 / (0.9 + 0.8 + 0.2) = 0.2 / 1.9 = 0.105, about one in ten.

Chapter 2 source: section "The Planes That Didn't Come Back".

Demonstration 4 of 4

The Brier score rewards saying what you believe

Can a forecaster improve a Brier score by hedging toward 50 percent or by sounding bolder?

The curve shows the average Brier score as the reports are pulled toward 0.5 (k below 1) or pushed away from it (k above 1). In expectation its lowest point is exactly k = 1, honesty, which is what makes the score proper; with 1000 drawn outcomes the curve's minimum falls close to 1, and among the four choices k = 1 scores best.

Equation: the Brier score is the average of the squared differences between the stated probability f i and the outcome o i

Scroll sideways for the whole equation

f is the stated probability for event i, o is 1 if it happened and 0 if not, n the number of forecasts. Each of 1000 events has a true probability between 0.1 and 0.9, the forecaster's belief; the report shades it by k.

Predict first. Which shading factor gives the lowest Brier score: hedging (k = 0.5), honesty (k = 1) or boldness (k = 1.25)?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The Brier score rewards saying what you believe. Brier score against the shading factor k; the curve is lowest at k = 1. The marked point at k = 0.50 scores 0.2059.
Shading factor k: 0.5: hedge
Constructed data: the chapter notebook's 1000 seeded events and outcomes; scores from the companion's Brier function.

Calculated values

k
0.50
Brier score
0.2059
Honest forecast (k = 1)
0.1899
Report when belief is 0.7
0.60
Expected score at belief 0.7
0.220

A forecaster who hedges toward 0.5 turns a belief of 0.7 into a report of 0.60. If the event happens 7 times in 10, the expected score is 0.7 x (1 - 0.60)^2 + 0.3 x 0.60^2 = 0.220, against 0.210 for reporting 0.7. Over the 1000 constructed events the Brier score is 0.2059, and honest reports score 0.1899.

Worked steps

  1. One forecast: say 0.8 and it rains, (0.8 - 1)^2 = 0.04; it stays dry, (0.8 - 0)^2 = 0.64.
  2. Belief 0.7 reported as 0.60: 0.7 x (1 - 0.60)^2 + 0.3 x 0.60^2 = 0.220.
  3. Average over 1000 events at k = 0.50: 0.2059.
  4. Honest reporting averages 0.1899; always saying 0.5 averages 0.2500.

Use the idea

Score forecasters with a proper rule such as the Brier score, so that their best strategy is to report what they believe rather than what sounds safe.

Where the conclusion applies

The forecaster's beliefs equal the true probabilities (seed 20260920, the notebook's draws), so any shading can only hurt. A real forecaster's beliefs are imperfect; the score still rewards reporting them honestly.

Check your understanding: You believe rain is 70 percent likely and report 60 percent. What is your expected Brier score, and how does it compare with reporting 70 percent?
0.7 x (1 - 0.6)^2 + 0.3 x 0.6^2 = 0.112 + 0.108 = 0.220, worse than 0.7 x 0.3^2 + 0.3 x 0.7^2 = 0.063 + 0.147 = 0.210.

Chapter 2 source: section "Part Four: The Methodology of Measured Uncertainty".