The Art and Science of Forecasting

Chapter 10

The Superforecaster

Good forecasting is a set of habits that can be written down, checked and scored.

Four demonstrations follow the chapter's toolkit: update the odds rather than the probability when evidence arrives, find which assumption drives a Fermi estimate, extremize an averaged forecast only after testing the multiplier on new events, and score a forecast journal one predeclared forecast per event against a baseline.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Multiply the odds, not the probability

Evidence twice as likely if the event will happen arrives. Does a 20 percent forecast become 40 percent?

The solid line is the forecast after each piece of evidence. On the odds scale, evidence for and against the event cancel exactly (ratios 2 then 0.5 return to 0.2); multiplying probabilities directly drifts and can pass 1.

Equation: the odds O equal p divided by one minus p

Equation: the posterior odds equal the prior odds times the likelihood ratio L R

Equation: the probability p equals the odds O divided by one plus O

Scroll sideways for the whole equation

p is the probability, O the odds p / (1 - p), and LR the likelihood ratio: how many times more likely the new evidence is if the event will happen than if it will not. The prior is 0.2; the four pieces of evidence have ratios 2, 0.5, 3 and 1.2.

Predict first. Starting at 0.2, evidence with a likelihood ratio of 2 arrives. What is the new probability?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Multiply the odds, not the probability. Probability path from 0.2 updated through 1 likelihood ratios on the odds scale, reaching 0.333.
Updates applied: 1, Rule: Multiply the odds
Constructed data: the chapter notebook's illustrative prior of 0.2 and likelihood ratios 2, 0.5, 3 and 1.2.

Calculated values

Updates applied
1
Likelihood ratio
2.0
Odds before
0.25
Odds after
0.50
Probability after
0.333
Probability if multiplied directly
0.40

Update 1 multiplies the odds: 0.25 x 2.0 = 0.50, then converts back: 0.50 / (1 + 0.50) = 0.333. The probability stays between 0 and 1 however strong the evidence.

Worked steps

  1. Prior odds: 0.2 / (1 - 0.2) = 0.25.
  2. Odds before update 1: 0.25.
  3. Multiply: 0.25 x 2.0 = 0.50.
  4. Back to probability: 0.50 / (1 + 0.50) = 0.333.

Use the idea

When a piece of news arrives, ask how much more likely it would be if the event will happen than if it will not, and update the odds; small ratios should make small moves.

Where the conclusion applies

The likelihood ratios are illustrative and the pieces of evidence are treated as independent given the outcome. Evidence that shares a source must not be multiplied twice, and an invented ratio for an ambiguous news report only dresses up a judgment.

Common wrong turn: Multiply the probability by the likelihood ratio
The rule multiplies the odds. Doubling 0.2 gives 0.40, but the correct posterior is 0.333, and a large enough ratio applied to a probability would pass 1.
Check your understanding: Your prior is 0.4 and evidence with a likelihood ratio of 3 arrives. What is the posterior probability?
Prior odds 0.4 / 0.6 = 0.667; 0.667 x 3 = 2.00; 2.00 / 3.00 = 0.667.

Chapter 10 source: section "The Methodology".

Demonstration 2 of 4

Which Fermi assumption matters most?

In the piano-tuner estimate, which guess should you check first?

Each bar is the range of the estimate when one input moves across its range. The highlighted bar is the input you chose and the dot its value at the chosen end. Long bars show where the uncertainty lives.

Equation: the number of tuners N equals households H times piano share f times tunings t times hours per tuning h, divided by working hours per tuner W

Scroll sideways for the whole equation

N is the number of tuners, H households (1.1 million), f the share with a piano (0.10), t tunings per piano per year (1), h hours per tuning (1.5) and W working hours per tuner per year (1,500). Each row varies one input across an illustrative range while the others stay at the base.

Predict first. Which input moves the estimate further: households between 1.0 and 1.2 million, or piano ownership between 5 and 15 percent?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Which Fermi assumption matters most?. Five horizontal ranges of the tuner estimate, one input varied at a time around the base of 110; piano ownership at its high value gives 165.0.
Input varied: Piano ownership, End of its range: High
Constructed data: the chapter's piano-tuner teaching inputs and the notebook's illustrative ranges for each input.

Calculated values

Input varied
piano ownership
Value used
0.15
Estimated tuners
165.0
Change from 110
55.0
Span of this row (low to high)
110.0
Widest span
110.0

With piano ownership at its high value: 1,100,000 x 0.15 x 1.0 x 1.5 / 1,500 = 165.0. That moves the estimate 165.0 - 110 = 55.0 tuners; its span of 110.0 tuners ties for the widest, so a wide row is the assumption worth checking first.

Worked steps

  1. Numerator: 1,100,000 x 0.15 x 1.0 x 1.5 = 247500 hours a year.
  2. Divide by 1,500 hours per tuner: 247500 / 1,500 = 165.0.
  3. Against the base: 165.0 - 110 = 55.0.

Use the idea

Write any estimate as a product of parts, give each part a plausible range, and spend your research time on the parts with the widest bars.

Where the conclusion applies

The ranges are illustrative sensitivity assumptions from the notebook, not measured uncertainty or confidence intervals, and the inputs are varied one at a time. The base inputs are the chapter's teaching numbers, not a count of real tuners.

Check your understanding: With 1.1 million households, 10 percent ownership, one tuning a year, 1.5 hours a tuning and 1,800 hours per tuner, how many tuners?
1,100,000 x 0.10 x 1 x 1.5 / 1,800 = 91.7 tuners.

Chapter 10 source: section "The Methodology".

Demonstration 3 of 4

Extremize, then test on new events

Averaged forecasts sit too close to 50 percent. How far should you push them, and how do you know?

Each curve is the Brier score for every multiplier: the bold one on the events you chose, the faint one on the other half. The dashed line is the multiplier picked using training events only.

Equation: the logit of p equals the log of p over one minus p

Equation: the logit of the extremized forecast q equals a times the logit of the average forecast p bar

Equation: the average Brier score equals one over N times the sum of the squared differences between each forecast and its outcome

Scroll sideways for the whole equation

p bar is the average of eight forecasters' probabilities, a the multiplier on the log-odds scale (a = 1 leaves the average alone, a above 1 pushes it toward 0 or 1), and q the extremized forecast. f and o in the Brier score are forecast and outcome (1 or 0) over N events.

Predict first. The multiplier 1.29 was chosen as best on the first 2,000 events. On the other 2,000, does it still beat the plain average?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Extremize, then test on new events. Brier score against log-odds multiplier on the held-out events (last 2,000); multiplier 1.29 scores 0.2152, the plain average 0.2172.
Log-odds multiplier a: 1.29, Events scored: Held out (last 2,000)
Constructed data: the chapter notebook's seeded 4,000 synthetic events and eight-expert averages, scored with the companion's Brier function.

Calculated values

Multiplier a
1.29
Events scored
held-out events (last 2,000)
Brier, extremized
0.2152
Brier, plain average
0.2172
Plain minus extremized
0.0020
An average of 0.70 becomes
0.749

On the held-out events (last 2,000), multiplier 1.29 scores 0.2152 against 0.2172 for the plain average: 0.2172 - 0.2152 = 0.0020, so this multiplier improves on the plain average by 0.0020. The multiplier 1.29 was chosen on the training events only; the held-out curve happens to be lowest near 1.42, and choosing again after looking would no longer be a fair test.

Worked steps

  1. logit of 0.70: ln(0.70 / 0.30) = 0.8473.
  2. Stretch: 1.29 x 0.8473 = 1.0930.
  3. Back to a probability: 0.749, further from 0.5 when a is above 1.
  4. Score gain on these events: 0.2172 - 0.2152 = 0.0020.

Use the idea

If you average several forecasters, choose the extremizing multiplier on past forecasts, then check it on outcomes you did not use to choose it before you trust it.

Where the conclusion applies

4,000 constructed events (seed 20260928), eight synthetic experts who each see a shrunk, noisy version of the true log-odds. Shared information between forecasters can leave little extra signal to recover; these events are a teaching case, not a measurement of real forecasters.

Check your understanding: An average forecast of 0.80 is extremized with a = 2 on the log-odds scale. What is the new probability?
Odds 0.80 / 0.20 = 4; squaring the odds gives 4^2 = 16; 16 / 17 = 0.941.

Chapter 10 source: section "The Methodology".

Demonstration 4 of 4

Write it down first, then count

How do you score a forecast journal fairly once the outcomes are in?

Each dot is a bin of recorded probabilities: how often those events happened against what was written down, with the number of events beside it. Small counts mean the dot can sit far from the line by chance.

Equation: the average Brier score equals one over N times the sum of the squared differences between each forecast and its outcome

Scroll sideways for the whole equation

f is the forecast recorded for an event, o its outcome (1 if it happened, 0 if not), and N the number of resolved events. Each event counts once: its latest forecast before a cutoff seven days ahead of the deadline.

Predict first. Scoring all 100 resolved events at their pre-cutoff forecasts, does the journal beat a constant 0.5?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Write it down first, then count. Calibration points with event counts for pre-cutoff forecasts over 100 resolved synthetic events; mean Brier 0.2042 against 0.2500 for the baseline.
Record scored: Pre-cutoff forecasts, Resolved events: 100
Constructed data: the chapter notebook's seeded synthetic event journal, scored with the companion's journal scoring function.

Calculated values

Resolved events scored
100
Record scored
pre-cutoff forecasts
Mean Brier score
0.2042
Baseline Brier (0.5)
0.2500
Baseline minus record
0.0458
Bins with events
5

Over the first 100 resolved events, pre-cutoff forecasts score 0.2042 against 0.2500 for the baseline: 0.2500 - 0.2042 = 0.0458, better than the baseline by 0.0458. Each event counts once, at its last forecast before the cutoff; 20 unresolved events and one late entry are left out.

Worked steps

  1. Score each event once: (f - o)^2, for example (0.7 - 1)^2 = 0.09.
  2. Average over 100 events: 0.2042.
  3. Baseline 0.5 always scores (0.5 - o)^2 = 0.25 per event: 0.2500.
  4. Margin: 0.2500 - 0.2042 = 0.0458.

Use the idea

Write each forecast with its date before the outcome is known, add revisions instead of overwriting, and score one predeclared forecast per event against a baseline you also wrote down in advance.

Where the conclusion applies

120 synthetic events (seed 20260928): 100 resolved, 20 unresolved and excluded, plus one deliberately late entry. The editable record is a teaching example, not tamper-proof storage, and 100 events cannot certify calibration.

Common wrong turn: A late update is just a better forecast
An entry made after the outcome is known scores perfectly and proves nothing. Only forecasts written before the cutoff count; the hindsight setting shows the free improvement it would give.
Check your understanding: An event has a forecast of 0.3 before the cutoff and 0.99 entered after it resolved yes. Which counts, and what is its Brier score?
Only 0.3 is eligible: (0.3 - 1)^2 = 0.49. The 0.99 entry is kept in the record but not scored.

Chapter 10 source: section "Practicing the Skill".