The Art and Science of Forecasting

Chapter 9

The Oracle Problem

Expertise and calibration are different properties, and only scoring forecasts against outcomes can tell them apart.

Four demonstrations follow the chapter's methodology: how to check calibration against outcomes, why the Brier score rewards stating your real belief, why the base rate is the benchmark any expert must beat, and how overconfidence shows up in a record grouped by stated confidence.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Do the numbers mean what they say?

When a forecaster says 70 percent, how would you check that it means 70 percent?

Each point is one bin: how often the events in it happened against what was forecast for them. Calibration is a property of a collection of forecasts, so a bin needs many events before its point means anything.

Equation: the observed frequency in bin b equals the average outcome of the events in that bin

Scroll sideways for the whole equation

Forecasts are grouped into bins one tenth wide. For each bin b holding n b events, the observed frequency is the share of those events that happened (outcome 1). A calibrated forecaster's points sit on the diagonal.

Predict first. A forecaster always says 50 percent, and about one third of the events happen. Is that forecaster calibrated?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Do the numbers mean what they say?. Reliability curve for the calibrated forecaster over 6000 constructed events: 9 bins, the fullest at forecast 0.25 against frequency 0.26.
Forecaster: Calibrated, Events scored: 6000
Constructed data: the chapter notebook's 6,000 synthetic events and its calibrated and overconfident forecasters, plus two constant forecasts.

Calculated values

Events
6000
Bins with more than 20 events
9
Fullest bin, mean forecast
0.25
Fullest bin, observed frequency
0.26
Forecast minus frequency
-0.01
Overall event rate
0.335

For the calibrated forecaster over 6000 events, the fullest bin holds 1289 events with mean forecast 0.25 and observed frequency 0.26: 0.25 - 0.26 = -0.01. Read the whole curve: points below the diagonal are forecasts that said more than happened.

Worked steps

  1. Fullest bin: 1289 events, mean forecast 0.25.
  2. Share that happened: 0.26.
  3. Gap: 0.25 - 0.26 = -0.01.

Use the idea

Keep a log of your stated probabilities and outcomes; once you have dozens per bin, plot this curve and see whether your 70 percent behaves like 70 percent.

Where the conclusion applies

6,000 constructed events whose true probabilities follow a beta(2, 4) distribution (seed 20260927); bins with 20 or fewer events are not drawn, as in the notebook. Real questions are not independent draws.

Check your understanding: Of 40 events you forecast at 80 percent, 26 happened. What is the observed frequency, and how far is it from your forecast?
26 / 40 = 0.65; 0.80 - 0.65 = 0.15, so those forecasts were too confident by 15 points.

Chapter 9 source: section "Calibration, Accuracy, and Resolution".

Demonstration 2 of 4

The Brier score rewards your real belief

If you believe 70 percent but say 60 percent to sound humble, what does the score do?

The curve is the average score for every stated probability when 7 of 10 (or 3 of 10) events happen. Its lowest point is exactly the real belief, so any distortion, toward 50 percent or away from it, raises the score.

Equation: the Brier score of one forecast equals f minus o, squared

Equation: the average Brier score equals one over N times the sum of the squared differences between each forecast and its outcome

Scroll sideways for the whole equation

f is the stated probability, o the outcome (1 if the event happened, 0 if not), N the number of forecasts. Lower is better: 0 is perfect, 0.25 is what a constant 50 percent always scores.

Predict first. With a real belief of 0.7, will stating 0.6 score better or worse than stating 0.7?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The Brier score rewards your real belief. U-shaped curve of average Brier score against stated probability, lowest at the real belief 0.7; stating 0.6 scores 0.22.
Stated probability: 0.6, Real belief: 0.7
Constructed data: ten events per state, happening in proportion to the real belief, scored with the companion's Brier function.

Calculated values

Stated probability
0.6
Events that happen (of 10)
7
Average Brier score
0.22
Score when stating the belief
0.21

Ten events at a real chance of 0.7: 7 happen. Stating 0.6 scores 0.7 x 0.16 + 0.3 x 0.36 = 0.22. That is 0.01 worse than stating the real belief, which scores 0.21.

Worked steps

  1. When the event happens: (1 - 0.6)^2 = 0.16.
  2. When it does not: (0.6 - 0)^2 = 0.36.
  3. Average: 0.7 x 0.16 + 0.3 x 0.36 = 0.22.

Use the idea

Score your own probability forecasts with the Brier score and state the number you actually believe; the score is built so that honesty is the winning strategy.

Where the conclusion applies

The events happen exactly in proportion to the belief (7 or 3 of 10), so the average equals the expected score. With few events, luck can make a distorted forecast score better on a particular day.

Common wrong turn: A hedged probability is the cautious choice
The chapter is explicit: stating 60 percent when you believe 70 percent makes your expected Brier score worse. The score penalizes distortion in either direction.
Check your understanding: You state 0.8 for an event that does not happen. What is your Brier score for it?
(0.8 - 0)^2 = 0.64.

Chapter 9 source: section "The Brier Score".

Demonstration 3 of 4

The base rate is the score to beat

How do you know whether a forecaster knows anything beyond what usually happens?

Every bar is one way of forecasting the same events. The base rate bar is the outside view; the calibrated forecaster has case knowledge and uses it honestly, the overconfident one has the same knowledge and overstates it.

Equation: the average Brier score equals one over N times the sum of the squared differences between each forecast and its outcome

Scroll sideways for the whole equation

Each forecast is scored on the same events with the average Brier score. The base rate forecast says one third every time, the share of events the generator produces on average: the outside view with no case knowledge.

Predict first. Over all 6000 events, does the overconfident forecaster still beat the base rate?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The base rate is the score to beat. Brier scores over 6000 constructed events: calibrated 0.1898, overconfident 0.2018, always 50 percent 0.2500, base rate 0.2229.
Forecast compared: Calibrated, Events scored: 6000
Constructed data: the chapter notebook's 6,000 synthetic events and its three forecasts, plus a constant 50 percent, scored with the companion's Brier function.

Calculated values

Events
6000
Event rate
0.335
Chosen forecast
0.1898
Base rate one third
0.2229
Base rate minus chosen
0.0331

Over 6000 events, the calibrated forecaster scores 0.1898 and the base rate scores 0.2229: 0.2229 - 0.1898 = 0.0331, so it beats the outside view. Any expert worth hearing should beat the number you get by asking only what usually happens.

Worked steps

  1. Chosen forecast: Brier 0.1898.
  2. Base rate one third: Brier 0.2229.
  3. Margin: 0.2229 - 0.1898 = 0.0331.

Use the idea

When you evaluate an expert, score the base rate on the same questions first; skill is the margin over it, not the score alone.

Where the conclusion applies

All forecasts are scored on identical events (seed 20260927). One third is the generator's known average rate; in practice the base rate must come from a reference class, and with 100 events the ranking can shift by chance.

Check your understanding: An analyst's Brier score is 0.230 and the base rate's is 0.215 on the same 200 questions. Does the analyst add skill?
No. 0.215 - 0.230 = -0.015: the analyst is worse than simply forecasting the base rate.

Chapter 9 source: section "Outside View vs. Inside View".

Demonstration 4 of 4

Confidence is a claim you can test

If an expert's probabilities are pushed toward the extremes, what happens to how often their confident calls come true?

Each point groups calls by how confident they were. On the dashed line confidence equals accuracy. Stretching the forecasts moves points right without moving them up, which is what overconfidence looks like in a record.

Equation: the forecast equals one half plus k times the true probability minus one half

Scroll sideways for the whole equation

p is the event's true probability, f the forecast, and k how far the forecaster stretches away from 0.5: k = 1 is honest, k = 1.6 is the notebook's overconfident forecaster (forecasts are kept between 0.01 and 0.99). Stated confidence is the probability given to whichever outcome the forecast favours.

Predict first. At the notebook's stretch of 1.6, how often are the forecaster's calls in the 90 percent and above bin right?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Confidence is a claim you can test. Fraction correct against stated confidence for stretch 1.6 over 6000 events; the top bin states 0.97 and is right 0.84 of the time.
Stretch away from 0.5, k: 1.6, Events scored: 6000
Constructed data: the chapter notebook's 6,000 synthetic events and its overconfident forecaster (stretch 1.6), with the stretch varied.

Calculated values

Stretch k
1.6
Events
6000
Top bin, events
2285
Top bin, stated confidence
0.97
Top bin, fraction correct
0.84
Confidence minus accuracy
0.13

In the most confident bin (2285 events) the forecaster claims 0.97 and is right 0.84 of the time: 0.97 - 0.84 = 0.13. The forecaster is overconfident: stretching by 1.6 pushes forecasts away from 0.5 without adding knowledge.

Worked steps

  1. Top bin: 2285 events, mean stated confidence 0.97.
  2. Fraction correct: 0.84.
  3. Overconfidence: 0.97 - 0.84 = 0.13.

Use the idea

Ask an expert for a track record grouped by confidence; a run of confident calls is only evidence when you can see how many there were and how often they came true.

Where the conclusion applies

Constructed events (seed 20260927) with the forecaster's knowledge fixed and only the stretch varied. A single confident success proves little: the bins need many events.

Check your understanding: An expert made 50 calls at 90 percent confidence and 38 came true. What is the gap between confidence and accuracy?
38 / 50 = 0.76; 0.90 - 0.76 = 0.14, so the expert was overconfident by 14 points.

Chapter 9 source: section "Overconfidence Revisited".