The Art and Science of Forecasting

Chapter 26

How to Be Ready Without Being Certain

A forecast earns its keep through the decision it informs and the score it keeps.

Four demonstrations follow the chapter: what a stated probability promises about many cases, how costs turn a probability into an action, why a kept score settles which forecaster to trust, and why the decision needs a score of its own beside the forecast's.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Seventy percent means thirty in a hundred do not happen

If you say 70 percent and the event does not happen, were you wrong?

Each dot groups the journal entries whose stated probability fell in one bin and plots the share that happened. A forecaster who says what they mean sits near the dashed line, and every bin below 1 contains events that did not happen.

Equation: the observed share y bar equals k events that happened out of n

Scroll sideways for the whole equation

n is the number of events given about the same stated probability, k the number that happened, and y-bar their observed share.

Predict first. In the first 200 journal entries, events were stated at about 0.7. Will every one of them happen?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Seventy percent means thirty in a hundred do not happen. Observed share of events that happened against stated probability for the first 200 journal entries; the bin near 0.7 is circled at 0.71, 17 of 24.
Stated probability: About 0.7, Journal entries scored: 200
Constructed data: the chapter notebook's 3,000 seeded synthetic events (cell 1), propensity drawn between 0.05 and 0.95, outcomes drawn from it.

Calculated values

Events in the journal
200
Events stated at 0.65 to 0.75
24
Events that happened
17
Observed share
0.71
Did not happen
7

Among 24 events stated at about 0.7, 17 happened: 17 / 24 = 0.71, close to the stated level. The other 7 did not happen, and that is what the stated probability promised: a single miss is the expected minority, not proof the forecast was wrong. Only a run of cases can show whether the stated level was miscalibrated.

Worked steps

  1. Events with stated probability from 0.65 to 0.75: 24.
  2. Of those, 17 happened.
  3. Observed share: 17 / 24 = 0.71, against about 0.7 stated.

Use the idea

Keep a journal of stated probabilities and outcomes, and judge a 70 percent call by how often your 70 percent calls come true, never by a single outcome.

Where the conclusion applies

The stated probabilities are the true propensities of seeded events (seed 20260944), so they are calibrated by construction; a real journal has to earn that. With 50 entries a bin holds a handful of events and its share is noisy.

Check your understanding: Of 40 events you called at 70 percent, 26 happened. What share is that, and is it far from 70 percent?
26 / 40 = 0.65, five points below 0.70: with 40 events that gap is within ordinary chance.

Chapter 26 source: section "Section Three: The Practice".

Demonstration 2 of 4

A probability becomes a decision through costs

Should you act on a 30 percent chance?

The falling line is the expected cost of acting, the rising line the expected cost of waiting. Where they cross is the threshold: above it acting is cheaper. The threshold comes from the costs, not from how confident anyone feels.

Equation: the expected cost of acting equals the false alarm cost times 1 minus p

Equation: the expected cost of waiting equals the miss cost times p

Equation: the threshold p star equals the false alarm cost over the false alarm cost plus the miss cost

Scroll sideways for the whole equation

p is the probability of the event, C FA the cost of acting when nothing happens (a false alarm) and C M the cost of not acting when it does (a miss). p star is the probability at which the two expected costs are equal.

Predict first. A false alarm costs 2 and a miss costs 8. Should you act at probability 0.3?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A probability becomes a decision through costs. Expected cost of acting falls and of waiting rises with probability; they cross at the threshold 0.20. At probability 0.3 acting costs 1.40 and waiting 2.40.
Costs (false alarm, miss): False alarm 2, miss 8, Probability of the event: 0.3
Constructed data: the chapter notebook's 3,000 seeded synthetic events (cells 1 and 2), false alarm cost 2 and miss cost 8, plus the equal and reversed costs the notebook's exercise asks for.

Calculated values

Threshold
0.20
Expected cost if you act
1.40
Expected cost if you wait
2.40
Decision at this probability
act
Mean realized cost at the threshold, 3,000 events
0.885
Cheapest threshold in hindsight
0.18

With false alarm cost 2 and miss cost 8, the threshold is 2 / (2 + 8) = 0.20. At probability 0.3, acting costs 2 x (1 - 0.3) = 1.40 and waiting costs 8 x 0.3 = 2.40, so acting is cheaper. On the notebook's events the cheapest threshold in hindsight is 0.18, beside 0.20 by sampling chance; picking that hindsight value and reporting its cost as future performance would overfit.

Worked steps

  1. Threshold: 2 / (2 + 8) = 0.20.
  2. Act: 2 x (1 - 0.3) = 1.40.
  3. Wait: 8 x 0.3 = 2.40.
  4. Decision at 0.3: act.

Use the idea

Before the forecast arrives, write down what a false alarm and a miss cost and the threshold they imply; then the probability decides the action, not the mood in the room.

Where the conclusion applies

The threshold is the cheapest rule only when the probabilities are calibrated and the costs are fixed and known per event. The events are the notebook's seeded propensities (seed 20260944).

Check your understanding: A false alarm costs 9 and a miss costs 1. Above what probability should you act?
9 / (9 + 1) = 0.9: act only when the event is at least 90 percent likely.

Chapter 26 source: section "Section Four: The Full Methodology".

Demonstration 3 of 4

Keeping score makes the comparison repeatable

How do you know which of two forecasters, or their average, to trust?

Each line is one source's Brier score on successive batches of 200 events. Batches move around; the running record is what shows a steady ordering.

Equation: the ensemble probability for event i is the average of model A and model B

Equation: the Brier score is the average over n events of the squared difference between probability p i and outcome y i

Scroll sideways for the whole equation

a and b are the two models' probabilities for event i, p-bar their equal average, y the outcome (1 if it happened, 0 if not). The Brier score BS averages the squared gaps over n events.

Predict first. Over all 3,000 events, will the equal ensemble score better (lower) than both models?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Keeping score makes the comparison repeatable. Brier score per 200-event batch for model A, model B and their equal average over 5 batches, with equal ensemble highlighted; overall scores 0.2087, 0.2041 and 0.1972.
Events in the journal: 1000, Source highlighted: Equal ensemble
Constructed data: the chapter notebook's synthetic journal (cell 3), two models built from the seeded propensities with independent noise, scored against the generated outcomes with the companion's Brier score.

Calculated values

Events scored
1000
Brier, model A
0.2087
Brier, model B
0.2041
Brier, equal ensemble
0.1972
Batches where the ensemble beats both
4 of 5

Over 1000 journal entries the equal ensemble scores 0.1972; the best of the others scores 0.2041, a difference of 0.1972 - 0.2041 = -0.0069. Lowest overall: equal ensemble. The ensemble beats both models in 4 of 5 batches: a record kept batch after batch, not one memorable call, is what settles which source to trust.

Worked steps

  1. Brier, model A: 0.2087; model B: 0.2041; ensemble: 0.1972.
  2. Equal ensemble minus best other: 0.1972 - 0.2041 = -0.0069.
  3. Ensemble best in 4 of 5 batches.

Use the idea

Write each probability down with its time stamp before the outcome, score every source the same way, and let the accumulated score, not the last surprise, decide which source gets weight.

Where the conclusion applies

Both models add independent noise to the same seeded propensities, so averaging helps here by construction. This is a synthetic journal with fixed equal weights, not a historical time stamped record.

Check your understanding: Two forecasters score Brier 0.21 and 0.19 over 500 events. Which is better, and by how much?
The second, lower is better: 0.21 - 0.19 = 0.02.

Chapter 26 source: section "Section Four: The Full Methodology".

Demonstration 4 of 4

Score the decision as well as the forecast

Can the same forecast lead to good and bad outcomes?

Each line is the running total cost of one decision rule applied to the same forecasts and outcomes. The forecast's accuracy does not change between the lines; the decisions do.

Equation: the threshold p star equals the false alarm cost over the false alarm cost plus the miss cost

Scroll sideways for the whole equation

With false alarm cost C FA = 2 and miss cost C M = 8, p star is 0.2. Each line adds 2 for every false alarm and 8 for every miss as the events resolve.

Predict first. With the ensemble, will threshold 0.8 cost more or less than threshold 0.2 over 3,000 events?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Score the decision as well as the forecast. Cumulative cost over 3,000 events using the equal ensemble at thresholds 0.2, 0.5 and 0.8, with 0.2 highlighted ending at 2786.
Threshold highlighted: 0.2, Forecast used: Equal ensemble
Constructed data: the chapter notebook's cell 4, the synthetic equal ensemble and outcomes under thresholds 0.2, 0.5 and 0.8; the true propensities added for comparison.

Calculated values

Forecast used
the equal ensemble
Brier score of the forecast
0.1934
Threshold
0.2
Total cost, 3,000 events
2786
Mean cost per event
0.929
Cheapest of the three thresholds
0.2

Using the equal ensemble with threshold 0.2, the total cost is 2786, or 2786 / 3000 = 0.929 per event. The Brier score is 0.1934 at every threshold: the forecast is the same, only the decision changed. Among the three thresholds 0.2 is cheapest, the one closest to the cost-derived 0.2, so a decision rule needs its own score next to the forecast's.

Worked steps

  1. Brier score of the forecast: 0.1934, unchanged by the threshold.
  2. Total cost at threshold 0.2: 2786.
  3. Per event: 2786 / 3000 = 0.929.

Use the idea

Report two scores for a forecasting process: the forecast's own score and the cost of the decisions taken with it, because either one can fail while the other looks fine.

Where the conclusion applies

Costs are fixed at 2 and 8 per event and every event is resolved. The events and forecasts are the notebook's seeded synthetic journal (seed 20260944), not decisions anyone actually took.

Check your understanding: Over 100 events a rule raises 6 false alarms (cost 2 each) and 3 misses (cost 8 each). What is its total cost?
6 x 2 + 3 x 8 = 12 + 24 = 36.

Chapter 26 source: section "Section Five: The Return".