Demonstration 1 of 4
Seventy percent means thirty in a hundred do not happen
If you say 70 percent and the event does not happen, were you wrong?
Each dot groups the journal entries whose stated probability fell in one bin and plots the share that happened. A forecaster who says what they mean sits near the dashed line, and every bin below 1 contains events that did not happen.
Scroll sideways for the whole equation
n is the number of events given about the same stated probability, k the number that happened, and y-bar their observed share.
Predict first. In the first 200 journal entries, events were stated at about 0.7. Will every one of them happen?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's 3,000 seeded synthetic events (cell 1), propensity drawn between 0.05 and 0.95, outcomes drawn from it.
Calculated values
- Events in the journal
- 200
- Events stated at 0.65 to 0.75
- 24
- Events that happened
- 17
- Observed share
- 0.71
- Did not happen
- 7
Among 24 events stated at about 0.7, 17 happened: 17 / 24 = 0.71, close to the stated level. The other 7 did not happen, and that is what the stated probability promised: a single miss is the expected minority, not proof the forecast was wrong. Only a run of cases can show whether the stated level was miscalibrated.
Worked steps
- Events with stated probability from 0.65 to 0.75: 24.
- Of those, 17 happened.
- Observed share: 17 / 24 = 0.71, against about 0.7 stated.
Use the idea
Keep a journal of stated probabilities and outcomes, and judge a 70 percent call by how often your 70 percent calls come true, never by a single outcome.
Where the conclusion applies
The stated probabilities are the true propensities of seeded events (seed 20260944), so they are calibrated by construction; a real journal has to earn that. With 50 entries a bin holds a handful of events and its share is noisy.
Check your understanding: Of 40 events you called at 70 percent, 26 happened. What share is that, and is it far from 70 percent?
Chapter 26 source: section "Section Three: The Practice".
Demonstration 2 of 4
A probability becomes a decision through costs
Should you act on a 30 percent chance?
The falling line is the expected cost of acting, the rising line the expected cost of waiting. Where they cross is the threshold: above it acting is cheaper. The threshold comes from the costs, not from how confident anyone feels.
Scroll sideways for the whole equation
p is the probability of the event, C FA the cost of acting when nothing happens (a false alarm) and C M the cost of not acting when it does (a miss). p star is the probability at which the two expected costs are equal.
Predict first. A false alarm costs 2 and a miss costs 8. Should you act at probability 0.3?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's 3,000 seeded synthetic events (cells 1 and 2), false alarm cost 2 and miss cost 8, plus the equal and reversed costs the notebook's exercise asks for.
Calculated values
- Threshold
- 0.20
- Expected cost if you act
- 1.40
- Expected cost if you wait
- 2.40
- Decision at this probability
- act
- Mean realized cost at the threshold, 3,000 events
- 0.885
- Cheapest threshold in hindsight
- 0.18
With false alarm cost 2 and miss cost 8, the threshold is 2 / (2 + 8) = 0.20. At probability 0.3, acting costs 2 x (1 - 0.3) = 1.40 and waiting costs 8 x 0.3 = 2.40, so acting is cheaper. On the notebook's events the cheapest threshold in hindsight is 0.18, beside 0.20 by sampling chance; picking that hindsight value and reporting its cost as future performance would overfit.
Worked steps
- Threshold: 2 / (2 + 8) = 0.20.
- Act: 2 x (1 - 0.3) = 1.40.
- Wait: 8 x 0.3 = 2.40.
- Decision at 0.3: act.
Use the idea
Before the forecast arrives, write down what a false alarm and a miss cost and the threshold they imply; then the probability decides the action, not the mood in the room.
Where the conclusion applies
The threshold is the cheapest rule only when the probabilities are calibrated and the costs are fixed and known per event. The events are the notebook's seeded propensities (seed 20260944).
Check your understanding: A false alarm costs 9 and a miss costs 1. Above what probability should you act?
Chapter 26 source: section "Section Four: The Full Methodology".
Demonstration 3 of 4
Keeping score makes the comparison repeatable
How do you know which of two forecasters, or their average, to trust?
Each line is one source's Brier score on successive batches of 200 events. Batches move around; the running record is what shows a steady ordering.
Scroll sideways for the whole equation
a and b are the two models' probabilities for event i, p-bar their equal average, y the outcome (1 if it happened, 0 if not). The Brier score BS averages the squared gaps over n events.
Predict first. Over all 3,000 events, will the equal ensemble score better (lower) than both models?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's synthetic journal (cell 3), two models built from the seeded propensities with independent noise, scored against the generated outcomes with the companion's Brier score.
Calculated values
- Events scored
- 1000
- Brier, model A
- 0.2087
- Brier, model B
- 0.2041
- Brier, equal ensemble
- 0.1972
- Batches where the ensemble beats both
- 4 of 5
Over 1000 journal entries the equal ensemble scores 0.1972; the best of the others scores 0.2041, a difference of 0.1972 - 0.2041 = -0.0069. Lowest overall: equal ensemble. The ensemble beats both models in 4 of 5 batches: a record kept batch after batch, not one memorable call, is what settles which source to trust.
Worked steps
- Brier, model A: 0.2087; model B: 0.2041; ensemble: 0.1972.
- Equal ensemble minus best other: 0.1972 - 0.2041 = -0.0069.
- Ensemble best in 4 of 5 batches.
Use the idea
Write each probability down with its time stamp before the outcome, score every source the same way, and let the accumulated score, not the last surprise, decide which source gets weight.
Where the conclusion applies
Both models add independent noise to the same seeded propensities, so averaging helps here by construction. This is a synthetic journal with fixed equal weights, not a historical time stamped record.
Check your understanding: Two forecasters score Brier 0.21 and 0.19 over 500 events. Which is better, and by how much?
Chapter 26 source: section "Section Four: The Full Methodology".
Demonstration 4 of 4
Score the decision as well as the forecast
Can the same forecast lead to good and bad outcomes?
Each line is the running total cost of one decision rule applied to the same forecasts and outcomes. The forecast's accuracy does not change between the lines; the decisions do.
Scroll sideways for the whole equation
With false alarm cost C FA = 2 and miss cost C M = 8, p star is 0.2. Each line adds 2 for every false alarm and 8 for every miss as the events resolve.
Predict first. With the ensemble, will threshold 0.8 cost more or less than threshold 0.2 over 3,000 events?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's cell 4, the synthetic equal ensemble and outcomes under thresholds 0.2, 0.5 and 0.8; the true propensities added for comparison.
Calculated values
- Forecast used
- the equal ensemble
- Brier score of the forecast
- 0.1934
- Threshold
- 0.2
- Total cost, 3,000 events
- 2786
- Mean cost per event
- 0.929
- Cheapest of the three thresholds
- 0.2
Using the equal ensemble with threshold 0.2, the total cost is 2786, or 2786 / 3000 = 0.929 per event. The Brier score is 0.1934 at every threshold: the forecast is the same, only the decision changed. Among the three thresholds 0.2 is cheapest, the one closest to the cost-derived 0.2, so a decision rule needs its own score next to the forecast's.
Worked steps
- Brier score of the forecast: 0.1934, unchanged by the threshold.
- Total cost at threshold 0.2: 2786.
- Per event: 2786 / 3000 = 0.929.
Use the idea
Report two scores for a forecasting process: the forecast's own score and the cost of the decisions taken with it, because either one can fail while the other looks fine.
Where the conclusion applies
Costs are fixed at 2 and 8 per event and every event is resolved. The events and forecasts are the notebook's seeded synthetic journal (seed 20260944), not decisions anyone actually took.
Check your understanding: Over 100 events a rule raises 6 false alarms (cost 2 each) and 3 misses (cost 8 each). What is its total cost?
Chapter 26 source: section "Section Five: The Return".