Demonstration 1 of 4
A proper score makes honesty the best report
If you believe an event has a 70 percent chance, can reporting 80 percent improve your expected Brier score?
The curve is the expected loss for every possible report under one fixed belief. Its lowest point sits exactly at the belief, so moving the report toward 0.5 or toward the extremes can only raise it. That is what the chapter means by a strictly proper rule.
Scroll sideways for the whole equation
p is your real belief that the event happens, r the probability you report. The Brier loss is (1 - r) squared if the event happens and r squared if it does not; the expected loss weights the two by p and 1 - p.
Predict first. With a belief of 0.7, does reporting 0.8 raise or lower the expected Brier loss?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter's own example of a 70 percent belief, with other reports and a 30 percent belief added; confirmed with the companion's Brier function on ten events in proportion to the belief.
Calculated values
- Belief p
- 0.7
- Report r
- 0.80
- Expected loss of the report
- 0.2200
- Expected loss of the honest report
- 0.2100
- Extra expected loss
- 0.0100
With belief 0.7, reporting 0.80 expects 0.7 x 0.0400 + 0.3 x 0.6400 = 0.2200. The honest report 0.7 expects 0.2100, so this distortion costs 0.2200 - 0.2100 = 0.0100.
Worked steps
- If the event happens: (1 - 0.80)^2 = 0.0400.
- If it does not: 0.80^2 = 0.6400.
- Weight by the belief: 0.7 x 0.0400 + 0.3 x 0.6400 = 0.2200.
Use the idea
When forecasts are scored with a proper rule, report the number you actually believe; the rule removes the incentive to inflate or hedge, though calibration still has to be checked against outcomes.
Where the conclusion applies
The expectation is taken under the forecaster's own belief. On a handful of real outcomes, luck can make a distorted report score better, and an honest but overconfident belief is still miscalibrated.
Common wrong turn: Overstating confidence looks prescient
Check your understanding: With a belief of 0.7, what is the expected Brier loss of reporting 0.6?
Chapter 17 source: section "Section Three: What Probabilistic Forecasting Actually Is".
Demonstration 2 of 4
An interval must earn its stated coverage
A forecaster calls an interval 80 percent. How do you find out whether it is?
Each point checks one nominal level against the outcomes. A narrow forecast sits below the dashed line at every level, a broad one above it. A prediction interval is a frequency statement about many cases like this one, which is why it can be checked only over many outcomes.
Scroll sideways for the whole equation
The outcomes y are standard normal. The forecast believes they have spread s, so its central interval at nominal coverage c runs from minus s times z to plus s times z, z being the standard normal quantile at (1 + c) / 2. Empirical coverage c hat is the share of the n = 20000 outcomes inside.
Predict first. A forecast whose spread is only 0.65 of the true one states 80 percent intervals. How often do they cover?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's 20,000 synthetic standard-normal outcomes (cell 3) and its three forecast scales.
Calculated values
- Forecast scale s
- 0.65
- Nominal coverage
- 0.80
- Outcomes inside the interval
- 11870 of 20000
- Empirical coverage
- 0.5935
- Empirical minus nominal
- -0.2065
At scale 0.65 the nominal 0.80 interval holds 11870 of 20000 outcomes: 11870 / 20000 = 0.5935, and 0.5935 - 0.80 = -0.2065. This forecast under-covers: the intervals are too narrow. Only the scale 1 forecast matches the outcomes it is scored on.
Worked steps
- Half-width: 0.65 times the standard normal quantile at 0.900.
- Count inside: 11870 of 20000, so 11870 / 20000 = 0.5935.
- Gap: 0.5935 - 0.80 = -0.2065.
Use the idea
Before trusting a stated 80 or 95 percent band, count how often past bands of the same kind held the outcome, level by level.
Where the conclusion applies
20,000 constructed independent standard-normal outcomes (seed 20260935) and intervals centred correctly; only the spread is wrong. Real errors are dependent and can be biased, and a few dozen outcomes cannot pin coverage down.
Check your understanding: Of 50 past 90 percent intervals, 38 held the outcome. What is the empirical coverage, and is the band too narrow or too wide?
Chapter 17 source: section "Section Four: The Full Methodology".
Demonstration 3 of 4
Calibrate the interval on an earlier block
How wide should a band be so that it misses about a share alpha of later outcomes?
A coefficient is fitted on the first block and frozen; its absolute one-step errors on a later calibration block set the radius; the band is then judged on a still later test block it never saw. The correction uses n + 1 and rounds up, which ordinary interpolated quantiles do not do.
Scroll sideways for the whole equation
n is the number of calibration residuals (80 here), alpha the target miss rate, and k the rank of the residual used as the band's radius after sorting them from smallest to largest; the brackets mean round up.
Predict first. With 80 calibration residuals and alpha 0.1, which residual sets the radius?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded AR(1) series (cell 7), its training, calibration and test blocks, with the miss rate varied.
Calculated values
- Miss rate alpha
- 0.10
- Calibration residuals
- 80
- Rank used
- 73
- Radius
- 1.8984
- Test outcomes covered
- 77 of 81
- Test coverage
- 0.9506
- Target 1 minus alpha
- 0.90
The corrected rank uses (80 + 1) x 0.90 = 72.90, rounded up to 73, so the radius is the 73th smallest of the 80 calibration residuals, 1.8984. On the later test block it covers 77 / 81 = 0.9506 against a target of 0.90. The residuals are serially dependent, so this is an empirical check, not a guarantee.
Worked steps
- Fit the coefficient on steps 0 to 79 and freeze it: 0.6212.
- Absolute one-step residuals on steps 80 to 159: 80 of them.
- Rank: (80 + 1) x 0.90 = 72.90, rounded up to 73; radius 1.8984.
- Test coverage on steps 160 to 240: 77 / 81 = 0.9506.
Use the idea
Keep three blocks in time order: fit, calibrate, test. Report the test coverage and the band width together, and never tune and claim coverage on the same outcomes.
Where the conclusion applies
A constructed AR(1) series with coefficient 0.7 and 240 shocks (seed 20260935). The finite-sample guarantee needs exchangeable scores; these residuals are serially dependent, so the chapter and the notebook assert no guarantee.
What this does not settle
The chapter gives split conformal its coverage guarantee only for exchangeable calibration and test scores, and notes that ordinary time-series dependence does not automatically satisfy those conditions. The coverage counted here is a result on one constructed series, not a guarantee.
Chapter 17 source: "Ordinary time-series dependence does not automatically satisfy those conditions".
Check your understanding: With 40 calibration residuals and alpha 0.2, which ranked residual does the corrected rule use?
Chapter 17 source: section "Section Four: The Full Methodology".
Demonstration 4 of 4
Width alone is not enough
Can a forecaster win an interval score by making the interval narrow, or by making it very wide?
The width term grows steadily with the interval, while the miss penalty shrinks as fewer outcomes fall outside. The sum is lowest near the interval that truly covers 80 percent, so neither gaming direction pays.
Scroll sideways for the whole equation
l and u are the interval's lower and upper ends, y the outcome, alpha the nominal miss rate (0.2 for an 80 percent interval). A plus subscript means keep the value if positive, else 0, so a miss costs 2 / alpha = 10 times its distance.
Predict first. Which of the offered widths gives the lowest mean score?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's 20,000 synthetic standard-normal outcomes (cells 3 and 5), scored at nominal 80 percent.
Calculated values
- Interval width
- 1.8000
- Coverage
- 0.6291
- Mean width term
- 1.8000
- Mean miss penalty
- 2.0649
- Mean interval score
- 3.8649
The interval from -0.9000 to 0.9000 covers 0.6291 of the outcomes. Its mean score is width plus miss penalty: 1.8000 + 2.0649 = 3.8649, higher than the 3.5532 scored at the true 80 percent width.
Worked steps
- Width term: 2 x 0.9000 = 1.8000.
- Each miss costs 2 / 0.2 = 10 times its distance outside; averaged over 20000 outcomes: 2.0649.
- Mean score: 1.8000 + 2.0649 = 3.8649.
Use the idea
Score interval forecasts with width and misses together, and report coverage beside mean width; a band that always covers by being huge is not a good forecast.
Where the conclusion applies
20,000 constructed standard-normal outcomes (seed 20260935) and intervals centred at 0. The minimum sits at the true quantiles only when the outcomes follow the distribution assumed; with skewed outcomes it moves.
Check your understanding: For an 80 percent interval from 8 to 12 and an outcome of 14, what is the interval score?
Chapter 17 source: section "Section Four: The Full Methodology".