The Art and Science of Forecasting

Chapter 17

Living with Probability

A probabilistic forecast is judged on honesty and usefulness: calibrated coverage first, then as much sharpness as the evidence allows.

Four demonstrations follow the chapter's methodology: why a proper score makes the honest probability the best report, how to check whether an interval earns its stated coverage, how a band is calibrated on an earlier block with the corrected rank, and why an interval score punishes both narrow and needlessly wide intervals.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

A proper score makes honesty the best report

If you believe an event has a 70 percent chance, can reporting 80 percent improve your expected Brier score?

The curve is the expected loss for every possible report under one fixed belief. Its lowest point sits exactly at the belief, so moving the report toward 0.5 or toward the extremes can only raise it. That is what the chapter means by a strictly proper rule.

Equation: the expected Brier score equals p times one minus r, squared, plus one minus p times r squared

Scroll sideways for the whole equation

p is your real belief that the event happens, r the probability you report. The Brier loss is (1 - r) squared if the event happens and r squared if it does not; the expected loss weights the two by p and 1 - p.

Predict first. With a belief of 0.7, does reporting 0.8 raise or lower the expected Brier loss?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A proper score makes honesty the best report. Expected Brier loss against the reported probability for belief 0.7: a U-shaped curve lowest at 0.7; the report 0.80 sits at 0.2200.
Reported probability r: 0.8, Real belief p: 0.7
Constructed data: the chapter's own example of a 70 percent belief, with other reports and a 30 percent belief added; confirmed with the companion's Brier function on ten events in proportion to the belief.

Calculated values

Belief p
0.7
Report r
0.80
Expected loss of the report
0.2200
Expected loss of the honest report
0.2100
Extra expected loss
0.0100

With belief 0.7, reporting 0.80 expects 0.7 x 0.0400 + 0.3 x 0.6400 = 0.2200. The honest report 0.7 expects 0.2100, so this distortion costs 0.2200 - 0.2100 = 0.0100.

Worked steps

  1. If the event happens: (1 - 0.80)^2 = 0.0400.
  2. If it does not: 0.80^2 = 0.6400.
  3. Weight by the belief: 0.7 x 0.0400 + 0.3 x 0.6400 = 0.2200.

Use the idea

When forecasts are scored with a proper rule, report the number you actually believe; the rule removes the incentive to inflate or hedge, though calibration still has to be checked against outcomes.

Where the conclusion applies

The expectation is taken under the forecaster's own belief. On a handful of real outcomes, luck can make a distorted report score better, and an honest but overconfident belief is still miscalibrated.

Common wrong turn: Overstating confidence looks prescient
The chapter is explicit: believing 70 percent and reporting 80 percent raises the expected Brier loss from 0.21 to 0.22. Lower Brier scores are better, so the distortion hurts.
Check your understanding: With a belief of 0.7, what is the expected Brier loss of reporting 0.6?
0.7 x 0.16 + 0.3 x 0.36 = 0.22, the same extra 0.01 as reporting 0.8.

Chapter 17 source: section "Section Three: What Probabilistic Forecasting Actually Is".

Demonstration 2 of 4

An interval must earn its stated coverage

A forecaster calls an interval 80 percent. How do you find out whether it is?

Each point checks one nominal level against the outcomes. A narrow forecast sits below the dashed line at every level, a broad one above it. A prediction interval is a frequency statement about many cases like this one, which is why it can be checked only over many outcomes.

Equation: empirical coverage c hat equals the share of the n outcomes whose absolute value is at most s times the normal quantile at one plus c over two

Scroll sideways for the whole equation

The outcomes y are standard normal. The forecast believes they have spread s, so its central interval at nominal coverage c runs from minus s times z to plus s times z, z being the standard normal quantile at (1 + c) / 2. Empirical coverage c hat is the share of the n = 20000 outcomes inside.

Predict first. A forecast whose spread is only 0.65 of the true one states 80 percent intervals. How often do they cover?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: An interval must earn its stated coverage. Empirical against nominal coverage for forecast scale 0.65 over 20000 constructed outcomes; at nominal 0.80 the coverage is 0.5935. The other two scales are drawn faintly.
Forecast spread s: 0.65 (too narrow), Nominal coverage c: 0.8
Constructed data: the chapter notebook's 20,000 synthetic standard-normal outcomes (cell 3) and its three forecast scales.

Calculated values

Forecast scale s
0.65
Nominal coverage
0.80
Outcomes inside the interval
11870 of 20000
Empirical coverage
0.5935
Empirical minus nominal
-0.2065

At scale 0.65 the nominal 0.80 interval holds 11870 of 20000 outcomes: 11870 / 20000 = 0.5935, and 0.5935 - 0.80 = -0.2065. This forecast under-covers: the intervals are too narrow. Only the scale 1 forecast matches the outcomes it is scored on.

Worked steps

  1. Half-width: 0.65 times the standard normal quantile at 0.900.
  2. Count inside: 11870 of 20000, so 11870 / 20000 = 0.5935.
  3. Gap: 0.5935 - 0.80 = -0.2065.

Use the idea

Before trusting a stated 80 or 95 percent band, count how often past bands of the same kind held the outcome, level by level.

Where the conclusion applies

20,000 constructed independent standard-normal outcomes (seed 20260935) and intervals centred correctly; only the spread is wrong. Real errors are dependent and can be biased, and a few dozen outcomes cannot pin coverage down.

Check your understanding: Of 50 past 90 percent intervals, 38 held the outcome. What is the empirical coverage, and is the band too narrow or too wide?
38 / 50 = 0.76, and 0.76 - 0.90 = -0.14: the band under-covers, so it is too narrow.

Chapter 17 source: section "Section Four: The Full Methodology".

Demonstration 3 of 4

Calibrate the interval on an earlier block

How wide should a band be so that it misses about a share alpha of later outcomes?

A coefficient is fitted on the first block and frozen; its absolute one-step errors on a later calibration block set the radius; the band is then judged on a still later test block it never saw. The correction uses n + 1 and rounds up, which ordinary interpolated quantiles do not do.

Equation: the rank k equals n plus one times one minus alpha, rounded up

Scroll sideways for the whole equation

n is the number of calibration residuals (80 here), alpha the target miss rate, and k the rank of the residual used as the band's radius after sorting them from smallest to largest; the brackets mean round up.

Predict first. With 80 calibration residuals and alpha 0.1, which residual sets the radius?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Calibrate the interval on an earlier block. Test block of 81 constructed values with a band of plus or minus 1.8984 around the frozen one-step forecast; 4 outcomes fall outside, coverage 0.9506 against target 0.90.
Target miss rate alpha: 0.1, Radius rule: Corrected rank
Constructed data: the chapter notebook's seeded AR(1) series (cell 7), its training, calibration and test blocks, with the miss rate varied.

Calculated values

Miss rate alpha
0.10
Calibration residuals
80
Rank used
73
Radius
1.8984
Test outcomes covered
77 of 81
Test coverage
0.9506
Target 1 minus alpha
0.90

The corrected rank uses (80 + 1) x 0.90 = 72.90, rounded up to 73, so the radius is the 73th smallest of the 80 calibration residuals, 1.8984. On the later test block it covers 77 / 81 = 0.9506 against a target of 0.90. The residuals are serially dependent, so this is an empirical check, not a guarantee.

Worked steps

  1. Fit the coefficient on steps 0 to 79 and freeze it: 0.6212.
  2. Absolute one-step residuals on steps 80 to 159: 80 of them.
  3. Rank: (80 + 1) x 0.90 = 72.90, rounded up to 73; radius 1.8984.
  4. Test coverage on steps 160 to 240: 77 / 81 = 0.9506.

Use the idea

Keep three blocks in time order: fit, calibrate, test. Report the test coverage and the band width together, and never tune and claim coverage on the same outcomes.

Where the conclusion applies

A constructed AR(1) series with coefficient 0.7 and 240 shocks (seed 20260935). The finite-sample guarantee needs exchangeable scores; these residuals are serially dependent, so the chapter and the notebook assert no guarantee.

What this does not settle

The chapter gives split conformal its coverage guarantee only for exchangeable calibration and test scores, and notes that ordinary time-series dependence does not automatically satisfy those conditions. The coverage counted here is a result on one constructed series, not a guarantee.

Chapter 17 source: "Ordinary time-series dependence does not automatically satisfy those conditions".

Check your understanding: With 40 calibration residuals and alpha 0.2, which ranked residual does the corrected rule use?
41 x 0.8 = 32.8, rounded up to 33: the 33rd smallest residual.

Chapter 17 source: section "Section Four: The Full Methodology".

Demonstration 4 of 4

Width alone is not enough

Can a forecaster win an interval score by making the interval narrow, or by making it very wide?

The width term grows steadily with the interval, while the miss penalty shrinks as fewer outcomes fall outside. The sum is lowest near the interval that truly covers 80 percent, so neither gaming direction pays.

Equation: the interval score S equals the width u minus l, plus two over alpha times how far y falls below l, plus two over alpha times how far y falls above u

Scroll sideways for the whole equation

l and u are the interval's lower and upper ends, y the outcome, alpha the nominal miss rate (0.2 for an 80 percent interval). A plus subscript means keep the value if positive, else 0, so a miss costs 2 / alpha = 10 times its distance.

Predict first. Which of the offered widths gives the lowest mean score?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Width alone is not enough. Mean interval score against width for constructed standard-normal outcomes at nominal 80 percent, lowest near the true width 2.5631; width 1.8000 scores 3.8649.
Interval half-width: 0.9 (width 1.8)
Constructed data: the chapter notebook's 20,000 synthetic standard-normal outcomes (cells 3 and 5), scored at nominal 80 percent.

Calculated values

Interval width
1.8000
Coverage
0.6291
Mean width term
1.8000
Mean miss penalty
2.0649
Mean interval score
3.8649

The interval from -0.9000 to 0.9000 covers 0.6291 of the outcomes. Its mean score is width plus miss penalty: 1.8000 + 2.0649 = 3.8649, higher than the 3.5532 scored at the true 80 percent width.

Worked steps

  1. Width term: 2 x 0.9000 = 1.8000.
  2. Each miss costs 2 / 0.2 = 10 times its distance outside; averaged over 20000 outcomes: 2.0649.
  3. Mean score: 1.8000 + 2.0649 = 3.8649.

Use the idea

Score interval forecasts with width and misses together, and report coverage beside mean width; a band that always covers by being huge is not a good forecast.

Where the conclusion applies

20,000 constructed standard-normal outcomes (seed 20260935) and intervals centred at 0. The minimum sits at the true quantiles only when the outcomes follow the distribution assumed; with skewed outcomes it moves.

Check your understanding: For an 80 percent interval from 8 to 12 and an outcome of 14, what is the interval score?
(12 - 8) + 10 x (14 - 12) = 24.

Chapter 17 source: section "Section Four: The Full Methodology".