The Art and Science of Forecasting

Chapter 23

The Epidemiologist's Dilemma

Knowing what an epidemic is doing now and forecasting what it will do next are different problems, and each needs its own honest check.

Four demonstrations follow the chapter on a constructed epidemic: why the newest reports always look like a decline, how R0 sets the turning point of an SIR epidemic, what the exposed stage of an SEIR model changes, and how to score a fitted epidemic forecast against persistence on days after its origin.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

A falling report count may be a delay

When the most recent days show fewer reported cases, is the epidemic shrinking?

Each day's cases arrive over a week: 15 percent on the day, the rest later. At any cutoff the newest days have had the least time to fill in, so the report line always bends down at the right edge. Dividing by completeness rescales each day toward its eventual size.

Equation: the nowcast for event day d equals the count reported so far, r d, divided by the completeness c a at its age

Equation: completeness at age a is the sum of the delay probabilities p k for k from 0 to a

Scroll sideways for the whole equation

r is the count reported so far for event day d, a the age of that day at the cutoff (0 for the cutoff day itself), p the share of cases reported k days late, c the completeness (share expected to have arrived by age a), and n-hat the nowcast of the eventual count.

Predict first. With reports cut off at day 50, the epidemic is still growing. Will the reported counts for the last week rise or fall?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A falling report count may be a delay. Event days 45 to 65: eventual cases, cases reported by day 65, and the delay-adjusted nowcast; reports for the cutoff day are 173 against an eventual 1261.
Reports cut off at day: 65, Lines shown: Reports and nowcast
Constructed data: the chapter notebook's seeded SIR epidemic (seed 20260941) reported through its known seven-day delay law, with the cutoff varied; the notebook uses day 65.

Calculated values

Reported for the cutoff day
173
Completeness at age 0
0.15
Nowcast for the cutoff day
1153.3
Eventual count for the cutoff day
1261
Reported, six days before the cutoff
1814

As of day 65, the reports for the last seven days fall from 1814 to 173, while the eventual counts for those days are falling (1814 to 1261). Only 15 percent of a day's cases arrive on the day, so the nowcast is 173 / 0.15 = 1153.3 against an eventual 1261. The nowcast line divides each recent report by its completeness. The newest day is the least certain: a few reports more or less move it most.

Worked steps

  1. Delay law: 0.15 on the day, then 0.25, 0.25, 0.15, 0.10, 0.06, 0.04.
  2. Cutoff day has age 0, completeness 0.15; day 64 has age 1, completeness 0.15 + 0.25 = 0.40.
  3. Nowcast for day 65: 173 / 0.15 = 1153.3.
  4. Nowcast for day 64: 542 / 0.40 = 1355.0, eventual 1377.

Use the idea

Before reading the last few days of any lagged count as a trend, divide each day by the share of reports that have usually arrived by its age, and keep the newest days visibly uncertain.

Where the conclusion applies

The delay law is known and stable here. If reporting speeds up or slows down, a historical completeness factor fails; and dividing a small count by 0.15 multiplies its chance noise by more than six.

Check your understanding: A day has 60 reports and its completeness at this age is 0.40. What is the simple nowcast, and what would it be if completeness were really 0.25?
60 / 0.40 = 150 cases; with completeness 0.25 it is 60 / 0.25 = 240 cases. The nowcast is only as good as the delay law.

Chapter 23 source: section "Section Three: Nowcasting and the Real-Time Estimation Problem".

Demonstration 2 of 4

Behaviour sets the curve, and R0 sets the turning point

What decides when an epidemic turns, and how many people it reaches?

New infections flow at beta x S x I / N and removals at 0.1 x I. The infectious count grows while the inflow is larger, which is while S / N is above 1/R0; it peaks when S / N falls to about 1/R0. That share is the immunity threshold seen from the other side.

Equation: new infections per day equal beta times S times I divided by N

Equation: R nought equals beta divided by gamma

Equation: the immunity threshold H equals 1 minus 1 over R nought

Scroll sideways for the whole equation

S, I and N are the susceptible, infectious and total populations (N is 100,000). beta is the transmission rate per day, gamma the removal rate (0.1 per day, a ten day infectious period), R0 the basic reproduction number and H the immunity threshold.

Predict first. At beta 0.3, R0 is 3. What share of people are still susceptible when the infectious count peaks?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Behaviour sets the curve, and R0 sets the turning point. SIR curve with beta 0.22 per day (infectious count), peaking on day 58; the other two transmission scenarios are drawn faintly.
Transmission rate beta per day: 0.22, Curve shown: Currently infectious
Constructed data: the chapter notebook's closed-population SIR scenarios (beta 0.15, 0.22 and 0.3 per day, removal 0.1 per day, 100 initial cases in 100,000).

Calculated values

Transmission rate beta (per day)
0.22
R0 = beta / 0.1
2.2
Immunity threshold 1 - 1/R0
0.545
Peak infectious
18745
Peak day
58
Share susceptible at the peak
0.463
Share ever infected by day 120
0.839

R0 = 0.22 / 0.1 = 2.2, so the threshold is 1 - 1/2.2 = 0.545. The infectious count peaks on day 58, when 0.463 of people are still susceptible, close to 1/R0 = 0.455: once that many are immune each case infects fewer than one other, and the epidemic declines. By day 120, 0.839 have been infected, more than the threshold, because the decline does not stop new infections at once.

Worked steps

  1. R0 = beta / gamma = 0.22 / 0.1 = 2.2.
  2. Threshold = 1 - 1/2.2 = 0.545.
  3. At the peak (day 58) the susceptible share is 0.463, near 1/R0 = 0.455.
  4. Ever infected by day 120: 0.839.

Use the idea

When reading scenario curves, ask what transmission rate each assumes: the curves are conditional on behaviour, and a small change in beta moves both the timing and the size of the peak.

Where the conclusion applies

Homogeneous mixing, a closed population, no reinfection and a fixed transmission rate. The chapter names each as an assumption real epidemics break; these curves are scenarios, not estimated probabilities.

Check your understanding: If removal is 0.1 per day and beta is 0.25, what are R0 and the immunity threshold?
R0 = 0.25 / 0.1 = 2.5, and the threshold is 1 - 1/2.5 = 1 - 0.4 = 0.6, so 60 percent.

Chapter 23 source: section "Section Four: The Full Methodology".

Demonstration 3 of 4

The exposed stage puts a lag between infection and infectiousness

What changes when infected people cannot transmit at first?

New infections enter E, not I, and leave E at rate sigma. Cases therefore become infectious after a delay, the exposed curve peaks first, and with the same beta every chain of infection runs slower than in SIR.

Equation: the change in exposed people per day equals new infections beta S I over N minus sigma times E

Equation: the change in infectious people per day equals sigma times E minus gamma times I

Scroll sideways for the whole equation

E is the number exposed (infected, not yet infectious), I the number infectious, S the susceptible, N 100,000. beta is 0.22 per day, gamma 0.1 per day, and sigma the rate at which exposed people become infectious.

Predict first. Raise sigma from 0.25 to 1, shortening the mean latent period from 4 days to 1. Does the infectious peak come earlier or later?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The exposed stage puts a lag between infection and infectiousness. SEIR exposed and infectious curves with latent rate 0.25: exposed peak day 83, infectious peak day 92; the SIR infectious curve with the same transmission rate peaks on day 58.
Latent rate sigma per day: 0.25 (4 days), as in the notebook
Constructed data: the chapter notebook's SEIR example (beta 0.22, sigma 0.25, removal 0.1, 50 exposed and 100 infectious in 100,000), with sigma varied.

Calculated values

Latent rate sigma (per day)
0.25
Mean latent period, days
4
Exposed peak day
83
Infectious peak day
92
Infectious peak
13301
SIR peak day with the same beta
58

With sigma 0.25 the mean latent period is 1 / 0.25 = 4 days. The exposed count peaks on day 83 and the infectious count on day 92, 92 - 83 = 9 days later. Against the SIR model with the same beta, the infectious peak moves 92 - 58 = 34 days later and falls to 13301, because each generation of infection now takes longer.

Worked steps

  1. Mean latent period: 1 / 0.25 = 4 days.
  2. Exposed peak on day 83, infectious peak on day 92: 9 days apart.
  3. SIR peak on day 58: the latent stage delays the infectious peak by 34 days.
  4. People are conserved: S + E + I + R stays 100,000 on every day.

Use the idea

When cases are used to infer transmission, remember that what is observed lags the infections that caused it; and do not set sigma from the symptom incubation period without checking for presymptomatic transmission.

Where the conclusion applies

Sigma is the rate of becoming infectious, not the inverse of the incubation period to symptoms; the chapter warns the two are not interchangeable. Transmission, removal and the closed population are assumed.

Check your understanding: If the mean latent period is 5 days, what is sigma?
sigma = 1 / 5 = 0.2 per day.

Chapter 23 source: section "Section Four: The Full Methodology".

Demonstration 4 of 4

Score the epidemic forecast after its origin

How should a mechanistic epidemic forecast be judged against a naive one?

At each origin the transmission rate is fitted to the counts known by then, the model is run forward, and both forecasts are scored only on later days. Here the fitted model wins at every origin and horizon, because it is the model that made the data; persistence falls further behind the further ahead it looks, since it cannot know the curve will turn.

Equation: the RMSE at horizon h is the square root of the mean over four origins of the squared forecast errors

Scroll sideways for the whole equation

o is a forecast origin (days 30, 45, 60 and 75), h the number of days ahead, and e the forecast minus the measured count. RMSE is the root mean squared error over the four origins at one horizon.

Predict first. From origin day 45 (near the peak), which forecast is closer 14 days ahead: the fitted SIR or persistence?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Score the epidemic forecast after its origin. Measured infectious counts with the forecast origin at day 45; the fitted SIR projection and the flat persistence forecast run 14 days ahead, with day 52 circled.
Forecast origin day: 45, Days ahead: 7
Constructed data: the chapter notebook's noisy synthetic infectious counts (SIR with beta 0.22 plus seeded noise of standard deviation 25) and its past-only fits at four origins.

Calculated values

Past-only fitted beta
0.2200
Measured count, day 52
17111.1
Fitted SIR absolute error
2.4
Persistence absolute error
4690.7
Fitted SIR RMSE over 4 origins
15.0
Persistence RMSE over 4 origins
3544.6

From day 45 the fit uses only counts up to day 45 and finds beta 0.2200 (the true value is 0.22). 7 days ahead, the fitted SIR misses by 2.4 and persistence by 4690.7, so the fitted SIR is closer: 4690.7 - 2.4 = 4688.3. Over all four origins the RMSE at this horizon is 15.0 against 3544.6. This is the most favourable case possible: the model is the true one, with removal and starting state known. Real reports are not infectious prevalence.

Worked steps

  1. Fit beta on days 0 to 45 only: 0.2200.
  2. Forecast error at 7 days: SIR 2.4, persistence 4690.7.
  3. Same horizon, four origins: RMSE = sqrt(mean of squared errors): SIR 15.0, persistence 3544.6.
  4. Scoring only future days keeps the forecast honest; the nowcast error of the first demonstration is a different question.

Use the idea

Score any epidemic model on days after its origin, against a simple benchmark, at several origins and horizons; never on the data it was fitted to.

Where the conclusion applies

The notebook calls this an unusually favourable, correctly specified experiment: the data come from the very model being fitted. With real reports, wrong structure or shifting behaviour, the mechanistic edge can vanish: a model fitted to flat counts, for example, can lose badly to persistence.

Check your understanding: At one horizon the four origin errors of a forecast are 10, minus 20, 20 and minus 10. What is its RMSE?
sqrt((100 + 400 + 400 + 100) / 4) = sqrt(250) = 15.8.

Chapter 23 source: section "Section Four: The Full Methodology".