The Art and Science of Forecasting

Chapter 8

The Delphi Room

Delphi removes the room's social pressure, but convergence is not the same as being right.

Four demonstrations follow the chapter: how anonymous median feedback narrows a panel, why that agreement need not mean accuracy, why a bigger panel cannot average away an error its members share, and how forecast value added scores whether each revision helped.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Anonymous feedback narrows the panel

When experts see the median and the interquartile range and may revise, what happens to their spread?

Each dot is one expert's anonymous estimate; the box marks the quartiles and the bar the median of the highlighted round. Because everyone moves toward the same median, the box tightens round by round.

Equation: the interquartile range equals the 75th percentile minus the 25th percentile

Equation: expert i's next estimate equals 1 minus w times their estimate, plus w times the round's median, plus a small private change

Scroll sideways for the whole equation

The lower and upper quartiles are the 25th and 75th percentiles of the 30 estimates; their difference is the interquartile range. In the revision rule each expert keeps 1 minus w of their own estimate, moves w of the way to the round's median, and adds a small private change (standard deviation 1).

Predict first. With the notebook's pull of 0.4, will the interquartile range in round 3 be less than half of round 1's?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Anonymous feedback narrows the panel. Thirty constructed expert estimates for rounds 1 to 5, with round 1 highlighted: median 116.80, interquartile range 26.43, all well above the dashed true value of 100.
Round highlighted: 1, Pull toward the median, w: 0.4
Constructed data: the chapter notebook's 30 synthetic experts and four feedback rounds (pull 0.4), with the pull varied.

Calculated values

Round
1
Median
116.80
Lower quartile
99.93
Upper quartile
126.36
Interquartile range
26.43
Median minus true value
16.80

Round 1: IQR = 126.36 - 99.93 = 26.43, against 26.43 in round 1. The median sits at 116.80, so median minus truth is 116.80 - 100 = 16.80. This is the first, independent round: no one has seen anyone else's number yet.

Worked steps

  1. Upper quartile 126.36, lower quartile 99.93.
  2. IQR: 126.36 - 99.93 = 26.43.
  3. Median minus truth: 116.80 - 100 = 16.80.

Use the idea

When you run a structured elicitation, keep the first round's estimates: they are the only independent ones, and later rounds can only be judged against them.

Where the conclusion applies

Experts are constructed (seed 20260926): they start 15 units above the truth with spread 18, and the pull toward the median is assumed, not measured from any panel. Real outliers may hold their ground with reasons, which this rule ignores.

Check your understanding: A round has quartiles 112.40 and 98.10. What is its interquartile range?
112.40 - 98.10 = 14.30.

Chapter 8 source: section "Delphi in the Wild".

Demonstration 2 of 4

Agreement is not accuracy

If a panel converges, has it become more accurate?

The solid line is disagreement and the dashed line is the group's error. Feedback drives the first down; the second stays where the shared bias put it, because the median the experts move toward is itself biased.

Equation: expert i's next estimate equals 1 minus w times their estimate, plus w times the round's median, plus a small private change

Scroll sideways for the whole equation

The shared starting bias is how far above the truth the whole panel leans before round 1. Spread is the interquartile range; the median's error is its distance from the true value 100.

Predict first. With a shared bias of 15, will five rounds of feedback cut the median's error by more than 5 units?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Agreement is not accuracy. Over five rounds the interquartile range falls from 26.43 to 3.83 while the median's distance from the truth stays near 16.02.
Shared starting bias: 15, Pull toward the median, w: 0.4
Constructed data: the chapter notebook's synthetic panel (shared bias 15, pull 0.4), with the bias and pull varied.

Calculated values

Shared starting bias
15
Round 1 spread
26.43
Round 5 spread
3.83
Round 1 median error
16.80
Round 5 median error
16.02
Improvement in median error
0.78

Spread falls from 26.43 to 3.83, but the median's error only changes by 16.80 - 16.02 = 0.78. Every expert started 15 units too high, and feedback cannot remove an error the whole group shares.

Worked steps

  1. Spread: 26.43 in round 1, 3.83 in round 5.
  2. Median error: 16.80 in round 1, 16.02 in round 5.
  3. Improvement: 16.80 - 16.02 = 0.78.

Use the idea

Never report a narrow final Delphi range as evidence of accuracy; compare the first-round and final medians with outcomes before crediting the process.

Where the conclusion applies

Constructed experts (seed 20260926) with a bias of 0, 15 or 30 units added to everyone. The chapter's limit applies: a panel that shares a misconception converges on a wrong answer with confidence.

What this does not settle

The procedure filters out social noise; it does not add knowledge that is not already in the group, and it cannot tell genuine consensus from collective error.

Chapter 8 source: "It does not add knowledge that is not already in the group".

Check your understanding: A panel's median error is 12.40 in round 1 and 11.90 in round 4, while its IQR falls from 20 to 5. How much did feedback improve the median?
12.40 - 11.90 = 0.50 units. The spread fell by three quarters; the accuracy barely changed.

Chapter 8 source: section "The Cognitive Problem Delphi Did Not Solve".

Demonstration 3 of 4

More experts cannot average away a shared error

If every expert's error has the same size, does it matter whether part of it is shared?

The solid line is a panel whose errors are independent: it falls like one over the square root of n. The dashed line is a panel whose members share part of their error: averaging removes the private part and leaves the shared part.

Equation: the error of an n expert average equals the square root of the shared variance plus the private variance over n

Scroll sideways for the whole equation

n is the number of experts. Each expert's error has standard deviation 15: either all private, or a common part with standard deviation 12 (sigma c) plus a private part with standard deviation 9 (sigma u), since 15^2 = 12^2 + 9^2. RMSE is the root mean squared error of the panel average over 500 simulated panels.

Predict first. With 50 experts sharing part of their error, will the panel average's error fall below 5?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: More experts cannot average away a shared error. Panel error against size: the independent line falls toward 2 while the shared-error line stays near 12; at 10 experts they are 4.91 and 12.51.
Experts on the panel: 10
Constructed data: the chapter notebook's repeated synthetic panels (500 per size), equal individual error 15 with or without a shared part of 12.

Calculated values

Panel size
10
Independent, simulated
4.91
Independent, 15 / sqrt(n)
4.74
Shared, simulated
12.51
Shared, sqrt(144 + 81 / n)
12.33

With 10 experts, independent errors of 15 average down to 15 / sqrt(10) = 4.74 (simulated 4.91). When 12 of each expert's error is common to all, the floor stays: sqrt(144 + 81 / 10) = 12.33 (simulated 12.51). Adding voices removes only the part of the error they do not share.

Worked steps

  1. Independent: 15 / sqrt(10) = 4.74.
  2. Shared part 12, private part 0.6 x 15 = 9: sqrt(12^2 + 9^2 / 10) = 12.33.
  3. Simulated over 500 panels: 4.91 and 12.51.

Use the idea

Before adding experts to a panel, ask whether they draw on independent information; a larger panel of people reading the same reports buys almost nothing.

Where the conclusion applies

Errors are normal and the shared part is identical for every member (seed 20260926 draws, 500 panels per size). Real panels share some information partly, which puts the floor between the two lines.

Check your understanding: Each expert's error has standard deviation 10, of which a common part has standard deviation 8 and a private part 6. What is the panel-average error with 4 experts?
sqrt(8^2 + 6^2 / 4) = sqrt(64 + 9) = sqrt(73) = 8.54.

Chapter 8 source: section "The Full Machinery".

Demonstration 4 of 4

Did the revisions add value?

How do you tell whether a round of revision made a forecast better or worse?

Each dot is one expert; the diamonds join the first and the chosen later median. Value added is measured against the outcome, never against the spread, so a revision can converge and still add nothing.

Equation: forecast value added equals the first median's absolute error minus the later median's absolute error

Scroll sideways for the whole equation

a is the outcome once known, m1 the first-round median and mR the median after round R. Forecast value added is the first-round error minus the later error: positive means the revision helped.

Predict first. For question 5 (outcome 140), will the final-round median have a smaller error than the first-round median?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Did the revisions add value?. Twelve constructed estimates per round for question 5; the median moves from 154.05 in round 1 to 146.88 in round 3, against an outcome of 140.
Question: Question 5, Later round compared: 3
Constructed data: the companion's seeded example panel for this chapter (8 questions, 12 experts, 3 rounds), scored with the companion's Delphi round comparison.

Calculated values

Outcome
140
Round 1 median
154.05
Round 3 median
146.88
Round 1 error
14.05
Round 3 error
6.88
Forecast value added
7.17

Question 5: FVA = |154.05 - 140| - |146.88 - 140| = 14.05 - 6.88 = 7.17, so the revisions added value by that measure. Across all 8 example questions the round 3 median beat the round 1 median on 8. This works only because round 1 was recorded before anyone saw the feedback.

Worked steps

  1. Round 1 error: |154.05 - 140| = 14.05.
  2. Round 3 error: |146.88 - 140| = 6.88.
  3. FVA: 14.05 - 6.88 = 7.17.

Use the idea

Record the statistical or first-round forecast and every adjusted version, then score each against outcomes; an adjustment that does not add value on average should not be made by default.

Where the conclusion applies

The example panel is constructed so its medians drift toward the outcomes; real adjustments often do the opposite. Eight questions are too few to judge any process.

Check your understanding: A first-round median of 210 and a final median of 196 forecast an outcome of 200. What is the forecast value added?
|210 - 200| - |196 - 200| = 10 - 4 = 6. The revision added 6 units of value.

Chapter 8 source: section "The Full Machinery".