The Art and Science of Forecasting

Chapter 11

The Crowd and the Ox

A crowd's aggregate can beat its members, but only when their errors are independent and the aggregation is done with care.

Four demonstrations follow the chapter: why the aggregate of many rough estimates lands close to the truth, why an error the crowd shares sets a floor that no crowd size removes, how the mean, median and trimmed mean respond to wild guesses, and when weights learned from past accuracy beat the simple average.

Most examples are constructed teaching data, generated with a fixed seed so that every number matches the chapter notebook. Where a demonstration uses a real historical series, such as the annual flow of the Nile, it says so and names the source. Nothing here is a forecast of any real market, product or person.

Demonstration 1 of 4

Every guess is off, the crowd is close

How can a crowd whose members miss by over a hundred pounds land within a few pounds of the truth?

Each bar counts estimates in a 40 lb band. Individual errors spread widely on both sides of the truth; because they are independent, the high and low ones largely offset in the mean or median.

Equation: x bar equals one over n times the sum of the n estimates x i

Scroll sideways for the whole equation

x i is the weight written by person i, n the number of estimates, and x bar their mean. The median is the middle estimate once they are sorted. Weights are in pounds (lb); the truth is 1200 lb.

Predict first. With all 800 estimates, is the crowd median closer to the truth than the average person in the crowd?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Every guess is off, the crowd is close. Histogram of 800 constructed weight estimates spread from about 700 to 1700 lb, with the truth at 1200 lb and the crowd median at 1186.9 lb.
Number of estimates: 800, Aggregate: Median
Constructed data: the chapter notebook's 800 synthetic weight estimates (cell 1), not the historical ox competition.

Calculated values

Estimates
800
Crowd median
1186.9 lb
Aggregate error
13.1 lb
Average individual error
142.8 lb

The median of the first 800 estimates is 1186.9 lb, below the truth: 1200.0 - 1186.9 = 13.1 lb. The average individual estimate misses by 142.8 lb, so the aggregate is much closer than a typical member of the crowd. Larger crowds shrink the typical aggregate error, but any one sample still lands a chance distance from the truth.

Worked steps

  1. Aggregate: median of 800 estimates = 1186.9 lb.
  2. Aggregate error: 1200.0 - 1186.9 = 13.1 lb.
  3. Average individual error: 142.8 lb.

Use the idea

When you need one number from several people, collect written estimates privately and report their median or mean alongside how far apart the individuals were.

Where the conclusion applies

800 constructed estimates drawn independently around 1200 lb with standard deviation 180 lb (seed 20260929). Every person is unbiased; a shared bias would move the aggregate with it.

Check your understanding: Five people guess 1150, 1180, 1210, 1260 and 1300 lb for an ox that weighs 1200 lb. What is the median, and how far off is it?
Sorted, the middle value is 1210 lb; 1210 - 1200 = 10 lb.

Chapter 11 source: section "Surowiecki's Framework".

Demonstration 2 of 4

Shared error does not average away

If everyone's errors are partly the same error, how much does adding more people help?

Each solid line is the measured error of the crowd mean as the crowd grows, for one error correlation; the dotted line is the formula for the chosen correlation. The independent part of each error shrinks with n; the shared part, rho, does not shrink at all.

Equation: the standard deviation of the crowd mean equals sigma times the square root of rho plus one minus rho over n

Scroll sideways for the whole equation

sigma is one person's error standard deviation (180 lb), rho the correlation between two people's errors, n the crowd size, and SD of x bar the spread of the crowd mean's error. RMSE is the root mean squared error of the crowd mean over 400 repeated panels.

Predict first. At correlation 0.7, does growing the crowd from 10 to 300 people cut the error of the mean by much?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Shared error does not average away. Error of the crowd mean against crowd size for correlations 0, 0.2 and 0.7; correlation 0.2 is highlighted and the crowd of 10 has error 92.90 lb.
Correlation of errors: 0.2, Crowd size: 10
Constructed data: the chapter notebook's repeated panels with correlated errors (cell 3), 400 panels per point.

Calculated values

Error correlation
0.2
Crowd size
10
Simulated RMSE (400 panels)
92.90
Formula
95.25
Floor as the crowd grows
80.50

With correlation 0.2 and 10 people: 180 x sqrt(0.2 + 0.8 / 10) = 95.25 lb, against 92.90 lb measured over 400 constructed panels. However large the crowd, the formula (dotted) never drops below 80.50 lb; single simulated points scatter around it.

Worked steps

  1. Shared part: 0.2; independent part: 0.8 / 10 = 0.08.
  2. Formula: 180 x sqrt(0.2 + 0.08) = 95.25 lb.
  3. Floor: 180 x sqrt(0.2) = 80.50 lb.

Use the idea

Before counting heads, ask where the estimates came from: people who read the same report add little beyond the first of them.

Where the conclusion applies

Each person's error has standard deviation 180 lb and the same correlation with everyone else's, as in the notebook (seed 20260929). The simulated points scatter around the formula because 400 panels is a finite sample.

Common wrong turn: A bigger crowd fixes any error
The chapter is explicit that the crowd cannot correct for errors it shares: with correlation rho above zero, the expected error of the mean never falls below sigma x sqrt(rho), however many people join.
Check your understanding: Individual error SD is 10, the correlation is 0.2 and the crowd has 10 people. What is the SD of the crowd mean?
10 x sqrt(0.2 + 0.8 / 10) = 10 x sqrt(0.28) = 5.29, not 10 / sqrt(10) = 3.16.

Chapter 11 source: section "The Full Machinery".

Demonstration 3 of 4

Which aggregate survives wild guesses?

When some estimates are wildly high, should the crowd be summarized by its mean, its median or a trimmed mean?

The left panel is one panel of 100 estimates; the right panel is the average miss of each aggregate over 300 such panels as the share of estimates raised by 1000 lb grows.

Equation: mean absolute error equals the average of the absolute differences between the truth y t and the aggregate y hat t

Scroll sideways for the whole equation

y is the truth (1200 lb) and y hat the aggregate of one panel of 100 estimates; MAE averages the absolute miss over 300 panels. The trimmed mean drops the 20 lowest and 20 highest estimates and averages the rest.

Predict first. With no wild guesses at all, which aggregate has the smallest error?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Which aggregate survives wild guesses?. Left, one constructed panel of 100 estimates with 10 raised by 1000 lb and its mean, median and trimmed mean marked. Right, aggregate MAE against the contaminated share: at 10 percent the mean scores 98.49, the median 14.75 and the trimmed mean 15.62 lb.
Contaminated estimates (of 100): 10
Constructed data: the chapter notebook's contaminated panels (cell 4), 300 panels of 100 estimates per contamination level.

Calculated values

Contaminated estimates
10 of 100
Mean MAE
98.49
Median MAE
14.75
Trimmed mean MAE
15.62
Lowest MAE
Median

10 of 100 estimates are raised by 1000 lb, which moves the mean by about 10 x 1000 / 100 = 100 lb; its MAE is 98.49 lb. The median scores 14.75 lb and the trimmed mean 15.62 lb, so the lowest is the median. The trimmed mean drops 20 estimates from each tail, which removes every outlier here.

Worked steps

  1. Shift of the mean: 10 x 1000 / 100 = 100 lb.
  2. Mean MAE 98.49, median 14.75, trimmed mean 15.62 lb.
  3. Lowest: median.

Use the idea

Report the mean, the median and a trimmed mean side by side; when they disagree, find the estimates that pull them apart before choosing one.

Where the conclusion applies

Constructed panels: 100 estimates around 1200 lb with standard deviation 100 lb, a share of them raised by 1000 lb (seed 20260929). This tests upward extremes only; it says nothing about a moderate bias shared by everyone.

What this does not settle

No aggregate here corrects a shared error: if every estimate is pushed the same way, the mean, the median and the trimmed mean all move with it. As the chapter puts it, the aggregate of biased individual estimates is a biased aggregate.

Chapter 11 source: "The aggregate of biased individual estimates is a biased aggregate".

Check your understanding: Four estimates are 90, 100, 110 and 1000. What are the mean, the median, and the mean after trimming one value from each tail?
Mean (90 + 100 + 110 + 1000) / 4 = 325; median (100 + 110) / 2 = 105; trimmed (100 + 110) / 2 = 105.

Chapter 11 source: section "The Full Machinery".

Demonstration 4 of 4

Learned weights or the simple average?

Should forecasts be combined with weights learned from past accuracy, or simply averaged?

Inverse error weights give most of the weight to the least noisy expert. Whether that beats the simple average depends on how stable the skill differences are and on judging the weights on questions they never saw.

Equation: the weight on expert j equals one over that expert's mean squared error, divided by the sum of one over every expert's mean squared error

Equation: mean absolute error equals the average of the absolute differences between the truth y t and the aggregate y hat t

Scroll sideways for the whole equation

w j is the weight on expert j, and MSE j that expert's mean squared error on earlier resolved questions. Equal weighting gives each of the 8 experts 1 / 8. MAE is the mean absolute error of the combined forecast.

Predict first. Scored on the same 80 questions used to learn the weights, does the learned combination look better or worse than on the 40 later questions?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Learned weights or the simple average?. Left, learned weights for eight constructed experts from 80 training questions, largest 0.626. Right, MAE on 40 later questions: equal weights 4.81, learned weights 3.31.
Questions used to learn the weights: 80, Scored on: 40 later questions
Constructed data: the chapter notebook's eight synthetic experts on 120 questions (cell 7); weights learned on the first questions, the last 40 kept for scoring.

Calculated values

Training questions
80
Scored on
40 later questions
Equal weights MAE
4.81
Learned weights MAE
3.31
Equal minus learned
1.50
Largest weight
0.626

Weights learned from 80 questions, scored on 40 later questions: equal weights 4.81, learned weights 3.31, so 4.81 - 3.31 = 1.50 and learned weights score lower. These experts differ persistently in skill, from noise SD 5 to 30. The weights were frozen before these questions were scored.

Worked steps

  1. Equal weight per expert: 1 / 8 = 0.125.
  2. Largest learned weight: 0.626, for the expert with noise SD 5.
  3. Gain: 4.81 - 3.31 = 1.50.

Use the idea

Start from the simple average; adopt learned weights only if they beat it on later questions held out before the weights were fitted.

Where the conclusion applies

Eight constructed experts whose noise SD runs from 5 to 30 and never changes (seed 20260929). Here that spread is large enough that learned weights win even from 3 questions; when skill differences are small or shift over time, weight estimation error can make equal weights the better choice, as the chapter describes.

What this does not settle

This generator gives the experts large, permanent skill differences, which favours learned weights. The chapter's point is that with limited data the estimated weights are noisy, so a simple average should be the baseline any weighting scheme has to beat out of sample.

Chapter 11 source: "The simple average: surprisingly hard to beat".

Check your understanding: Two forecasters had MSE 4 and 16 on past questions. What inverse MSE weights do they get?
1 / 4 = 0.25 and 1 / 16 = 0.0625; total 0.3125. Weights 0.25 / 0.3125 = 0.8 and 0.0625 / 0.3125 = 0.2.

Chapter 11 source: section "The Full Machinery".