Demonstration 1 of 4
Every guess is off, the crowd is close
How can a crowd whose members miss by over a hundred pounds land within a few pounds of the truth?
Each bar counts estimates in a 40 lb band. Individual errors spread widely on both sides of the truth; because they are independent, the high and low ones largely offset in the mean or median.
Scroll sideways for the whole equation
x i is the weight written by person i, n the number of estimates, and x bar their mean. The median is the middle estimate once they are sorted. Weights are in pounds (lb); the truth is 1200 lb.
Predict first. With all 800 estimates, is the crowd median closer to the truth than the average person in the crowd?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's 800 synthetic weight estimates (cell 1), not the historical ox competition.
Calculated values
- Estimates
- 800
- Crowd median
- 1186.9 lb
- Aggregate error
- 13.1 lb
- Average individual error
- 142.8 lb
The median of the first 800 estimates is 1186.9 lb, below the truth: 1200.0 - 1186.9 = 13.1 lb. The average individual estimate misses by 142.8 lb, so the aggregate is much closer than a typical member of the crowd. Larger crowds shrink the typical aggregate error, but any one sample still lands a chance distance from the truth.
Worked steps
- Aggregate: median of 800 estimates = 1186.9 lb.
- Aggregate error: 1200.0 - 1186.9 = 13.1 lb.
- Average individual error: 142.8 lb.
Use the idea
When you need one number from several people, collect written estimates privately and report their median or mean alongside how far apart the individuals were.
Where the conclusion applies
800 constructed estimates drawn independently around 1200 lb with standard deviation 180 lb (seed 20260929). Every person is unbiased; a shared bias would move the aggregate with it.
Check your understanding: Five people guess 1150, 1180, 1210, 1260 and 1300 lb for an ox that weighs 1200 lb. What is the median, and how far off is it?
Chapter 11 source: section "Surowiecki's Framework".
Demonstration 2 of 4
Shared error does not average away
If everyone's errors are partly the same error, how much does adding more people help?
Each solid line is the measured error of the crowd mean as the crowd grows, for one error correlation; the dotted line is the formula for the chosen correlation. The independent part of each error shrinks with n; the shared part, rho, does not shrink at all.
Scroll sideways for the whole equation
sigma is one person's error standard deviation (180 lb), rho the correlation between two people's errors, n the crowd size, and SD of x bar the spread of the crowd mean's error. RMSE is the root mean squared error of the crowd mean over 400 repeated panels.
Predict first. At correlation 0.7, does growing the crowd from 10 to 300 people cut the error of the mean by much?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's repeated panels with correlated errors (cell 3), 400 panels per point.
Calculated values
- Error correlation
- 0.2
- Crowd size
- 10
- Simulated RMSE (400 panels)
- 92.90
- Formula
- 95.25
- Floor as the crowd grows
- 80.50
With correlation 0.2 and 10 people: 180 x sqrt(0.2 + 0.8 / 10) = 95.25 lb, against 92.90 lb measured over 400 constructed panels. However large the crowd, the formula (dotted) never drops below 80.50 lb; single simulated points scatter around it.
Worked steps
- Shared part: 0.2; independent part: 0.8 / 10 = 0.08.
- Formula: 180 x sqrt(0.2 + 0.08) = 95.25 lb.
- Floor: 180 x sqrt(0.2) = 80.50 lb.
Use the idea
Before counting heads, ask where the estimates came from: people who read the same report add little beyond the first of them.
Where the conclusion applies
Each person's error has standard deviation 180 lb and the same correlation with everyone else's, as in the notebook (seed 20260929). The simulated points scatter around the formula because 400 panels is a finite sample.
Common wrong turn: A bigger crowd fixes any error
Check your understanding: Individual error SD is 10, the correlation is 0.2 and the crowd has 10 people. What is the SD of the crowd mean?
Chapter 11 source: section "The Full Machinery".
Demonstration 3 of 4
Which aggregate survives wild guesses?
When some estimates are wildly high, should the crowd be summarized by its mean, its median or a trimmed mean?
The left panel is one panel of 100 estimates; the right panel is the average miss of each aggregate over 300 such panels as the share of estimates raised by 1000 lb grows.
Scroll sideways for the whole equation
y is the truth (1200 lb) and y hat the aggregate of one panel of 100 estimates; MAE averages the absolute miss over 300 panels. The trimmed mean drops the 20 lowest and 20 highest estimates and averages the rest.
Predict first. With no wild guesses at all, which aggregate has the smallest error?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's contaminated panels (cell 4), 300 panels of 100 estimates per contamination level.
Calculated values
- Contaminated estimates
- 10 of 100
- Mean MAE
- 98.49
- Median MAE
- 14.75
- Trimmed mean MAE
- 15.62
- Lowest MAE
- Median
10 of 100 estimates are raised by 1000 lb, which moves the mean by about 10 x 1000 / 100 = 100 lb; its MAE is 98.49 lb. The median scores 14.75 lb and the trimmed mean 15.62 lb, so the lowest is the median. The trimmed mean drops 20 estimates from each tail, which removes every outlier here.
Worked steps
- Shift of the mean: 10 x 1000 / 100 = 100 lb.
- Mean MAE 98.49, median 14.75, trimmed mean 15.62 lb.
- Lowest: median.
Use the idea
Report the mean, the median and a trimmed mean side by side; when they disagree, find the estimates that pull them apart before choosing one.
Where the conclusion applies
Constructed panels: 100 estimates around 1200 lb with standard deviation 100 lb, a share of them raised by 1000 lb (seed 20260929). This tests upward extremes only; it says nothing about a moderate bias shared by everyone.
What this does not settle
No aggregate here corrects a shared error: if every estimate is pushed the same way, the mean, the median and the trimmed mean all move with it. As the chapter puts it, the aggregate of biased individual estimates is a biased aggregate.
Chapter 11 source: "The aggregate of biased individual estimates is a biased aggregate".
Check your understanding: Four estimates are 90, 100, 110 and 1000. What are the mean, the median, and the mean after trimming one value from each tail?
Chapter 11 source: section "The Full Machinery".
Demonstration 4 of 4
Learned weights or the simple average?
Should forecasts be combined with weights learned from past accuracy, or simply averaged?
Inverse error weights give most of the weight to the least noisy expert. Whether that beats the simple average depends on how stable the skill differences are and on judging the weights on questions they never saw.
Scroll sideways for the whole equation
w j is the weight on expert j, and MSE j that expert's mean squared error on earlier resolved questions. Equal weighting gives each of the 8 experts 1 / 8. MAE is the mean absolute error of the combined forecast.
Predict first. Scored on the same 80 questions used to learn the weights, does the learned combination look better or worse than on the 40 later questions?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's eight synthetic experts on 120 questions (cell 7); weights learned on the first questions, the last 40 kept for scoring.
Calculated values
- Training questions
- 80
- Scored on
- 40 later questions
- Equal weights MAE
- 4.81
- Learned weights MAE
- 3.31
- Equal minus learned
- 1.50
- Largest weight
- 0.626
Weights learned from 80 questions, scored on 40 later questions: equal weights 4.81, learned weights 3.31, so 4.81 - 3.31 = 1.50 and learned weights score lower. These experts differ persistently in skill, from noise SD 5 to 30. The weights were frozen before these questions were scored.
Worked steps
- Equal weight per expert: 1 / 8 = 0.125.
- Largest learned weight: 0.626, for the expert with noise SD 5.
- Gain: 4.81 - 3.31 = 1.50.
Use the idea
Start from the simple average; adopt learned weights only if they beat it on later questions held out before the weights were fitted.
Where the conclusion applies
Eight constructed experts whose noise SD runs from 5 to 30 and never changes (seed 20260929). Here that spread is large enough that learned weights win even from 3 questions; when skill differences are small or shift over time, weight estimation error can make equal weights the better choice, as the chapter describes.
What this does not settle
This generator gives the experts large, permanent skill differences, which favours learned weights. The chapter's point is that with limited data the estimated weights are noisy, so a simple average should be the baseline any weighting scheme has to beat out of sample.
Chapter 11 source: "The simple average: surprisingly hard to beat".
Check your understanding: Two forecasters had MSE 4 and 16 on past questions. What inverse MSE weights do they get?
Chapter 11 source: section "The Full Machinery".