Demonstration 1 of 4
Sampling fills the bank, but a blind pick ignores it
How many samples does it take to have a correct answer somewhere, what does a random pick from them deliver, and what if the chance is concentrated on a few problems?
Coverage is one minus the chance that every sample is wrong, so it climbs fastest over the first few samples, faster for larger p, and flattens as k grows. The blind pick returns a correct sample with probability p on average, whatever k is, because the k in the fraction m / k cancels. The gap between the two lines is the most a selector could ever add, and the right panel shows the added coverage from each sample shrinking geometrically. Choose the overconfident setting to see the chapter's mechanism: the average chance of one sample is unchanged, so the single answer looks as good, but chance concentrated on some problems and absent on others leaves a thin bank whose coverage stops rising long before Equation (16.1) says.
Scroll sideways for the whole equation
p is the chance that one sample is correct. k is the number of samples drawn, assumed independent. Cov(k) is coverage: the chance at least one of the k samples is correct. Sel(k) is the chance the returned sample is correct; with no way to rank, the system returns one at random. m is the number of correct samples among the k. The right panel shows the extra coverage bought by sample number k + 1. In the overconfident setting, half the problems have chance 2p per sample and half have chance 0, so the average is still p.
Predict first. At p = 0.05 with every problem sharing p, how likely is it that ten samples hold a correct answer?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's p = 0.2, 0.05 and 0.3 arithmetic (coverage 0.89 and 0.99 at 10 and 20 samples for p = 0.2, 0.40 and 0.99 at 10 and 100 for p = 0.05, gains 0.30, 0.07 and under 0.001 at samples 1, 5 and 20 for p = 0.3), computed directly from the equations; the overconfident setting is defined for this reader.
Calculated values
- Coverage Cov(k)
- 0.8926
- Blind selection Sel(k)
- 0.20
- Gap a selector could close
- 0.6926
- Gain from sample 11
- 0.0215
- Coverage at 20 samples
- 0.9885
- Gain from sample 5
- 0.0819
- Gain from sample 20
- 0.0029
Cov(10) = 1 - (1 - 0.20)^10 = 1 - 0.1074 = 0.8926. A blind selector returns one sample at random, so it delivers p = 0.20 however large k is. The gap a selector could close is 0.8926 - 0.20 = 0.6926. The next sample adds 0.20 x (1 - 0.20)^10 = 0.20 x 0.1074 = 0.0215. That equals Cov(11) - Cov(10) = 0.9141 - 0.8926 = 0.0215.
Worked steps
- Chance one sample is wrong: 1 - p = 1 - 0.20 = 0.80.
- All 10 samples wrong: 0.80^10 = 0.1074.
- Coverage Cov(10) = 1 - 0.1074 = 0.8926.
- A blind pick delivers p = 0.20 whatever k is (Equation 16.2: the k in m / k cancels).
- Gap a selector could close: 0.8926 - 0.20 = 0.6926.
- Sample 11 adds p x (1 - p)^10 = 0.20 x 0.1074 = 0.0215.
Use the idea
Before buying more samples for a task, ask what picks among them. If the answer is nothing, the extra samples raise the number of correct candidates but not the number delivered. Measure bank coverage directly, by sampling repeated banks under one declared procedure, rather than inferring it from p.
Where the conclusion applies
Within each problem, samples are independent and share one correctness chance. The overconfident setting is a constructed illustration of the chapter's statement that p can be near zero for many problems, not a model of any training run; the chapter's Figure 16.1 is schematic and compares different sampling procedures, so it does not isolate candidate count as the cause. The values of p are constructed; they do not describe any real model.
Common wrong turn: More samples deliver a better answer
What this does not settle
The chapter's results come from an unrefereed preprint reporting on its authors' own models on one dataset, and the thirty-fold figure is their comparison rather than an established rate. Equation (16.1) assumes independent samples, which an implementation must verify rather than assume. The verifier's discrimination is measured against final-answer correctness, and the authors record the false positives that produces. Chapter 26 takes up what coverage can and cannot bound.
Chapter 16 source: "What this does not settle".
Check your understanding: At p = 0.2 with every problem sharing p, what is the coverage of five samples, and what does a blind pick from them deliver?
Chapter 16 source: section "What sampling buys". Demonstration C16-D01.
Demonstration 2 of 4
The selector turns coverage into delivery
How much of the coverage reaches the user under a budget and a deadline, and what changes when the samples share one failure or the limits are tight?
Equation (16.3) multiplies two chances: a correct sample must be present, and the selector must find it. A perfect selector (1.0) delivers the whole coverage; a weaker one delivers a fraction. If the samples share one failure, coverage stays at p however many are drawn, and a bank that holds a correct sample holds only correct ones, so the selector adds nothing. The right panel adds the resources: an allocation that is over the budget or the deadline stays visible but cannot be chosen, even when its coverage is the highest, and the gap between coverage and delivery tells which part, the generator or the selector, is holding the result back.
Scroll sideways for the whole equation
p is the chance one sample is correct. Cov(k) is the chance a correct sample is in the bank. The selector success shown is the chance the selector ranks a correct sample first when the bank holds one; the book calls it pi(k), and here it is held constant for banks of two or more. With one sample there is nothing to rank. Sel(k) is the chance the returned answer is correct. Blind selection returns p. Cost of n samples is n times the sample cost plus the selector cost (no selector cost for one sample); time is n times the sample time plus the selector time. If the samples share one failure, a bank that holds a correct sample holds only correct samples, so the selector is certain to pick a correct one and its setting is not used.
Predict first. With independent samples and a selector that succeeds 0.5 of the time when a correct sample exists, will five samples beat a blind pick of 0.40?
Choose an example
Scroll sideways for the whole figure
Constructed example: the laboratory's sample-allocation function with the notebook's default (p = 0.4, selector 0.9, counts 1, 2, 3 and 5, limits 6 and 6), changed (one shared failure) and transfer cases (p = 0.7, selector 0.5, costs 2 and 3, limits 3 and 6, counts 1, 2 and 4), and the workbook's question about an allocation set with no feasible row (a budget of 0.5, defined for this reader).
Calculated values
- Feasible allocations
- n = 1, n = 2, n = 3, n = 5
- Chosen allocation
- n = 5
- Coverage of the chosen allocation
- 0.9222
- Selected success of the chosen allocation
- 0.8300
- Blind selection
- 0.40
- Coverage minus delivery at n = 5
- 0.0922
Cov(5) = 1 - (1 - 0.40)^5 = 1 - 0.07776 = 0.92224. Sel(5) = 0.92224 x 0.90 + (1 - 0.92224) x 0 = 0.8300; n = 5: cost 5 x 1 + 1 = 6, time 5 x 1 + 1 = 6. Other rows: n = 1 is feasible with Sel = 0.4000 and cost 1; n = 2 is feasible with Sel = 0.5760 and cost 3; n = 3 is feasible with Sel = 0.7056 and cost 4. At the chosen n = 5, coverage is 0.922 and delivery 0.830, a gap of 0.092: the selector still costs part of the coverage, but delivery is more than half of it, so neither factor alone explains the shortfall.
Worked steps
- Coverage: Cov(5) = 1 - (1 - 0.40)^5 = 0.92224.
- Delivery: Sel(5) = 0.92224 x 0.90 = 0.8300.
- n = 5: cost 5 x 1 + 1 = 6, time 5 x 1 + 1 = 6; budget 6, deadline 6 s.
- n = 1: cost 1 x 1 = 1, time 1 x 1 = 1 (no selector): feasible, Sel = 0.4000.
- n = 2: cost 2 x 1 + 1 = 3, time 2 x 1 + 1 = 3: feasible, Sel = 0.5760.
- n = 3: cost 3 x 1 + 1 = 4, time 3 x 1 + 1 = 4: feasible, Sel = 0.7056.
- Rule: highest selected success among the feasible rows; ties go to lower cost.
Use the idea
Report coverage and delivery as separate numbers on the same tasks. If coverage is high and delivery low, work on the selector; if coverage itself is low, no selector can help. Keep infeasible rows in the record, as the companion's probe does, and say when none is feasible instead of exceeding a hard limit.
Where the conclusion applies
A constant selector success is a constructed simplification. The book stresses that a real selector's success must be measured on banks of the size and kind in use, because it depends on ties, on the number of correct answers and on dependence. The shared-failure state is an extreme case chosen to show the boundary; there the selector setting has no effect. Time is serial and the costs are constructed. The rule that a shared failure makes selection equal coverage is the laboratory's and the workbook's convention; the chapter itself states only that dependence breaks Equation (16.1). The no-feasible state uses a budget of 0.5, defined for this reader.
Common wrong turn: A selector that ranks well pairwise delivers well
Check your understanding: With independent samples, p = 0.4 and selector success 0.75, what is Sel(3)?
Chapter 16 source: section "The selector needs its own measured success rate". Demonstration C16-D02.
Demonstration 3 of 4
Length, volume or a selector: spending sixty units
With 60 units of compute, does it matter more to think longer, to sample more, or to pay for a selector?
Without a selector, Equation (16.2) says delivery is just p, so fifteen long samples deliver the long-chain p and sixty short samples deliver 0.20, however many samples were paid for. With a selector, Equation (16.3) multiplies a coverage near one by the selector success, so the selector's quality sets the result. Once coverage is near one, Equation (16.4) says another sample adds almost nothing, which is why, at the chapter's 0.35, the nine units left over by allocation 4 are not best spent on more samples.
Scroll sideways for the whole equation
One unit is one short generation. A long reasoning chain costs 4 units and raises the chance one sample is correct from 0.20 to the chosen value. Checking a sample with the selector costs a quarter of a unit. Cov is coverage, Sel the chance the returned answer is correct, and the selector success pi is the chance it ranks a correct sample first. Bars show Sel; diamonds show Cov.
Predict first. Using the book's selector values, which allocation comes out highest?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's sixty-unit construction (Figure 16.3), including its stipulated selector values 0.73 and 0.785, with the other selector values and the second long-chain p defined for this reader.
Calculated values
- All on length
- 0.350
- All on volume
- 0.200
- Volume + selector
- 0.730
- Length + selector
- 0.781
- Best allocation
- Length + selector
Allocation 1: blind, so Sel = p = 0.35. Allocation 2: Sel = p = 0.20. Allocation 3: Cov(48) = 1 - 0.8^48 = 0.99998, so Sel = 0.99998 x 0.73 = 0.730. Allocation 4: Cov(12) = 1 - 0.65^12 = 0.9943, so Sel = 0.9943 x 0.785 = 0.781. Spending: allocation 4 uses 12 x 4 + 12 x 0.25 = 51 units and leaves 9 unspent; the other three use all 60. A thirteenth long sample would add 0.35 x 0.65^12 = 0.0020 to coverage (Equation 16.4), small next to the 0.9943 already held. Both selector allocations beat both blind ones: the worse selector allocation (0.730) is above the better blind one (0.350).
Worked steps
- Allocation 1: 15 long samples, blind: Sel = p = 0.35.
- Allocation 2: 60 short samples, blind: Sel = p = 0.20.
- Allocation 3: Cov(48) = 1 - 0.8^48 = 0.99998; Sel = 0.99998 x 0.73 = 0.730.
- Allocation 4: Cov(12) = 1 - 0.65^12 = 0.9943; Sel = 0.9943 x 0.785 = 0.781.
- Units: allocation 4 uses 12 x 4 + 12 x 0.25 = 51 of 60, leaving 9.
- A thirteenth long sample adds 0.35 x 0.65^12 = 0.0020 (Equation 16.4).
- Highest delivered: Length + selector.
Use the idea
When splitting a fixed budget, price every lever in one unit and compute delivery for each combination. Do the arithmetic before arguing about whether longer reasoning or more samples is the better lever.
Where the conclusion applies
The costs, the two p values and the selector successes are constructed; the book states them to show the ordering, not as results for any real system. Selector success is held constant across bank sizes here. A selector that ranks correct answers below wrong ones, as in the 0.15 state, can do worse than not selecting.
Common wrong turn: Without a good selector, think less
Check your understanding: Short samples cost 1 unit and long ones 4. With 20 short samples verified at 0.25 each, how many units are used, and what is Sel if pi = 0.6 and p = 0.2?
Chapter 16 source: section "Three levers, one budget". Demonstration C16-D03.
Demonstration 4 of 4
What a completed task costs, and who is allowed to be slow
Which procedure is cheapest per authorized success once a completion floor and a deadline rule some out, and what does the ratio do when nothing succeeds?
Equation (16.5) divides everything spent by the number of real successes, so cheap procedures that fail often can still look cheap per success while missing the owner's requirements. A deadline and a completion floor are separate constraints and remove procedures before any ratio is compared. If none remains, the ratio does not choose for the owner. The ratio also needs a denominator: with two observed attempts and no success, Equation (16.5) has no value, and a printable replacement would change the metric.
Scroll sideways for the whole equation
N is the number of attempted tasks (100 here). The cost of run i, written c with subscript i in the equation, counts failed attempts too. The indicator in the denominator is 1 when run i ended in an authorized, confirmed completion and 0 otherwise, so the denominator counts those runs. Cost per success is total cost divided by that count. Completion rate is the share of attempts that succeed; response time is seconds to answer. A failed task is also given a loss of 20 in one metric. The Pareto frontier keeps the procedures that no other procedure beats on completion, cost and time at once.
Predict first. With a 3 second deadline and a completion floor of 0.75, which procedure is the only feasible one?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's three procedures (costs 1, 1.5, 2; completion 60, 80, 85 percent; response 1, 2, 5 seconds; failure loss 20) and workbench problem E.4 (two attempts costing 2 and 3, neither successful), with the deadline and completion floor varied. The costs and completion rates come from the section named above; the response times, deadlines and the loss of 20 per failure come from the section "Waiting changes the allocation".
Calculated values
- Cost per success A
- 1.67
- Cost per success B
- 1.88
- Cost per success C
- 2.35
- Feasible procedures
- B, C
- Cheapest per success overall
- A
- Cheapest per success among feasible
- B
- Lowest cost plus failure loss (20 per failure)
- C
- Procedures on the Pareto frontier
- A, B, C
Over 100 attempts, Equation (16.5) gives A: 100 x 1 / (100 x 0.60) = 100 / 60 = 1.6667; B: 150 / 80 = 1.8750; C: 200 / 85 = 2.3529. With a 10 second deadline and at least 0.75 completion, the feasible set is B, C; the lowest cost per success among them is B (the lowest of all three is A, which is not feasible). Adding a loss of 20 per failure, cost plus failure loss = cost + (1 - rate) x 20 gives A: 1 + 0.40 x 20 = 9.0, B: 1.5 + 0.20 x 20 = 5.5, C: 2 + 0.15 x 20 = 5.0, so C is lowest among the feasible. No procedure is at least as good as another on completion, cost and time at once (A is cheapest and fastest, C completes most), so all three stay on the Pareto frontier; the frontier does not pick one for this deadline.
Worked steps
- A: total cost 100 x 1 = 100, successes 100 x 0.60 = 60, so 100 / 60 = 1.6667.
- B: total cost 100 x 1.5 = 150, successes 100 x 0.80 = 80, so 150 / 80 = 1.8750.
- C: total cost 100 x 2 = 200, successes 100 x 0.85 = 85, so 200 / 85 = 2.3529.
- Feasible needs completion at least 0.75 and time at most 10 s: B, C.
- Cheapest per success among the feasible: B (1.8750).
- With a loss of 20 per failure: A 9.0, B 5.5, C 5.0; lowest among the feasible is C.
Use the idea
When reporting an efficiency ratio, count every attempt's cost, count only runs that met the declared completion contract, and state the constraints the ratio does not capture. Report a zero denominator as undefined, with the attempted-task count and the total cost.
Where the conclusion applies
Three constructed procedures with known costs, completion rates and response times. Real rates must be estimated on matching tasks, and a ratio from a small sample is a different object from the underlying rate. The two observed attempts describe a sample, not the population ratios. The Pareto statement is about the three procedures as stated; it does not pick one for a particular deadline or authority contract.
Common wrong turn: The lowest cost per success picks the procedure
Check your understanding: A procedure costs 3 per attempt and completes 0.5 of 100 attempts. What is its cost per success?
Chapter 16 source: section "What the budget actually buys". Demonstration C16-D04.