Demonstration 1 of 4
A bank is a finite object, and its pool depends on how it was sampled
For fixed pools of ten programs per task, how often does a random subset contain a pass, how does that differ from an independent-draw guess, and how does the sampling rule change the answer?
Count the subsets with no pass, divide by all subsets, and subtract from one. This is exact for the fixed pool, because subsets are drawn without replacement from programs that already exist. The independent-draw formula treats every pick as a fresh draw with the pool's pass rate, which is a different statistical question. A selector with no correctness signal succeeds only at the pass share c / n. The chapter's temperature point shows in the two banks: the concentrated bank is better at k = 1, the dispersed bank overtakes it as k grows, so the ceiling is not a property of the weights alone.
Scroll sideways for the whole equation
n is the number of programs generated for a task (10 here) and c the number that pass its tests. k is the size of the subset considered. C(n, k) counts the ways to choose k of n programs. The fraction C(n - c, k) / C(n, k) is the share of subsets with no pass, so one minus it is the chance a subset holds at least one pass. Equation (26.1) averages this over the tasks. The independent-draw formula is 1 - (1 - c / n)^k, which treats every pick as a fresh draw. The low and high temperature banks are constructed: low concentrates passes on a few tasks, high spreads them.
Predict first. With 2 passing programs among 10 and subsets of size 3, is the exact coverage above or below the independent-draw value 0.488?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's ten-candidate worked subset calculation (2 passing, size 3) and two constructed four-task banks that differ in how concentrated the passes are.
Calculated values
- Tasks in the bank
- 1
- Passing programs per task (of 10)
- 2
- Exact coverage at k = 3
- 0.5333
- Independent-draw formula
- 0.4880
- No-signal selector success
- 0.200
- Coverage not collected by that selector
- 0.3333
Coverage = 1 - C(8,3) / C(10,3) = 1 - 56 / 120 = 0.5333, so 64 of the 120 subsets hold at least one pass. A selector that picks one program of the subset with no correctness signal succeeds with probability 2 / 10 = 0.20, which leaves 0.5333 - 0.20 = 0.3333 uncollected. Putting the single-program rate 0.20 into an independent-draw formula gives 1 - (1 - 0.20)^3 = 0.4880, a different quantity that does not equal the exact fixed-pool value.
Worked steps
- One task, n = 10 programs, c = 2 pass, subsets of size k = 3: C(10,3) = 120 subsets.
- Subsets by number of passes: 0 passes: 56, 1 passes: 56, 2 passes: 8.
- Subsets with no pass: C(8,3) = 56.
- Coverage = 1 - 56 / 120 = 0.5333.
- No-signal selector: c / n = 2 / 10 = 0.20; uncollected 0.5333 - 0.20 = 0.3333.
- Independent-draw formula: 1 - (1 - 0.20)^3 = 0.4880.
Use the idea
When a report gives a pass@k number, ask for the pool size n, the per-task pass counts, the subset size k and the sampling settings. Without them you cannot tell an exact pool calculation from an extrapolation, or compare two banks made at different temperatures.
Where the conclusion applies
Pools of ten programs, labels fixed by the declared tests, subsets chosen uniformly without replacement. The two four-task banks are constructed to show the direction the chapter reports for temperature; they are not measurements. The result says nothing about programs outside the pool, other prompts or adaptive retries. If the ten programs repeat one wrong approach the labels are strongly correlated, and the pool still has exactly this coverage but little useful variety. The chapter's pools of 200 programs per task are too large for a pencil check and are not shown.
Common wrong turn: Pass@k is pass@1 pushed through independent draws
Check your understanding: If only 1 of the 10 programs passes and k = 4, what is the exact coverage?
Chapter 26 source: section "Stage two: a bank is a finite object". Demonstration C26-D01.
Demonstration 2 of 4
Coverage is a ceiling, and each rung can only lose success
Can any selector return a correct candidate more often than the bank contains one, and where does each task's success get lost on the way to deployment?
For each task the ceiling is 1 if any candidate passes and 0 otherwise, and a chosen candidate can pass only if one exists. So task by task the selector cannot exceed the ceiling, and the weighted sums keep that order. Deployment can only remove tasks again. A failed task therefore sits at one rung: absent from the bank (generation), present but missed (selection), or selected and passing but denied (deployment). Each calls for a different next step.
Scroll sideways for the whole equation
Cov is oracle coverage: the weighted share of tasks for which at least one candidate in the bank passes. Sel is the weighted share of tasks for which the selector's chosen candidate passes. Deployed is the weighted share of tasks whose selected candidate passes and which are allowed to deploy. The weights are the task distribution. An oracle selector is shown the pass labels; the notebook selector is not.
Predict first. In the notebook's default case (three specialists, selector always takes candidate 1, every task deployable), what are Cov, Sel and deployed?
Choose an example
Scroll sideways for the whole figure
Constructed example: the laboratory's default, changed and transfer banks, computed with the laboratory's own bank-ceiling function, plus a two-candidate prefix of the default bank.
Calculated values
- Bank
- 3 candidates, three specialists
- Oracle coverage Cov
- 1.00
- Actual selection Sel
- 0.20
- Deployed success
- 0.20
- Lost to selection (Cov minus Sel)
- 0.80
- Lost to deployment (Sel minus deployed)
- 0.00
- Coverage as candidates are added
- 0.20 / 0.50 / 1.00
- Where each task stands
- task 1 delivered, task 2 missed by selector, task 3 missed by selector
Cov = 0.2 x 1 + 0.3 x 1 + 0.5 x 1 = 1.00. The selector takes the candidate the notebook case chooses, 1, 1, 1 by task, so Sel = 0.2 x 1 + 0.3 x 0 + 0.5 x 0 = 0.20. Deployed = 0.2 x 1 + 0.3 x 0 + 0.5 x 0 = 0.20 because every task is allowed to deploy. Cov - Sel = 1.00 - 0.20 = 0.80 is lost to selection and Sel - deployed = 0.20 - 0.20 = 0.00 to deployment.
Worked steps
- Weights 0.2, 0.3, 0.5; a task is covered if any candidate in the bank passes it.
- Cov = 0.2 x 1 + 0.3 x 1 + 0.5 x 1 = 1.00.
- Selector choices by task: 1, 1, 1; Sel = 0.2 x 1 + 0.3 x 0 + 0.5 x 0 = 0.20.
- Deployed = 0.2 x 1 + 0.3 x 0 + 0.5 x 0 = 0.20.
- Ladder: 1.00 then 0.20 then 0.20; no rung is higher than the one before it.
- Where each task stands: 1 delivered, 2 missed, 3 missed.
Use the idea
Before buying a better ranker, compute coverage on the bank you already have. If coverage is low, the generator, its context or its tools are the limit. Keep permission filtering out of correctness labels so a denied correct answer is not counted as a model error.
Where the conclusion applies
Small constructed banks with fixed task weights and pass labels from one evaluator. The oracle selector uses information a deployed system may not have. The ceiling holds only for this bank and evaluator; a new candidate or a different evaluator changes the object being bounded. Adding a candidate cannot lower coverage, because a covered task stays covered.
Common wrong turn: High coverage proves the deployed selector can recover the answer
Check your understanding: If only the heavy task (weight 0.5) were solved by anything in the bank, what is the largest Sel any selector could reach?
Chapter 26 source: section "Stage three: oracle coverage is an information ceiling". Demonstration C26-D02.
Demonstration 3 of 4
What actual selection leaves behind
Where does a coverage and selection pair sit against the ceiling, how large is the gap, and what does the gap point at?
The chapter reports one experiment with 100 candidates per problem: 37.7 percent for a single sample, 44.5 percent for choosing by mean log-probability, and 77.5 percent when unit tests choose. Subtract the practical rule from the unit-test result to get the gap. No selector can sit above the line Sel = Cov. Its distance below the line is what needs explaining: a large gap points at selection or evidence, a gap near zero or a low ceiling points back at generation, context, tools or budget.
Scroll sideways for the whole equation
Cov(k) is the success when the oracle (the unit tests) chooses among k candidates and Sel(k) the success of the practical rule, both in percent of problems solved. The gap Cov minus Sel is measured in percentage points and cannot be negative. The share collected is a hypothetical fraction of that gap that a new selector might win back. Codex-S and Codex-12B figures are the chapter's reported percentages, used as given.
Predict first. Starting from Codex-S with mean log-probability selection (44.5 against 77.5), what success would collecting half the gap give?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's reported Codex percentages used as given, two constructed coverage and selection pairs, and a hypothetical collected share defined for this reader.
Calculated values
- Point
- Codex-S, mean log-probability
- Cov (ceiling, percent)
- 77.5
- Sel (delivered, percent)
- 44.5
- Gap Cov minus Sel (points)
- 33.0
- Gap as a share of the ceiling
- 0.43
- Hypothetical share collected
- 0.50
- New selected success
- 61.0
- Gap that remains
- 16.5
Gap = Cov - Sel = 77.5 - 44.5 = 33.0 points, which is not negative, as Equation (26.3) says. If a new selector collected a hypothetical share 0.50 of it, success would be 44.5 + 0.50 x 33.0 = 61.0, leaving 77.5 - 61.0 = 16.5 points. The chapter reports 37.7 for a single Codex-S sample, so mean log-probability gained 44.5 - 37.7 = 6.8 over one sample and left 33.0 points of the test-selected result uncollected. A wide gap points at validation, ranking or the evidence interface. The share is a planning what-if: it becomes a result only when a selector is built and measured on the same bank, tasks and evaluator.
Worked steps
- Ceiling Cov = 77.5 and delivered Sel = 44.5, both in percent.
- Gap = 77.5 - 44.5 = 33.0 points, and Sel stays on or below the line Sel = Cov.
- Hypothetical gain = 0.50 x 33.0 = 16.5.
- New Sel = 44.5 + 16.5 = 61.0.
- Gap that remains = 77.5 - 61.0 = 16.5.
Use the idea
Report delivered selection and the same-bank unit-test result side by side. Ask what part of the gap a practical intervention could collect under the deployment contract, then compare that gain with cost, latency, error, safety risk and permissions.
Where the conclusion applies
The Codex percentages are the chapter's reported figures for one experiment, used as given and not re-measured, and they hold only for that generator, bank and tests. The collected share is a planning what-if with no data behind it. The two constructed pairs exist only to show the other two readings. Collecting any of the gap needs evidence available before the outcome is known, which a deployment may not have.
Common wrong turn: The oracle number is a forecast of the product score
What this does not settle
A finite candidate bank can set a ceiling for selectors with declared information, and the gap to actual selection identifies an engineering question rather than an achieved capability. The chapter does not show that an oracle-conditioned score transfers to deployment.
Chapter 26 source: "What this does not settle".
Check your understanding: A practical selector reaches 50.0 percent against 77.5 for unit tests. A new verifier is hoped to collect 40 percent of the gap. What is the new success?
Chapter 26 source: section "Stage four: what actual selection leaves behind". Demonstration C26-D03.
Demonstration 4 of 4
A saturation forecast depends on the curve
If three observed coverage points are fitted by three different curves, do they agree about the limit, and about whether a larger bank or a better selector has more room?
Three points fix three numbers, so every family fits them with no error. Beyond the observed range the families separate, and their limits come out near 0.65, 0.74 and 0.86. The data cannot say which family is right, so the limit reflects the choice of curve. The same holds for the stopping question: how much a larger bank adds depends on the family, and the room below the ceiling depends on how far the selector is from it.
Scroll sideways for the whole equation
k is the number of candidates per task and the curve value is coverage. A is the limit the curve approaches. The three formulas, in order, are the exponential, hyperbolic and power-law families, and e is Euler's number (about 2.718). A, B and the shape number (tau, h or a) are all fitted so that each curve passes exactly through the three observed points (10, 0.40), (25, 0.55) and (50, 0.63). Sel at k = 50 is a hypothetical selected success defined for this reader, 0.60 or 0.30.
Predict first. All three curves hit the same three points. Which family projects the highest limit?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's constructed coverage points (0.40, 0.55, 0.63 at 10, 25, 50 samples) fitted exactly by three curve families, with a hypothetical selected success defined for this reader.
Calculated values
- Family
- Exponential approach
- Shape parameter (tau)
- 16.7
- Projected limit A
- 0.653
- Curve at k = 100
- 0.652
- Limits of all three families
- 0.65 / 0.74 / 0.86
- Spread of the limits (points)
- 21
- Room from a larger bank (50 to 100)
- 0.022
- Room below the ceiling at k = 50
- 0.330
A = 0.63 + B x e^(-50/16.7) = 0.63 + 0.461 x 0.0501 = 0.653. The three families all pass through 0.40, 0.55 and 0.63, yet their limits are 0.65, 0.74 and 0.86, a spread of 0.86 - 0.65 = 0.21. Stopping question: this curve reads 0.652 at k = 100, so doubling the bank adds 0.652 - 0.63 = 0.022. With a hypothetical selected success of 0.30 at k = 50, room below the ceiling is 0.63 - 0.30 = 0.330; all three families agree that the selector has the larger room at this selector value. The observed range cannot choose among the families, so the limit and the growth are forecasts that depend on the family, not observed results.
Worked steps
- Observed coverage: 0.40 at k = 10, 0.55 at k = 25, 0.63 at k = 50.
- Exponential approach fitted exactly through them: shape tau = 16.7.
- A = 0.63 + B x e^(-50/16.7) = 0.63 + 0.461 x 0.0501 = 0.653.
- Curve at k = 100 is 0.652, so a larger bank adds 0.652 - 0.63 = 0.022.
- Room below the ceiling at k = 50: 0.63 - 0.30 = 0.330 (Sel is a hypothetical value).
- Compare the two rooms, then price latency and work before choosing; the comparison does not decide by itself.
Use the idea
Use a saturation curve to plan capacity inside a stated budget, and report the fitted family with it. When generation adds little coverage while selection stays far below coverage, the next increment of budget has two destinations: more candidates or a better way to judge the ones already present. Do not read the limit as a bound on what a changed prompt, tool or generator could reach.
Where the conclusion applies
The chapter's constructed fixed-generator saturation curve: three made-up coverage points on one declared task bank. With more points or a physical reason for one family, the choice could narrow, but this example has neither. A changed generator, selector or task distribution gives a different curve, so no limit here is a bound on an open design space. The selected success values are hypothetical, and the comparison of rooms does not price latency, cost or work.
Common wrong turn: A curve that fits every point has found the limit
What this does not settle
This chapter does not identify a model's maximum capability, predict gains from unlimited compute or tools, or prove that an oracle-conditioned score transfers to deployment.
Chapter 26 source: "What this does not settle".
Check your understanding: Two fits agree on all three points but project limits of 0.65 and 0.86. By how much do they differ?
Chapter 26 source: section "A saturation forecast". Demonstration C26-D04.