The Mathematics of AI Agents, laboratory reader ยท Chapter 26

How Much More Could the System Become?

A finite bank of candidates sets a conditional ceiling, and the gap to what a selector actually returns is an engineering question, not a forecast.

These four demonstrations follow the chapter's staged argument. A bank of candidates is a finite object with an exact coverage calculation that moves with the sampling rule. Coverage is a ceiling that no selector can pass, and success is lost at each rung from the bank to deployment. The gap below the ceiling points at an engineering question, and a curve fitted to a few points projects a limit that depends on the curve you pick.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

A bank is a finite object, and its pool depends on how it was sampled

For fixed pools of ten programs per task, how often does a random subset contain a pass, how does that differ from an independent-draw guess, and how does the sampling rule change the answer?

Count the subsets with no pass, divide by all subsets, and subtract from one. This is exact for the fixed pool, because subsets are drawn without replacement from programs that already exist. The independent-draw formula treats every pick as a fresh draw with the pool's pass rate, which is a different statistical question. A selector with no correctness signal succeeds only at the pass share c / n. The chapter's temperature point shows in the two banks: the concentrated bank is better at k = 1, the dispersed bank overtakes it as k grows, so the ceiling is not a property of the weights alone.

Equation (26.1), written in LaTeX: \widehat{\operatorname{Cov}}_{\mathcal K}(k)= \frac{1}{|\mathcal K|}\sum_{i\in\mathcal K}(1 - \frac{\binom{n_i-c_i}{k}}{\binom{n_i}{k}}), 1\le k\le \min_{i\in\mathcal K}n_i

Scroll sideways for the whole equation

n is the number of programs generated for a task (10 here) and c the number that pass its tests. k is the size of the subset considered. C(n, k) counts the ways to choose k of n programs. The fraction C(n - c, k) / C(n, k) is the share of subsets with no pass, so one minus it is the chance a subset holds at least one pass. Equation (26.1) averages this over the tasks. The independent-draw formula is 1 - (1 - c / n)^k, which treats every pick as a fresh draw. The low and high temperature banks are constructed: low concentrates passes on a few tasks, high spreads them.

Predict first. With 2 passing programs among 10 and subsets of size 3, is the exact coverage above or below the independent-draw value 0.488?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A bank is a finite object, and its pool depends on how it was sampled. Left, coverage at k = 3 for each of 1 task as bars with independent-draw diamonds. Right, three bars for the bank average: exact 0.533, independent-draw 0.488 and a no-signal selector 0.200.
Bank of ten-program pools: Chapter's worked pool (one task, 2 passing), Subset size k: 3
Constructed example: the book's ten-candidate worked subset calculation (2 passing, size 3) and two constructed four-task banks that differ in how concentrated the passes are.

Calculated values

Tasks in the bank
1
Passing programs per task (of 10)
2
Exact coverage at k = 3
0.5333
Independent-draw formula
0.4880
No-signal selector success
0.200
Coverage not collected by that selector
0.3333

Coverage = 1 - C(8,3) / C(10,3) = 1 - 56 / 120 = 0.5333, so 64 of the 120 subsets hold at least one pass. A selector that picks one program of the subset with no correctness signal succeeds with probability 2 / 10 = 0.20, which leaves 0.5333 - 0.20 = 0.3333 uncollected. Putting the single-program rate 0.20 into an independent-draw formula gives 1 - (1 - 0.20)^3 = 0.4880, a different quantity that does not equal the exact fixed-pool value.

Worked steps

  1. One task, n = 10 programs, c = 2 pass, subsets of size k = 3: C(10,3) = 120 subsets.
  2. Subsets by number of passes: 0 passes: 56, 1 passes: 56, 2 passes: 8.
  3. Subsets with no pass: C(8,3) = 56.
  4. Coverage = 1 - 56 / 120 = 0.5333.
  5. No-signal selector: c / n = 2 / 10 = 0.20; uncollected 0.5333 - 0.20 = 0.3333.
  6. Independent-draw formula: 1 - (1 - 0.20)^3 = 0.4880.

Use the idea

When a report gives a pass@k number, ask for the pool size n, the per-task pass counts, the subset size k and the sampling settings. Without them you cannot tell an exact pool calculation from an extrapolation, or compare two banks made at different temperatures.

Where the conclusion applies

Pools of ten programs, labels fixed by the declared tests, subsets chosen uniformly without replacement. The two four-task banks are constructed to show the direction the chapter reports for temperature; they are not measurements. The result says nothing about programs outside the pool, other prompts or adaptive retries. If the ten programs repeat one wrong approach the labels are strongly correlated, and the pool still has exactly this coverage but little useful variety. The chapter's pools of 200 programs per task are too large for a pencil check and are not shown.

Common wrong turn: Pass@k is pass@1 pushed through independent draws
The chapter says the calculation is not one direct sequence of k trials and not an independent-draw extrapolation from pass@1. Estimating pass@k as one minus one minus the empirical pass@1, raised to k, is biased; the finite-pool combinatorial form is the exact subset probability.
Check your understanding: If only 1 of the 10 programs passes and k = 4, what is the exact coverage?
1 - C(9,4) / C(10,4) = 1 - 126 / 210 = 0.40, which equals k / n = 4 / 10.

Chapter 26 source: section "Stage two: a bank is a finite object". Demonstration C26-D01.

Demonstration 2 of 4

Coverage is a ceiling, and each rung can only lose success

Can any selector return a correct candidate more often than the bank contains one, and where does each task's success get lost on the way to deployment?

For each task the ceiling is 1 if any candidate passes and 0 otherwise, and a chosen candidate can pass only if one exists. So task by task the selector cannot exceed the ceiling, and the weighted sums keep that order. Deployment can only remove tasks again. A failed task therefore sits at one rung: absent from the bank (generation), present but missed (selection), or selected and passing but denied (deployment). Each calls for a different next step.

Equation (26.2), written in LaTeX: \operatorname{Sel}(k) \le \operatorname{Cov}(k)

Scroll sideways for the whole equation

Cov is oracle coverage: the weighted share of tasks for which at least one candidate in the bank passes. Sel is the weighted share of tasks for which the selector's chosen candidate passes. Deployed is the weighted share of tasks whose selected candidate passes and which are allowed to deploy. The weights are the task distribution. An oracle selector is shown the pass labels; the notebook selector is not.

Predict first. In the notebook's default case (three specialists, selector always takes candidate 1, every task deployable), what are Cov, Sel and deployed?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Coverage is a ceiling, and each rung can only lose success. Left, a grid of candidates by tasks marking passes and the chosen candidate, with each task labeled delivered, missed, missed. Right, a descending ladder of three bars: Cov 1.00, Sel 0.20, deployed 0.20.
Bank: Three specialists (notebook default), Selector: The notebook case's selector, Deployment permission: Every task allowed
Constructed example: the laboratory's default, changed and transfer banks, computed with the laboratory's own bank-ceiling function, plus a two-candidate prefix of the default bank.

Calculated values

Bank
3 candidates, three specialists
Oracle coverage Cov
1.00
Actual selection Sel
0.20
Deployed success
0.20
Lost to selection (Cov minus Sel)
0.80
Lost to deployment (Sel minus deployed)
0.00
Coverage as candidates are added
0.20 / 0.50 / 1.00
Where each task stands
task 1 delivered, task 2 missed by selector, task 3 missed by selector

Cov = 0.2 x 1 + 0.3 x 1 + 0.5 x 1 = 1.00. The selector takes the candidate the notebook case chooses, 1, 1, 1 by task, so Sel = 0.2 x 1 + 0.3 x 0 + 0.5 x 0 = 0.20. Deployed = 0.2 x 1 + 0.3 x 0 + 0.5 x 0 = 0.20 because every task is allowed to deploy. Cov - Sel = 1.00 - 0.20 = 0.80 is lost to selection and Sel - deployed = 0.20 - 0.20 = 0.00 to deployment.

Worked steps

  1. Weights 0.2, 0.3, 0.5; a task is covered if any candidate in the bank passes it.
  2. Cov = 0.2 x 1 + 0.3 x 1 + 0.5 x 1 = 1.00.
  3. Selector choices by task: 1, 1, 1; Sel = 0.2 x 1 + 0.3 x 0 + 0.5 x 0 = 0.20.
  4. Deployed = 0.2 x 1 + 0.3 x 0 + 0.5 x 0 = 0.20.
  5. Ladder: 1.00 then 0.20 then 0.20; no rung is higher than the one before it.
  6. Where each task stands: 1 delivered, 2 missed, 3 missed.

Use the idea

Before buying a better ranker, compute coverage on the bank you already have. If coverage is low, the generator, its context or its tools are the limit. Keep permission filtering out of correctness labels so a denied correct answer is not counted as a model error.

Where the conclusion applies

Small constructed banks with fixed task weights and pass labels from one evaluator. The oracle selector uses information a deployed system may not have. The ceiling holds only for this bank and evaluator; a new candidate or a different evaluator changes the object being bounded. Adding a candidate cannot lower coverage, because a covered task stays covered.

Common wrong turn: High coverage proves the deployed selector can recover the answer
The chapter says no high coverage number proves that a deployed selector can recover the available answer, and no stronger reranker can fully repair a generator whose bounded banks rarely contain a correct answer. Both mistakes are blocked by the inequality Sel <= Cov.
Check your understanding: If only the heavy task (weight 0.5) were solved by anything in the bank, what is the largest Sel any selector could reach?
Cov = 0.2 x 0 + 0.3 x 0 + 0.5 x 1 = 0.5, so Sel is at most 0.5 for every selector.

Chapter 26 source: section "Stage three: oracle coverage is an information ceiling". Demonstration C26-D02.

Demonstration 3 of 4

What actual selection leaves behind

Where does a coverage and selection pair sit against the ceiling, how large is the gap, and what does the gap point at?

The chapter reports one experiment with 100 candidates per problem: 37.7 percent for a single sample, 44.5 percent for choosing by mean log-probability, and 77.5 percent when unit tests choose. Subtract the practical rule from the unit-test result to get the gap. No selector can sit above the line Sel = Cov. Its distance below the line is what needs explaining: a large gap points at selection or evidence, a gap near zero or a low ceiling points back at generation, context, tools or budget.

Equation (26.3), written in LaTeX: \operatorname{Cov}(k)-\operatorname{Sel}(k) \ge 0

Scroll sideways for the whole equation

Cov(k) is the success when the oracle (the unit tests) chooses among k candidates and Sel(k) the success of the practical rule, both in percent of problems solved. The gap Cov minus Sel is measured in percentage points and cannot be negative. The share collected is a hypothetical fraction of that gap that a new selector might win back. Codex-S and Codex-12B figures are the chapter's reported percentages, used as given.

Predict first. Starting from Codex-S with mean log-probability selection (44.5 against 77.5), what success would collecting half the gap give?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: What actual selection leaves behind. A square plot of selected success against oracle coverage with the line Sel = Cov and the region above it hatched. The point sits at coverage 77.5 and selection 44.5; a vertical segment shows the gap and a hollow marker at 61.0 shows the hypothetical share collected.
Coverage and selection pair: Codex-S, mean log-probability (44.5 of 77.5), Hypothetical share of the gap collected: 0.5
Constructed example: the chapter's reported Codex percentages used as given, two constructed coverage and selection pairs, and a hypothetical collected share defined for this reader.

Calculated values

Point
Codex-S, mean log-probability
Cov (ceiling, percent)
77.5
Sel (delivered, percent)
44.5
Gap Cov minus Sel (points)
33.0
Gap as a share of the ceiling
0.43
Hypothetical share collected
0.50
New selected success
61.0
Gap that remains
16.5

Gap = Cov - Sel = 77.5 - 44.5 = 33.0 points, which is not negative, as Equation (26.3) says. If a new selector collected a hypothetical share 0.50 of it, success would be 44.5 + 0.50 x 33.0 = 61.0, leaving 77.5 - 61.0 = 16.5 points. The chapter reports 37.7 for a single Codex-S sample, so mean log-probability gained 44.5 - 37.7 = 6.8 over one sample and left 33.0 points of the test-selected result uncollected. A wide gap points at validation, ranking or the evidence interface. The share is a planning what-if: it becomes a result only when a selector is built and measured on the same bank, tasks and evaluator.

Worked steps

  1. Ceiling Cov = 77.5 and delivered Sel = 44.5, both in percent.
  2. Gap = 77.5 - 44.5 = 33.0 points, and Sel stays on or below the line Sel = Cov.
  3. Hypothetical gain = 0.50 x 33.0 = 16.5.
  4. New Sel = 44.5 + 16.5 = 61.0.
  5. Gap that remains = 77.5 - 61.0 = 16.5.

Use the idea

Report delivered selection and the same-bank unit-test result side by side. Ask what part of the gap a practical intervention could collect under the deployment contract, then compare that gain with cost, latency, error, safety risk and permissions.

Where the conclusion applies

The Codex percentages are the chapter's reported figures for one experiment, used as given and not re-measured, and they hold only for that generator, bank and tests. The collected share is a planning what-if with no data behind it. The two constructed pairs exist only to show the other two readings. Collecting any of the gap needs evidence available before the outcome is known, which a deployment may not have.

Common wrong turn: The oracle number is a forecast of the product score
The chapter warns against calling the best offline score a roadmap forecast, and against describing the assistant as a 77.5-percent system before one exists under the deployment contract. The condition is the result.
What this does not settle

A finite candidate bank can set a ceiling for selectors with declared information, and the gap to actual selection identifies an engineering question rather than an achieved capability. The chapter does not show that an oracle-conditioned score transfers to deployment.

Chapter 26 source: "What this does not settle".

Check your understanding: A practical selector reaches 50.0 percent against 77.5 for unit tests. A new verifier is hoped to collect 40 percent of the gap. What is the new success?
Gap = 77.5 - 50.0 = 27.5. New success = 50.0 + 0.4 x 27.5 = 61.0, a hypothesis until measured.

Chapter 26 source: section "Stage four: what actual selection leaves behind". Demonstration C26-D03.

Demonstration 4 of 4

A saturation forecast depends on the curve

If three observed coverage points are fitted by three different curves, do they agree about the limit, and about whether a larger bank or a better selector has more room?

Three points fix three numbers, so every family fits them with no error. Beyond the observed range the families separate, and their limits come out near 0.65, 0.74 and 0.86. The data cannot say which family is right, so the limit reflects the choice of curve. The same holds for the stopping question: how much a larger bank adds depends on the family, and the room below the ceiling depends on how far the selector is from it.

Equation, written in LaTeX: A - B e^{-k/\tau}

Equation, written in LaTeX: A - B/(k+h)

Equation, written in LaTeX: A - B k^{-a}

Scroll sideways for the whole equation

k is the number of candidates per task and the curve value is coverage. A is the limit the curve approaches. The three formulas, in order, are the exponential, hyperbolic and power-law families, and e is Euler's number (about 2.718). A, B and the shape number (tau, h or a) are all fitted so that each curve passes exactly through the three observed points (10, 0.40), (25, 0.55) and (50, 0.63). Sel at k = 50 is a hypothetical selected success defined for this reader, 0.60 or 0.30.

Predict first. All three curves hit the same three points. Which family projects the highest limit?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A saturation forecast depends on the curve. Left, three curves fitted through the same three points on a log axis, with the exponential approach highlighted and limits near 0.65, 0.74 and 0.86. Right, two bars: room from doubling the bank to 100, 0.022, and room below the ceiling for a selector at 0.30, 0.330.
Highlighted curve family: Exponential approach, Hypothetical selected success at k = 50: 0.30 (selector far below it)
Constructed example: the chapter's constructed coverage points (0.40, 0.55, 0.63 at 10, 25, 50 samples) fitted exactly by three curve families, with a hypothetical selected success defined for this reader.

Calculated values

Family
Exponential approach
Shape parameter (tau)
16.7
Projected limit A
0.653
Curve at k = 100
0.652
Limits of all three families
0.65 / 0.74 / 0.86
Spread of the limits (points)
21
Room from a larger bank (50 to 100)
0.022
Room below the ceiling at k = 50
0.330

A = 0.63 + B x e^(-50/16.7) = 0.63 + 0.461 x 0.0501 = 0.653. The three families all pass through 0.40, 0.55 and 0.63, yet their limits are 0.65, 0.74 and 0.86, a spread of 0.86 - 0.65 = 0.21. Stopping question: this curve reads 0.652 at k = 100, so doubling the bank adds 0.652 - 0.63 = 0.022. With a hypothetical selected success of 0.30 at k = 50, room below the ceiling is 0.63 - 0.30 = 0.330; all three families agree that the selector has the larger room at this selector value. The observed range cannot choose among the families, so the limit and the growth are forecasts that depend on the family, not observed results.

Worked steps

  1. Observed coverage: 0.40 at k = 10, 0.55 at k = 25, 0.63 at k = 50.
  2. Exponential approach fitted exactly through them: shape tau = 16.7.
  3. A = 0.63 + B x e^(-50/16.7) = 0.63 + 0.461 x 0.0501 = 0.653.
  4. Curve at k = 100 is 0.652, so a larger bank adds 0.652 - 0.63 = 0.022.
  5. Room below the ceiling at k = 50: 0.63 - 0.30 = 0.330 (Sel is a hypothetical value).
  6. Compare the two rooms, then price latency and work before choosing; the comparison does not decide by itself.

Use the idea

Use a saturation curve to plan capacity inside a stated budget, and report the fitted family with it. When generation adds little coverage while selection stays far below coverage, the next increment of budget has two destinations: more candidates or a better way to judge the ones already present. Do not read the limit as a bound on what a changed prompt, tool or generator could reach.

Where the conclusion applies

The chapter's constructed fixed-generator saturation curve: three made-up coverage points on one declared task bank. With more points or a physical reason for one family, the choice could narrow, but this example has neither. A changed generator, selector or task distribution gives a different curve, so no limit here is a bound on an open design space. The selected success values are hypothetical, and the comparison of rooms does not price latency, cost or work.

Common wrong turn: A curve that fits every point has found the limit
The chapter says every curve fits the observed points, the projections spread over more than twenty points, and three points cannot choose among families. The asymptote is a forecast that depends on the fitted family, not an observed result.
What this does not settle

This chapter does not identify a model's maximum capability, predict gains from unlimited compute or tools, or prove that an oracle-conditioned score transfers to deployment.

Chapter 26 source: "What this does not settle".

Check your understanding: Two fits agree on all three points but project limits of 0.65 and 0.86. By how much do they differ?
0.86 - 0.65 = 0.21, a spread of 21 points that the three observed points cannot resolve.

Chapter 26 source: section "A saturation forecast". Demonstration C26-D04.