The Mathematics of AI Agents, laboratory reader ยท Chapter 24

How Much Capability Have We Extracted?

A score describes a system, a task sample and a measurement procedure together, and it needs its boundary to be read.

These four demonstrations follow the chapter's rule that a number travels with its contract. They show how one score depends on its denominator, its sampling and its target mixture, how a reported onset depends on a threshold, what a best tested configuration can claim after paying for the search, and why a mean success rate does not fix repeated-run reliability. Every number is a constructed teaching value.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

Observations, denominators and estimands

What do these observations support, and which quantities change when the sample size, the pairing or the target mixture changes while the records stay the same?

Every score is an average over a stated denominator, so the first thing to read is what was counted. A gap between two independent banks has noise that adds, and that noise falls with the number of tasks. When the records are paired, only the pairs both procedures attempted enter the difference, and when records do not overlap the matched difference does not exist. A target mixture changes the estimand without touching a single observation, and a mixture with exposed tasks can raise the aggregate without any gain on unfamiliar tasks.

Equation (24.1), written in LaTeX: \widehat U_{\mathcal V}(S)= \frac{1}{N_{\mathrm{eval}}}\sum_{j=1}^{N_{\mathrm{eval}}} u_j.

Equation (24.2), written in LaTeX: \Delta_{\mathrm{match}}(\theta)= \widehat U_{\mathcal V_{\mathrm{bench}}}(\{\theta\})- \widehat U_{\mathcal V_{\mathrm{match}}}(\{\theta\}).

Scroll sideways for the whole equation

A score is the average of the scored outcomes u_j over N_eval evaluation instances (Equation 24.1), for one named system, task draw, scoring rule and contract V. Delta_match is the benchmark pass proportion minus the matched-bank pass proportion for the same fixed model. The two-bank case has independent tasks, so the standard error of the gap is sqrt(p1(1 - p1)/n + p2(1 - p2)/n) and the 95% interval is the gap plus or minus 1.96 standard errors. The record cases come from the laboratory's paired design: a procedure's rate is its successes over its own runs, a matched difference uses only task/run pairs both procedures attempted, and a target mixture weights task-specific rates and needs every positively weighted task observed. The aged-benchmark case mixes success on unfamiliar tasks (0.4) with success on exposed tasks (1.0) at an exposed share of 20 percent.

Predict first. In the two-bank case on 'The difference between two scores', the same 0.07 gap is shown at 100, 400 and 1600 tasks per bank. At 100 tasks does the interval include zero?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Observations, denominators and estimands. Two bars, benchmark 0.80 and matched 0.73, each with a 95 percent interval, from 400 independent tasks per bank.
Observations: Two independent banks (the chapter's table), What to read off: Scores and their denominators
Constructed example: the book's teaching counts (320 and 292 passes out of 400, not GSM1k measurements), the laboratory's default, changed and transfer record cases evaluated with the laboratory's own function, and the chapter's constructed aged-benchmark mixture (0.4 unfamiliar, 1.0 exposed, 20 percent exposed).

Calculated values

Benchmark score
0.80
Matched score
0.73
Tasks in each bank
400
Half-width, benchmark
0.0392
Half-width, matched
0.0435

Benchmark = (1 / 400) x 320 = 0.80; matched = (1 / 400) x 292 = 0.73, the average of the scored outcomes in each bank (Equation 24.1). Half-widths: 1.96 x sqrt(0.80 x 0.20 / 400) = 0.0392 and 1.96 x sqrt(0.73 x 0.27 / 400) = 0.0435. Each score is only meaningful with its denominator and its contract.

Worked steps

  1. Benchmark: 320 passes of 400 tasks, so 320 / 400 = 0.80.
  2. Matched: 292 passes of 400 tasks, so 292 / 400 = 0.73.
  3. Benchmark half-width = 1.96 x sqrt(0.80 x 0.20 / 400) = 0.0392.
  4. Matched half-width = 1.96 x sqrt(0.73 x 0.27 / 400) = 0.0435.

Use the idea

Before reporting that a procedure improved, report each denominator, the number of matched pairs, the interval and the target mixture, so a reader can tell a real discrepancy from a thin sample, a missing cell or a changed population.

Where the conclusion applies

The two-bank interval needs independent samples with adequately populated pass and fail counts and does not cover choosing the largest gap after inspecting many models. The paired standard error and the Wilson intervals need independent representative pairs or runs, which repeated runs of the same task may violate; identifiers declare pairing, not randomization. The aged-benchmark mixture is the chapter's declared construction, not a measured contamination effect.

Common wrong turn: An interval that excludes zero identifies contamination
The chapter says the interval excludes zero under its sampling model but does not identify contamination, memorization, changed reasoning or an unmatched task feature as the cause. Matching observed features is not randomized assignment of training exposure.
What this does not settle

A matched held-out set can strengthen one comparison while leaving unmeasured task differences, selection effects and transfer limits open. The bounded conclusion is that an evaluation supports claims about the declared system, protocol, sample and task distribution that produced its score.

Chapter 24 source: "What this does not settle".

Check your understanding: Work this one by hand (800 tasks per bank is not shown). Benchmark 0.80 and matched 0.73 on 800 tasks each: what is the standard error of the gap, and does the interval exclude zero? And what is the aged benchmark's aggregate if 50 percent of its tasks were exposed?
SE = sqrt(0.80 x 0.20 / 800 + 0.73 x 0.27 / 800) = sqrt(0.0002 + 0.000246) = sqrt(0.000446) = 0.0211, between the 400-task value 0.0299 and the 1600-task value 0.0149, so the interval 0.07 plus or minus 1.96 x 0.0211 = [0.029, 0.111] excludes zero. At a 50 percent exposed share the aggregate is 0.5 x 1 + 0.5 x 0.4 = 0.70, which is 0.30 above the unfamiliar-task rate of 0.40.

Chapter 24 source: section "The matched question is not the same question twice". Demonstration C24-D01.

Demonstration 2 of 4

An onset belongs to its threshold

If a score rises gradually with scale, how much does the reported onset of a capability depend on the chosen threshold and on which sizes were tested?

The score curve is smooth, but a pass indicator turns on at the first member that reaches the threshold. Raising the threshold moves that crossing to the right. Testing fewer members can move it further, and can remove it from the tested set altogether. The onset is a property of the curve, the threshold and the sampling of sizes together.

Equation (24.3), written in LaTeX: \operatorname{Onset}_{\zeta}= \inf\{m: g(m)\ge \zeta\}.

Scroll sideways for the whole equation

m numbers the members of a model family in their declared scale order. g(m) is a continuous score under one fixed protocol. zeta is the threshold. Onset is the smallest m whose score reaches zeta; inf means the smallest such value in the tested set. If the set is empty there is no onset.

Predict first. At threshold 0.70 with only members 1, 3, 5 and 7 tested, is there an onset?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: An onset belongs to its threshold. Left: a rising score curve over eight family members with a threshold line at 0.50; onset at member 6. Right: pass indicators that jump from 0 to 1 at the onset.
Threshold zeta: 0.5, Family members tested: All eight members
Constructed example: eight family scores defined for this reader (0.04, 0.09, 0.17, 0.30, 0.46, 0.58, 0.66, 0.71), not a measured scale profile.

Calculated values

Threshold
0.50
Members tested
1, 2, 3, 4, 5, 6, 7, 8
Onset
m = 6
Highest tested score
0.71
Onset with all eight tested
m = 6

Smallest tested m with g(m) >= 0.50 is m = 6, where g(6) = 0.58; margin = 0.58 - 0.50 = 0.08. The member before it in this tested set is m = 5 with g = 0.46, 0.04 short of the threshold. The score curve rises gradually; only the declared threshold turns it into a pass. A different threshold or a sparser tested set can move this onset.

Worked steps

  1. Threshold zeta = 0.50; tested members: 1, 2, 3, 4, 5, 6, 7, 8.
  2. Scores of the tested members: g(1) = 0.04, g(2) = 0.09, g(3) = 0.17, g(4) = 0.30, g(5) = 0.46, g(6) = 0.58, g(7) = 0.66, g(8) = 0.71.
  3. The first tested member with g(m) >= zeta is m = 6.
  4. With all eight members the onset would be m = 6.

Use the idea

When someone reports that an ability appears at a certain size, ask for the continuous curve, the threshold, and the onset at two or three nearby thresholds. If the onset moves a lot, say the onset is threshold-sensitive.

Where the conclusion applies

One constructed curve of eight ordered members under one stable protocol, with no sampling noise on g. A crossing is an operational record, not a certified phase transition. Real scores carry uncertainty that could move a crossing by itself.

Common wrong turn: The onset is a phase transition in the system
The chapter says Equation (24.3) cannot certify a phase transition: it records an operational crossing that depends on the threshold, the metric, the family order and the protocol, and practical importance does not make the threshold a law of nature.
What this does not settle

The chapter leaves a model's maximum capability, a universal measure of reasoning and a complete cause of any benchmark gap open. A tested set leaves untested prompts, tools, budgets, selectors and future model states open.

Chapter 24 source: "What this does not settle".

Check your understanding: With all eight members tested, what is the onset at threshold 0.60? And with only the even members tested?
g(6) = 0.58 is below 0.60 and g(7) = 0.66 is at or above it, with margin 0.66 - 0.60 = 0.06, so the onset is m = 7. With only members 2, 4, 6 and 8 tested the first score at or above 0.60 is g(8) = 0.71, so the onset moves to m = 8.

Chapter 24 source: section "Scale profiles and extraction profiles are different maps". Demonstration C24-D02.

Demonstration 3 of 4

A lower frontier is evidence, not a ceiling

If several configurations are tested, what can we claim for the best one after paying for having looked at all of them?

Each configuration's lower bound is its observed rate minus a margin. To make all bounds hold together, the allowed failure probability is split across the tested configurations, which widens the margin. The frontier is the largest of those lower bounds. Configurations that were not tested do not appear in the maximum.

Equation (24.4), written in LaTeX: \Pr[\forall \xi\in\mathcal C:\operatorname{Lower}^{\mathrm{sim}}_{\alpha_{\mathrm{tail}}}(\xi)\le U_{\mathcal V}(S_\xi)]\ge 1-\alpha_{\mathrm{tail}}, \operatorname{LCF}_{\alpha_{\mathrm{tail}}}(\mathcal C)=\max_{\xi\in\mathcal C}\operatorname{Lower}^{\mathrm{sim}}_{\alpha_{\mathrm{tail}}}(\xi).

Scroll sideways for the whole equation

C is the finite set of tested configurations, xi one of them, and U its true success rate under the stated contract. Lower is a bound from a joint procedure that holds for every configuration at once with probability at least 1 minus alpha. LCF is the largest of those bounds. Here alpha is 0.05 and n is the number of tasks per configuration. In Equation (24.4), V is the declared evaluation contract (task distribution, evaluation unit, scoring rule, component partition, information and action interfaces, budget, horizon and baseline), S_xi is the set of components enabled in configuration xi, and Lower^sim is instantiated here as the observed rate minus the margin t.

Predict first. Keep 400 tasks but test only the first two configurations. Does the margin t shrink or grow compared with testing four?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A lower frontier is evidence, not a ceiling. Horizontal bars for 4 tested configurations, each with a solid lower bound and a hatched gap to its estimate, a dashed frontier line at 0.67, and a dotted bar marking the untested configurations as no evidence.
Tasks per configuration: 400, Configurations tested: 4
Constructed example: four configuration success rates (0.60, 0.68, 0.72, 0.74) defined for this reader, not results for any real system.

Calculated values

Configurations tested
4
Simultaneous margin t
0.074
Margin if one configuration were tested alone
0.061
Frontier value (LCF)
0.666
Configuration that sets it
Tools and selector

Margin t = sqrt(ln(4/0.05) / (2 x 400)) = sqrt(4.382 / 800) = 0.074. Lower bound for Tools and selector = 0.74 - 0.074 = 0.666, the largest of the 4 jointly supported bounds, so LCF = 0.666. Testing one configuration alone would use sqrt(ln(1/0.05) / (2 x 400)) = 0.061; the extra 0.013 is the price of having looked at 4. The frontier is evidence of what was achieved, not a ceiling.

Worked steps

  1. Each configuration gets alpha / 4 = 0.0125 of the failure budget.
  2. Margin t = sqrt(ln(4/0.05) / (2 x 400)) = 0.074.
  3. Lower bounds: Direct prompt 0.60 - 0.074 = 0.526, Retrieval 0.68 - 0.074 = 0.606, Retrieval and tools 0.72 - 0.074 = 0.646, Tools and selector 0.74 - 0.074 = 0.666.
  4. LCF is the largest lower bound, 0.666, set by Tools and selector.

Use the idea

If a program tries several prompts, tools or budgets, report how many were tried and use a margin that pays for the search, instead of presenting the highest score as if it were the only planned comparison.

Where the conclusion applies

The margin used here is one way to build the joint bound: a one-sided Hoeffding deviation for independent tasks scored between 0 and 1, with the failure budget split equally by a union bound. Other designs give other margins. If tasks are dependent, the independence-based margin is no longer justified (and may be too small under positive dependence).

Common wrong turn: The best tested score is the system's maximum
The chapter calls the frontier a lower frontier: the best tested result is evidence that at least that much was extracted under one tested configuration, and it leaves a finite tested set open to better untested configurations.
What this does not settle

A tested lower frontier leaves untested prompts, tools, budgets, selectors and future model states open. An evaluation supports claims about the declared system, protocol, sample and task distribution that produced its score.

Chapter 24 source: "What this does not settle".

Check your understanding: With 1600 tasks and four configurations tested at alpha 0.05, what is the margin t, and what is the frontier for a best estimate of 0.74?
t = sqrt(ln(4/0.05) / (2 x 1600)) = sqrt(4.382 / 3200) = 0.037, so LCF = 0.74 - 0.037 = 0.703.

Chapter 24 source: section "A frontier is lower evidence, not a final ceiling". Demonstration C24-D03.

Demonstration 4 of 4

Same average, different reliability

If two banks of tasks both average 0.50 per run, do they give the same chance that repeated runs succeed, and what do the observed runs of one task give?

The task is drawn once and then repeated, so the runs share the task's difficulty. Averaging p^k over tasks weights the easy task heavily, giving more than the mean squared. Averaging 1 - (1 - p)^k over tasks gives less than the mean-rate formula. The gap grows with the spread between tasks and vanishes when the tasks are the same. For a recorded bank the subset fractions count the observed runs directly, and they differ from plugging the observed rate into the formulas.

Equation, written in LaTeX: 1-(1-.9)^3=.999

Equation, written in LaTeX: .9^3=.729

Scroll sideways for the whole equation

The bank has two equally weighted tasks with single-run success probabilities p1 and p2. Runs are independent given the task. k is the number of runs. The chapter writes the single-task cases for one run probability 0.9 and three runs: success at least once, 1 - (1 - p)^k, and success on all runs, p^k. This demonstration applies each to each task, averages, and compares with applying it to the mean rate. The recorded bank is one task with 4 observed runs and 3 successes: a k-run subset contains only successes with fraction C(3,k) / C(4,k) and at least one success with fraction 1 - C(1,k) / C(4,k), where C(n,k) counts the ways to choose k of n runs.

Predict first. With tasks at 0.20 and 0.80 and two runs, is the true chance that both runs succeed above or below 0.25, the square of the 0.50 mean?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Same average, different reliability. Four bars: all 2 runs succeed 0.340 against the mean-rate value 0.250, and at least one succeeds 0.660 against 0.750.
Bank of tasks: Two tasks at 0.20 and 0.80 (the chapter's), Runs per task (subset size for the recorded bank): 2
Constructed example: the chapter's two-task construction (0.2 and 0.8, mean 0.5) with other spreads and run counts added for this reader, and the recorded bank of three successes and one failure from Mathematical Workbench E.5, part 3.

Calculated values

Mean single-run success
0.50
All runs, average of tasks
0.340
All runs, mean-rate formula
0.250
At least one, average of tasks
0.660
At least one, mean-rate formula
0.750

All 2 runs: (0.20^2 + 0.80^2) / 2 = (0.0400 + 0.6400) / 2 = 0.340, while 0.50^2 = 0.250. At least one: ((1 - 0.80^2) + (1 - 0.20^2)) / 2 = (0.3600 + 0.9600) / 2 = 0.660, while 1 - (1 - 0.50)^2 = 0.750. Squaring the mean rate misses by 0.090 for all-runs success (it understates) and the coverage-style formula misses by 0.090 for at-least-one success (it overstates). The task is drawn once and then repeated, which is the dependence the mean rate hides.

Worked steps

  1. Task rates 0.20 and 0.80; mean single-run rate 0.50.
  2. All 2 runs per task: 0.20^2 = 0.0400 and 0.80^2 = 0.6400; average 0.340.
  3. At least one in 2 runs: 0.3600 and 0.9600; average 0.660.
  4. Mean-rate formulas: 0.50^2 = 0.250 and 1 - (1 - 0.50)^2 = 0.750.

Use the idea

When a report gives one average success rate and then claims reliability over repeated runs, ask for the task-level repeat results. The pooled average cannot answer a question about repetition.

Where the conclusion applies

Two constructed tasks, equal weights, runs independent given the task and a frozen procedure. It fails when runs share a cached error, quota or evaluator fault, and it says nothing about a selector that must choose among the runs before the answer is known. The subset fractions are unbiased only under independent, identically distributed trials for that fixed task; otherwise they are descriptive fractions of the bank, and if fewer than k runs exist the subset metric is unavailable.

Common wrong turn: Squaring the mean success rate gives repeated-run reliability
The chapter says raising an average success rate to a power assumes away the variation that repetition is meant to expose: squaring the mean gives 0.25 where the task-first average gives 0.34, and applying coverage to the mean gives 0.75 where the task-first average gives 0.66.
What this does not settle

A matched held-out set can strengthen one comparison while leaving unmeasured task differences, selection effects and transfer limits open. The bounded conclusion is that an evaluation supports claims about the declared system, protocol, sample and task distribution that produced its score.

Chapter 24 source: "What this does not settle".

Check your understanding: Tasks at 0.10 and 0.90, three runs: what is the true chance that all three succeed, and the mean-rate value? And for a recorded bank of 3 successes in 4 runs, what fraction of two-run subsets contains only successes?
(0.1^3 + 0.9^3)/2 = (0.001 + 0.729)/2 = 0.365, against 0.5^3 = 0.125. Recorded bank: C(3,2) / C(4,2) = 3 / 6 = 0.5 of the two-run subsets contain only successes, and all 6 contain at least one success.

Chapter 24 source: section "The same average can describe different reliability". Demonstration C24-D04.