The Mathematics of AI Agents, laboratory reader ยท Chapter 1

When a Score Becomes a Skill

A reported capability depends on the model and also on the rule that reads its outputs, so a jump on a chart can come from the ruler.

These four demonstrations follow the chapter's argument about measurement. The first moves a pass line over fixed scores, the second changes the rules and the loop around an unchanged model, the third shows how raising a probability to the power of a target length turns small gains into large ones, and the fourth audits a claim before anyone attributes it to the model.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

The same scores, a moving cutoff

If no score changes, can moving the pass line move the place where a skill seems to arrive?

The graded scores on the left never change. The verdicts on the right come from comparing each score with the cutoff, as in the indicator inside Equation (1.2). A smooth rise therefore produces a sudden verdict wherever it crosses the line, and a different line puts the sudden change at a different checkpoint.

Equation (1.2), written in LaTeX: \widehat U_{\mathcal V}(\{\theta\})=\frac{1}{n}\sum_{i=1}^{n}\mathbf{1}\{\operatorname{eval}_\tau(\hat z_i)=1\}

Scroll sideways for the whole equation

z-hat_i is the output generated for prompt i and n is the number of items being scored. eval_tau(z-hat_i) is the evaluator's verdict for that output, 1 for success and 0 for failure, and tau is its cutoff. The hat on U marks an estimate: the average verdict. U with subscript V is the score under a declared evaluation contract V (task, scoring rule and evaluator settings, fixed in advance); theta is the model's parameters, and the braces list the enabled components, here only the frozen model. In this demonstration each item is a checkpoint with a graded score between 0 and 1, and the evaluator passes it when the score reaches the cutoff. A score exactly equal to the cutoff passes. A verdict change is a pair of neighbouring checkpoints where one fails and the other passes.

Predict first. With the gradual scores, move the cutoff from 0.50 to 0.55. At which checkpoint does the first pass appear?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The same scores, a moving cutoff. Left, graded scores rising across 5 checkpoints with a dashed cutoff line at 0.50. Right, the verdicts 00011, first passing at scale 4.
Score set: Gradual rise, five checkpoints, Pass cutoff: 0.5
Constructed example: the five scores and the larger-rise four-score set with its cutoff of 0.70 are the laboratory's declared inputs for this chapter (default and transfer cases), computed with its threshold function; the other cutoffs are values defined for the reader.

Calculated values

Verdicts
00011
Checkpoints passing
2 of 5
Verdict changes between checkpoints
1
First passing scale
4
Largest rise per scale unit
0.040 per scale unit

Verdicts at cutoff 0.50 (0.42 < 0.50 gives 0, 0.46 < 0.50 gives 0, 0.49 < 0.50 gives 0, 0.52 >= 0.50 gives 1, 0.56 >= 0.50 gives 1). Passes = 0 + 0 + 0 + 1 + 1 = 2 of 5. The largest rise per scale unit is (0.46 - 0.42) / (2 - 1) = 0.040 score per scale unit. The verdict jumps from fail to pass once, between scale 3 and scale 4, where the graded score moved by 0.03, against 0.04 for the largest rise in score between neighbouring checkpoints in this set. First pass for each cutoff offered on this score set: 0.50 gives scale 4, 0.55 gives scale 5, 0.60 gives no pass, 0.70 gives no pass. The scores never moved; only the line did. The first-pass date moves only when the line passes over the score that first clears it; while the line stays inside a gap between scores, the date stays put.

Worked steps

  1. Graded scores at scales 1, 2, 3, 4, 5: 0.42, 0.46, 0.49, 0.52, 0.56.
  2. Cutoff = 0.50. A score passes when it is at least the cutoff; equality passes.
  3. Verdicts: 0.42 < 0.50 gives 0, 0.46 < 0.50 gives 0, 0.49 < 0.50 gives 0, 0.52 >= 0.50 gives 1, 0.56 >= 0.50 gives 1.
  4. Passes = 0 + 0 + 0 + 1 + 1 = 2 of 5.
  5. Largest rise per scale unit = (0.46 - 0.42) / (2 - 1) = 0.040 score per scale unit.
  6. First pass at scale 4.

Use the idea

When a dashboard turns green, ask for the underlying graded scores and the cutoff behind the colour. If moving the cutoff a little moves the date of arrival, the date describes the cutoff as much as the system.

Where the conclusion applies

The same task, output boundary and environment at every checkpoint, and scores that are exact rather than noisy estimates. Five or four constructed points cannot show whether the true curve is smooth between them. A genuinely sharp change inside the model would also produce a jump, and this picture cannot rule that out.

Common wrong turn: A score that crosses a pass line is a new internal state
The chapter warns that a curve that looks sudden is not a discontinuity and a score that crosses a pass line is not a new internal state. Here every graded score is identical in every state; only the line moves.
What this does not settle

The re-scoring result covers the evaluations Schaeffer and colleagues examined, not every reported curve. These constructed points cannot show whether a true curve between them is smooth or sharp.

Chapter 1 source: "What this does not settle".

Check your understanding: With the larger-rise scores (0.1, 0.3, 0.6, 0.8 at scales 10, 20, 40, 80) and a cutoff of 0.70, which checkpoints pass, and what is the first graded step?
Only the last one: 0.8 >= 0.7, while 0.6 < 0.7. The first step is (0.3 - 0.1) / (20 - 10) = 0.02 score per scale unit.

Chapter 1 source: section "What did change". Demonstration C01-D01.

Demonstration 2 of 4

Three places a change can enter, and the loop

If the model's probabilities never change, how much can the reported score move?

Equation (1.3) joins three stages by two arrows: the model's law, the decoder and the evaluator. The left panel is the first stage, Equation (1.1), and is frozen. The decoder decides which response is returned, and the evaluator decides which responses count. The score in the right panel combines the two, so it cannot say which stage moved. Allowing twenty attempts adds the loop: the chance that some attempt passes climbs toward one with the law untouched.

Equation (1.1), written in LaTeX: K_\theta(z\mid c)=\Pr_\theta(Z=z\mid C=c)

Equation (1.3), written in LaTeX: K_\theta(z\mid c)\longrightarrow\hat z=\operatorname{dec}(K_\theta(\cdot\mid c))\longrightarrow y=\operatorname{eval}_\tau(\hat z,c)

Scroll sideways for the whole equation

K_theta(z | c) is the model's probability for output z given context c, Equation (1.1): the whole law over responses, frozen in this demonstration. dec picks one output from that law: greedy decoding takes the most probable one, sampling draws one at random according to the probabilities. eval_tau judges the chosen output, 1 for success and 0 for failure. The expected score adds up, over the responses, the chance the decoder returns each one times its verdict. With twenty attempts, an attempt succeeds with the total probability of the accepted responses and the score is the chance that at least one of twenty independent attempts succeeds.

Predict first. Keep greedy decoding on the chapter's law and switch the evaluator from accepting only A to rewarding a request for evidence. What happens to the score?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Three places a change can enter, and the loop. Left, bars for the model law 0.45, 0.35, 0.20 marking which responses are judged successes. Right, the share of the score each response contributes, total 1.00.
Decoding rule: Greedy (most probable), Evaluator: Accepts only A, Model law (A, B, evidence): Chapter: 0.45, 0.35, 0.20
Constructed example: the chapter's constructed distribution of 0.45, 0.35 and 0.20 over three responses and its twenty-attempt budget at probability 0.20; the second law is defined for the reader.

Calculated values

Model law (A, B, evidence)
0.45, 0.35, 0.20
Output returned
answer A
Evaluator
accepts only A
Score
1.00

Greedy decoding returns answer A, the response with the largest probability (0.45). The evaluator accepts only A. Score = 1.00 x 1 + 0.00 x 0 + 0.00 x 0 = 1.00. The model law (0.45, 0.35, 0.20) is the same for every decoder and evaluator choice in this law setting. Only the rule that picks an output and the rule that judges it changed, and the reported number moved with them.

Worked steps

  1. Model law for A, B, evidence: 0.45, 0.35, 0.20. It does not change with the decoder or the evaluator.
  2. Greedy decoding returns answer A, the response with the largest probability (0.45).
  3. The evaluator accepts only A: verdicts A 1, B 0, evidence 0.
  4. Score = 1.00 x 1 + 0.00 x 0 + 0.00 x 0 = 1.00.

Use the idea

When two reports of the same model disagree, ask which stage differs: the decoding settings, the scoring rule or the number of attempts. Report the stage that changed, not only the final number.

Where the conclusion applies

One prompt, three possible responses and independent attempts. The probabilities 0.45, 0.35 and 0.20 and the twenty attempts are the chapter's constructed values; the flat law 0.30, 0.30, 0.40 is a value defined for the reader. Real decoders and evaluators have more settings, attempts can be correlated, and several settings can differ at once.

Common wrong turn: The model solved it
The model is the component that produces the visible text, so it collects the credit for a result several components produced together. Here the law is fixed and the reported score still ranges from 0.00 to 1.00.
Check your understanding: Suppose the model law were 0.30, 0.30 and 0.40 for A, B and evidence, and the evaluator rewards evidence. What do greedy decoding and one sampled draw score?
Greedy returns evidence, the largest at 0.40, so it scores 1. One draw scores 0.30 x 0 + 0.30 x 0 + 0.40 x 1 = 0.40.

Chapter 1 source: section "Three places a change can enter". Demonstration C01-D02.

Demonstration 3 of 4

How exact match manufactures a cliff

How much does a small gain per token change an exact-match score, and how does target length decide it?

Equation (1.4) multiplies q by itself n times, so a short target amplifies a gain less than a long target does, and a long target can amplify it enormously. The left panel draws the four target lengths of Figure 1.4, with the selected one thick. The right panel compares the exact multiple with the first-order rule: the rule is close while n times the rise is small and fails once the rise is applied many times. The absolute slope can stay below one while the relative gain is large.

Equation (1.4), written in LaTeX: U=q^{n}, \text{so} \log U = n\log q

Equation, written in LaTeX: nq^{n-1}

Scroll sideways for the whole equation

q is the probability that one token is right, n is the number of tokens in the target, and U is the exact-match score: the probability that all n tokens are right, assuming tokens succeed independently. The logarithm form says the length n multiplies the per-token gain. The relative rise is q1 / q0 - 1. The first-order multiple 1 + n x rise is the chapter's elasticity remark: a small relative increase in q raises U by about n times as much. The absolute slope n q^(n-1) is the derivative dU/dq, how many points of score one extra point of q buys.

Predict first. For a target of 40 tokens, how many times larger is the score at q = 0.95 than at q = 0.90?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How exact match manufactures a cliff. Left, four curves of U = q^n for n = 5, 20, 40 and 100 with the curve for n = 40 drawn thick and its two marked points at q = 0.90 and q = 0.95. Right, a log-scale bar chart comparing the first-order multiple 3.2 with the exact multiple 8.7.
Target length n (tokens): 40, Per-token gain: q from 0.90 to 0.95
Constructed example: the per-token values 0.90, 0.95 and 0.99 and the target lengths of Table 1.1, computed from Equation (1.4); the coding-assistant case is the chapter's twenty-line example. The budget of 100 trials behind the expected-passes readout is a value defined for the reader.

Calculated values

Score at q = 0.90
0.0148
Score at q = 0.95
0.1285
Gain per token
5 percentage points
Gain in task score
a factor of 8.7
First-order multiple (1 + n x rise)
3.22
Absolute slope at start
0.657
Expected passes in 100 trials
1.5 then 12.9

0.90^40 = 0.0148 and 0.95^40 = 0.1285, so the ratio is 0.1285 / 0.0148 = 8.7. The model gained 5 points per token, yet the exact-match score changed by a factor of 8.7. The scores are rounded for display and the ratio is computed from the unrounded values. A short target amplifies a gain less than a long target does. In logarithms the gap is n x ln(0.95 / 0.90) = 40 x 0.0541 = 2.16, so the length multiplies the per-token gain. First-order rule: 1 + n x rise = 1 + 40 x 0.0556 = 3.22, more than 10 percent away from the exact 8.69: with n x rise = 2.22 the rise is applied 40 times and compounds. The absolute slope at q = 0.90 is n x q^(n - 1) = 40 x 0.90^39 = 0.657, below one, so at the starting q a one-point gain in q buys less than one point of score even though the relative gain is large. The slope at q = 0.95 is 5.411, and over the whole step the average slope is 2.275. Over 100 trials the expected passes are 1.5 then 12.9, which differ by 11.4, so the improvement is not hidden by low scores.

Worked steps

  1. Target length n = 40; per-token success goes from q = 0.90 to q = 0.95.
  2. Start: U0 = 0.90^40 = 0.0148.
  3. End: U1 = 0.95^40 = 0.1285.
  4. Ratio = 0.1285 / 0.0148 = 8.7, from the unrounded values.
  5. Logarithms: n x ln(0.95 / 0.90) = 40 x 0.0541 = 2.16.
  6. Relative rise in q = 0.95 / 0.90 - 1 = 0.0556.
  7. First-order multiple = 1 + 40 x 0.0556 = 3.22, against the exact 8.69.

Use the idea

Before reading a flat region at the start of a benchmark curve as no progress, compute what the score would be if per-token quality had improved smoothly. Then decide whether the benchmark can resolve those small values.

Where the conclusion applies

Tokens succeed independently and each has the same probability q. Real text violates this, so the mechanism is robust but the exact factors are constructed. The expected passes in 100 trials are expected counts, not observed results, and the first-order rule is a small-change approximation.

Common wrong turn: A flat observed curve means nothing was happening
The chapter says a flat observed region need not mean nothing was happening, and gives the success probabilities 0.000027 and 0.005921 at n = 100 for q = 0.90 and 0.95. For a budget of 100 trials, a number defined for the reader, the expected passes are 0.0 then 0.6, both under one.
What this does not settle

Equation (1.4) assumes tokens succeed independently, which real text violates; the mechanism is robust, the exact factors are constructed.

Chapter 1 source: "What this does not settle".

Check your understanding: A coding assistant gets each line right with probability 0.90. What is the chance a 20-line function is entirely right, and how does it change at 0.99 per line?
0.90^20 = 0.1216, about one in eight. 0.99^20 = 0.8179, about four in five, a factor of 0.8179 / 0.1216 = 6.7 from a nine-point gain per line.

Chapter 1 source: section "Why exact match manufactures a cliff". Demonstration C01-D03.

Demonstration 4 of 4

Audit a claim before attributing it

Which declarations must be fixed before an emergence claim can be tested, and what does the six-question audit say about two claims?

The left panel is Figure 1.2: each declaration beside the question it settles. Leaving a row open leaves its dispute alive, as in the published disagreements about the scale axis and the metric. The right panel applies the six audit questions. For the re-scored outputs a counterfactual was executed, so the audit assigns the cliff to the readout; for a bare comparison of two different models there is no matched counterfactual, so the correct output is no verdict.

Equation (1.2), written in LaTeX: \widehat U_{\mathcal V}(\{\theta\})=\frac{1}{n}\sum_{i=1}^{n}\mathbf{1}\{\operatorname{eval}_\tau(\hat z_i)=1\}

Equation (1.3), written in LaTeX: K_\theta(z\mid c)\longrightarrow\hat z=\operatorname{dec}(K_\theta(\cdot\mid c))\longrightarrow y=\operatorname{eval}_\tau(\hat z,c)

Scroll sideways for the whole equation

The five declarations of Figure 1.2 are the model family, the scale axis, the metric with its baseline, the class of predictors and the tolerance. A claim is checkable only once all five are fixed. The six audit questions are: what are the components, what couples them, what is the state, what is the observable, what is nonlinear, and what is the counterfactual. U with the braces of Equation (1.2) is the score of the system built from the enabled components, and the arrows of Equation (1.3) are the couplings between stages.

Predict first. Leave the scale axis open and fix the other four declarations. Is the emergence claim checkable?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Audit a claim before attributing it. Left, five declarations with all marked fixed. Right, the six audit rows for re-scored arithmetic outputs and the verdict: assigned to the readout.
Declaration left open: None (suppose all five fixed), Claim audited: Re-scored arithmetic outputs
Constructed example: the five declarations of Figure 1.2 and the two worked audits of the chapter's emergence audit section, shown as the chapter states them.

Calculated values

Declarations fixed (assumed)
5 of 5
Open declaration
none
Checkable in principle
yes
Audit rows answered
6 of 6
Audit verdict
assigned to the readout

Declarations fixed = 5 - 0 = 5 of 5; open: none. The claim is checkable only once all five are fixed; here all five are supposed fixed, so it would be checkable in principle. Audit rows the chapter answers for this claim = 6 - 0 = 6 of 6. The curve flattened, so the cliff is assigned to the readout, on the strength of an executed intervention.

Worked steps

  1. Figure 1.2 lists five declarations: model family, scale axis, metric and baseline, predictor class, tolerance.
  2. Open declarations: 0. Fixed = 5 - 0 = 5.
  3. The claim is checkable only when all five are fixed. Here all five are supposed fixed, so it would be.
  4. The audit asks six questions of the claim: components, coupling, state, observable, nonlinear, counterfactual.
  5. Rows the chapter answers for 're-scored arithmetic outputs': 6 - 0 = 6.
  6. The curve flattened, so the cliff is assigned to the readout, on the strength of an executed intervention.

Use the idea

Before crediting the model for a capability, write down the five declarations and the six audit answers. A blank counterfactual row means the claim has not been tested, whatever the chart shows.

Where the conclusion applies

The audit rows are the chapter's own worked answers; for the second claim the chapter answers only the components, the coupling and the counterfactual, and this demonstration leaves the other rows blank rather than invent them. Only one declaration is opened at a time here, and the model family and tolerance rows are not offered as the open one.

Common wrong turn: The most visible component earns the credit
The chapter names the common error: a system-level success is attributed to the most visible component, almost always the language model, even though the agent can supply the state, the loop, the permissions and the consequences.
What this does not settle

The three-condition definition of emergence is a working tool for this book rather than a settled position in the field.

Chapter 1 source: "What this does not settle".

Check your understanding: In the larger-model comparison, which audit row decides that no verdict is possible, and why?
The counterfactual row: there is no matched smaller system differing in exactly one thing, because the two models differ in parameters, data and training recipe. 3 rows answered of 6, and the missing one is the one that decides.

Chapter 1 source: section "An emergence audit". Demonstration C01-D04.