The Mathematics of AI Agents, laboratory reader ยท Chapter 2

How to Measure Emergence

A total score cannot say what two components add together. Four matched measurements can.

These four demonstrations follow the chapter's instrument: the interaction contrast built from four matched scores. The first reads the book's cases, the workbench pair and two ways the same four numbers can mislead; the second shows why the neither cell matters; the third walks the maple-river route and a removal control that leaks; the fourth asks how many trials the contrast needs and what a search over many pairs does.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

Four cells, one contrast

Given the four scores for neither, p only, i only and both, how much did the pair add beyond its two isolated gains, and when may that be read as a share?

Equation (2.1) takes the score with both operations and subtracts what each operation reaches alone. That removes the neither score twice, so it is added back once. What is left is the part of the joint score that the two separate scores cannot explain. The steps on the right build the joint score in that order, and the bar for Gamma starts where the additive prediction ends. The run control shows two ways the same four numbers can mislead: a joint run with a bigger budget, and a different score scale.

Equation (2.1), written in LaTeX: \Gamma(p;i)=U(\{\theta,p,i\})-U(\{\theta,p\})-U(\{\theta,i\})+U(\{\theta\})

Equation (2.2), written in LaTeX: \begin{aligned}U(\{\theta,p,i\})-U(\{\theta\})&=[U(\{\theta,p\})-U(\{\theta\})]\\& +[U(\{\theta,i\})-U(\{\theta\})]+\Gamma(p;i).\end{aligned}

Equation (2.3), written in LaTeX: \rho_\Gamma=\frac{\Gamma(p;i)}{U(\{\theta,p,i\})-U(\{\theta\})}

Scroll sideways for the whole equation

U is a score where higher is better. In U({theta, ...}), theta is the frozen model and the braces list the enabled components. U({theta}) is the system with neither operation, U({theta, p}) with only the operation p that prepares a clue, U({theta, i}) with only the operation i that uses the clue, and U({theta, p, i}) with both. Gamma is the interaction contrast. The ratio rho divides Gamma by the joint gain (both minus neither); it is a fraction between 0 and 1 only with a positive joint gain, nonnegative isolated gains and a nonnegative Gamma. A budget is the resource allowance of one cell, in the units of the reader (10 for each cell unless the joint run is doubled to 20).

Predict first. Choose the case with both operations at the additive 0.35 (singles 0.30 and 0.25, neither 0.20). Before you read the numbers, predict whether Gamma is positive, zero or negative.

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Four cells, one contrast. Left, a two by two grid of the four scores neither 0.20, p only 0.30, i only 0.25, both 0.80. Right, a step chart building the joint score from the neither score, the two gains and Gamma = 0.45.
Four-cell scores: Table 2.1: 0.20, 0.30, 0.25, 0.80, Run and scale: Matched budgets
Constructed example: the baseline 0.20 and Table 2.1 scores, the second (negative) example and the workbench pair I.1 and I.3 (0.40, 0.52, 0.49 with the joint score 0.73 under a doubled allowance and 0.58 after a matched rerun) are the book's own constructed numbers; the additive case is the book's additive prediction 0.35. Computed with the laboratory's four-cell function.

Calculated values

Joint gain (both minus neither)
0.60
Additive prediction
0.35
Interaction contrast Gamma
0.45
Sign of Gamma
positive
Ratio Gamma / joint gain
0.75
Valid as a fraction
yes
Budgets matched
yes

Gamma = 0.80 - 0.30 - 0.25 + 0.20 = 0.45. The additive prediction is 0.30 + 0.25 - 0.20 = 0.35, and the pair scored 0.80. Joint gain = 0.80 - 0.20 = 0.60. Gamma is positive: the pair delivers 0.45 more than the two isolated gains (0.10 + 0.05 = 0.15) predict, on this task and score only. Ratio = 0.45 / 0.60 = 0.75; the three conditions hold, so it is a fraction between 0 and 1. All four budgets are equal, so the four cells are a matched comparison; they still do not identify a mechanism.

Worked steps

  1. Isolated gains: p = 0.30 - 0.20 = 0.10; i = 0.25 - 0.20 = 0.05.
  2. Joint gain = 0.80 - 0.20 = 0.60.
  3. Gamma = 0.80 - 0.30 - 0.25 + 0.20 = 0.45.
  4. Equation (2.2) check: 0.10 + 0.05 + 0.45 = 0.60, the joint gain.
  5. Ratio = 0.45 / 0.60 = 0.75; the three conditions hold, so it is a fraction between 0 and 1.
  6. All four budgets are equal: a matched comparison, which still does not identify a mechanism.

Use the idea

When a team reports that two components together beat each alone, ask for the neither score, compute Gamma, check that the budgets match and say which scale the scores are on. Only then can anyone say whether the pair is better than its parts or merely the sum of them.

Where the conclusion applies

Everything except the two switched operations is held fixed: the same model, task distribution, scoring rule and removal controls. The scores are constructed teaching values and stipulated means, not noisy estimates. If the four conditions quietly differ in other ways, the arithmetic still balances but the number answers a different question. Cubing the scores is a change of scale defined for this reader.

Common wrong turn: The highest score means the pair cooperates
In the chapter's negative example the pair has the best number in the table, 0.70, which a leaderboard would report, yet its improvement is smaller than the additive reference predicts. Only the four-cell comparison shows it.
What this does not settle

A positive value identifies a positive additive interaction on the declared score scale, without identifying its mechanism.

Chapter 2 source: "What this does not settle".

Check your understanding: In the second example the singles are 0.55 and 0.50 and both operations score 0.70. What is Gamma, and may the ratio be called a fraction?
0.70 - 0.55 - 0.50 + 0.20 = (-0.15). The joint gain is 0.70 - 0.20 = 0.50, so the ratio is (-0.15) / 0.50 = (-0.30), a signed diagnostic and not a fraction because Gamma is negative. The pair has the highest score of the four cells yet scores 0.70 rather than the additive 0.85.

Chapter 2 source: section "Writing the four cells". Demonstration C02-D01.

Demonstration 2 of 4

Why the neither cell matters

Two systems both score 0.80 with the pair on. Can their contrasts differ, and what goes wrong if the neither score is skipped?

Both single-operation scores already contain the baseline, so a system that starts high has little room left to credit to the pair. Equation (2.1) adds the baseline back once, which keeps ordinary competence from counting as cooperation. Without that term the remainder equals Gamma minus the baseline, so it moves with the baseline rather than with the pair.

Equation (2.1), written in LaTeX: \Gamma(p;i)=U(\{\theta,p,i\})-U(\{\theta,p\})-U(\{\theta,i\})+U(\{\theta\})

Equation (2.2), written in LaTeX: \begin{aligned}U(\{\theta,p,i\})-U(\{\theta\})&=[U(\{\theta,p\})-U(\{\theta\})]\\& +[U(\{\theta,i\})-U(\{\theta\})]+\Gamma(p;i).\end{aligned}

Scroll sideways for the whole equation

System A has baseline 0.20 and single-operation scores 0.30 and 0.25. System B has the chosen baseline, and each single-operation score sits 0.02 and 0.01 above it. The baseline is the neither score, U({theta}), where theta is the frozen model. An intact score is the score with both operations on. The three-cell remainder means the score with both minus the two single scores, leaving out the neither score.

Predict first. With System B's baseline at 0.70, is its four-cell contrast larger or smaller than A's 0.45? Then switch to three cells: what sign does B's estimate take?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Why the neither cell matters. Left, two bars both reaching 0.80, system A from a baseline of 0.20 and system B from a baseline of 0.70. Right, the estimate each method reports: Four-cell Gamma gives 0.45 for A and 0.07 for B.
Baseline of System B (neither score): 0.70, Estimate of the pair's contribution: Four cells, Equation (2.1)
Constructed example: System A and System B at baseline 0.70 are the book's Figure 2.3 values; the other baselines are defined for this reader. Computed with the laboratory's four-cell function.

Calculated values

System A baseline
0.20
System B baseline
0.70
System A, four-cell Gamma
0.45
System B, four-cell Gamma
0.07
System A, three cells
0.25
System B, three cells
-0.63

System A: Gamma = 0.80 - 0.30 - 0.25 + 0.20 = 0.45. System B: Gamma = 0.80 - 0.72 - 0.71 + 0.70 = 0.07. Both systems score 0.80, but B starts at 0.70 and A at 0.20, so B's pair adds 0.80 - 0.70 = 0.10 of improvement against 0.60 for A. B's contrast 0.07 is smaller than A's 0.45. Check: 0.07 is the joint gain 0.10 minus the isolated gains 0.02 + 0.01 = 0.03. Look at the baseline before comparing contrasts.

Worked steps

  1. System A: Gamma = 0.80 - 0.30 - 0.25 + 0.20 = 0.45.
  2. System B: baseline 0.70, singles 0.72 and 0.71; Gamma = 0.80 - 0.72 - 0.71 + 0.70 = 0.07.
  3. Three cells, leaving out the baseline: A = 0.80 - 0.30 - 0.25 = 0.25; B = 0.80 - 0.72 - 0.71 = -0.63.
  4. Each three-cell number is Gamma minus the baseline: A 0.45 - 0.20 = 0.25; B 0.07 - 0.70 = -0.63.

Use the idea

When comparing two systems with the same headline score, put the baseline next to the contrast. A large intact score with a high baseline can hide a small contribution from the pair.

Where the conclusion applies

The same score scale and the same controls for both systems. System A and the 0.70 case of System B are the book's Figure 2.3 values; the other baselines are defined for this reader. A smaller contrast does not make a system worse: the choice between systems also depends on reliability and cost.

Common wrong turn: Equal intact scores mean equal cooperation
The chapter's Figure 2.3 gives two systems the same intact score 0.80 with contrasts 0.45 and 0.07; the baselines, 0.20 and 0.70, differ.
What this does not settle

Neither interaction magnitude is a score to maximize. Choosing between the two systems should depend on marginal benefit, reliability, and cost under the intended use.

Chapter 2 source: "Neither interaction magnitude is a score to maximize".

Check your understanding: System B has baseline 0.50, so its singles are 0.52 and 0.51 and the pair scores 0.80. What are its four-cell and three-cell numbers?
Four cells: 0.80 - 0.52 - 0.51 + 0.50 = 0.27. Three cells: 0.80 - 0.52 - 0.51 = (-0.23), which is 0.27 - 0.50. The three-cell number is negative only because the baseline was removed twice.

Chapter 2 source: section "Why all four cells". Demonstration C02-D02.

Demonstration 3 of 4

The maple-river route and a leaking removal control

What do the two operations do in the maple-river route, and what happens to the contrast when the control for p still lets p's influence through?

In the route of Figure 2.1 neither operation is the task: a tag nobody reads predicts nothing, and a search for a tag never written finds nothing. Only the cell with both promotes river. A faithful control for p removes the tag, so the i-only cell scores low. A pattern-preserving ablation reuses the attention pattern that p helped create, so part of p's influence survives inside it and it cannot serve as a clean condition with p disabled. In the leak model defined here, the p-off cells then score too high, the contrast shrinks in proportion, and at full leakage it vanishes while Equation (2.2) still balances.

Equation (2.1), written in LaTeX: \Gamma(p;i)=U(\{\theta,p,i\})-U(\{\theta,p\})-U(\{\theta,i\})+U(\{\theta\})

Equation (2.2), written in LaTeX: \begin{aligned}U(\{\theta,p,i\})-U(\{\theta\})&=[U(\{\theta,p\})-U(\{\theta\})]\\& +[U(\{\theta,i\})-U(\{\theta\})]+\Gamma(p;i).\end{aligned}

Scroll sideways for the whole equation

P is the earlier operation that tags river as following maple; I is the later operation that looks back for the tag at the second maple. p and i are the book's names for these two declared operations. The four cells switch each one on or replace it by its control. The surviving share is the part of p's influence that still reaches the later operation through the control, for example through a reused attention pattern; the cells with p off then score that share of the way toward the matching cell with p on. It is a leak model defined for this reader: the chapter says only that the influence survives.

Predict first. If all of p's influence survives the control (share 1), what is the measured Gamma?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The maple-river route and a leaking removal control. Left, the tokens maple, river, maple with arrows for the two operations; the neither cell scores 0.200. Right, a falling line of the measured contrast against the share of p's influence that survives, now at Gamma = 0.450.
Cell shown: Neither operation, Share of p's influence that survives the control: 0 (clean control)
Constructed example: the four scores are the book's Table 2.1 and the route is the chapter's maple river example of Figure 2.1; the surviving-share model is defined for this reader.

Calculated values

Selected cell
neither
Score of the selected cell
0.200
Influence of p that survives
0.00
Neither / p only / i only / both
0.200, 0.300, 0.250, 0.800
Measured Gamma
0.450
Clean-control Gamma
0.45

Cell neither: No route at all. Its measured score is 0.200. With 0.00 of p's influence surviving the control on the cells where p is off, neither = 0.20 + 0.00 x (0.30 - 0.20) = 0.200, and i only = 0.25 + 0.00 x (0.80 - 0.25) = 0.250. Gamma = 0.80 - 0.30 - 0.250 + 0.200 = 0.450, which is 0.45 x (1 - 0.00) = 0.450. The control is clean, so the contrast is the Table 2.1 value 0.45.

Worked steps

  1. Selected cell: neither. P is off and I is off.
  2. No route at all.
  3. Cells with p off borrow 0.00 of p's effect: neither = 0.20 + 0.00 x (0.30 - 0.20) = 0.200.
  4. i only = 0.25 + 0.00 x (0.80 - 0.25) = 0.250.
  5. Gamma = 0.80 - 0.30 - 0.250 + 0.200 = 0.450.
  6. Check: 0.45 x (1 - 0.00) = 0.450.

Use the idea

Before running the four cells, write down what replaces each operation and test that the replacement really removes the information under test. Declare the control and use the same one in all four cells.

Where the conclusion applies

Table 2.1's four scores are taken as the clean-control values. The leak model, a straight mix between the clean cell and the matching p-on cell, is defined for this reader and is not a measurement. The chapter says zero, mean and resampled removal give three different numbers and none is the true one; this demonstration does not model them.

Common wrong turn: An easier ablation answers the cooperation question
The chapter says an easier ablation answers a different question, whether a component's output matters once the rest of the computation is fixed, and that the second one silently answering for the first is the most common way a cooperation claim goes wrong.
What this does not settle

The contrast is defined for one pair on one task under one protocol; it is silent about every other route in the model.

Chapter 2 source: "What this does not settle".

Check your understanding: With half of p's influence surviving, what do the neither and i-only cells read, and what is Gamma?
Neither: 0.20 + 0.5 x (0.30 - 0.20) = 0.25. I only: 0.25 + 0.5 x (0.80 - 0.25) = 0.525. Gamma = 0.80 - 0.30 - 0.525 + 0.25 = 0.225, which is 0.45 x (1 - 0.5).

Chapter 2 source: section "What the experiment has to hold still". Demonstration C02-D03.

Demonstration 4 of 4

Is the contrast clearly different from zero?

How many trials per cell does it take before a contrast of 0.05 or 0.15 is clearly different from zero, and what does searching many pairs do?

Equation (2.1) adds and subtracts four measured cells. If their errors are independent, their variances add, so the standard error of Gamma is twice that of one cell, against about 1.4 times for the difference of two intact scores. The same sample can therefore resolve a gap between intact scores while leaving a contrast of that size unresolved; a true amount of 0.15 at 100 trials shows it. Examining many pairs adds a second problem: some will clear zero by chance alone.

Equation (2.1), written in LaTeX: \Gamma(p;i)=U(\{\theta,p,i\})-U(\{\theta,p\})-U(\{\theta,i\})+U(\{\theta\})

Scroll sideways for the whole equation

n is the number of trials in each of the four cells. An intact score is the score with both operations on, and theta in Equation (2.1) is the frozen model. The per-trial spread of a score is set to 0.4 for every cell. A standard error is the typical size of the error in a measured average: 0.4 divided by the square root of n for one cell. The interval is the true amount plus or minus 1.96 standard errors, which a pair with no real contrast clears about 0.05 of the time. The chapter's model with 384 attention heads has 384 x 383 / 2 = 73,536 pairs.

Predict first. At 400 trials per cell, does a true contrast of 0.05 stand clear of zero? What about 0.15?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Is the contrast clearly different from zero?. Left, standard errors for one cell, a difference of two intact scores and the interaction contrast at 100 trials per cell. Right, intervals around a true amount of 0.15; the Gamma interval runs from -0.007 to 0.307.
Trials per cell: 100, True interaction contrast: 0.15, Pairs examined: One pair, named in advance
Constructed example: the spread 0.4, the trial counts and the true contrasts are values defined for this reader; the cells use Table 2.1 for the three lower scores. The factor of 2 and the 384 heads with roughly seventy thousand pairs are the chapter's own.

Calculated values

Standard error of one cell
0.040
Standard error of Gamma
0.080
Standard error of a difference of two intact scores
0.057
Interval for Gamma
-0.007 to 0.307
Smallest Gamma that clears zero
0.157
Interval excludes zero
no
Pairs searched
1
Chance clears among pairs with no real contrast
about 0.05 of 1

One cell: 0.4 / sqrt(100) = 0.040. Gamma adds four independent cells, so its variance is 4 x 0.040^2 and its standard error is 2 x 0.040 = 0.080, against sqrt(2) x 0.040 = 0.057 for a difference of two intact scores. Interval: 0.15 - 1.96 x 0.080 = -0.007 up to 0.15 + 1.96 x 0.080 = 0.307. The lower end -0.007 is not above zero, so a contrast of 0.15 is not clear of zero here, even though a same-size difference of intact scores (lower end 0.039) is resolved. One pair named in advance: a pair with no real contrast would clear zero about 1 x 0.05 = 0.05 of the time.

Worked steps

  1. One cell: 0.4 / sqrt(100) = 0.040.
  2. Gamma has four independent cells: 2 x 0.040 = 0.080.
  3. Interval half-width = 1.96 x 0.080 = 0.157.
  4. Interval for Gamma = 0.15 +/- 0.157, from -0.007 to 0.307.
  5. Pairs searched: 1; chance clears with no real contrast = 1 x 0.05 = 0.05.

Use the idea

Before running a four-cell comparison, decide the smallest contrast that matters and size the trials so that its interval would exclude zero. Name the pair before measuring it and report the contrast with an interval, not alone.

Where the conclusion applies

Independent cells with equal spread, a normal approximation and a spread of 0.4 defined for this reader. The interval shown is centred on the true amount; a real run would scatter around it. Cells measured on the same prompts are correlated, which changes the factor of 2 but not the lesson. The count of chance clears treats pairs as independent, which pairs sharing heads are not.

Common wrong turn: A striking pair found by search is a finding about the model
The chapter says finding one pair with a striking contrast after examining thousands of candidates is a fact about search, not about the model.
What this does not settle

A contrast is a single measurement of a single pair on a single task. Systems have many pairs.

Chapter 2 source: "What one measurement cannot do".

Check your understanding: With 1600 trials per cell, what is the smallest contrast whose interval excludes zero?
One cell: 0.4 / sqrt(1600) = 0.010. Gamma: 2 x 0.010 = 0.020. Smallest contrast = 1.96 x 0.020 = 0.039, so 0.05 clears zero and 0.03 does not.

Chapter 2 source: section "What the experiment has to hold still". Demonstration C02-D04.