The Mathematics of AI Agents, laboratory reader ยท Chapter 19

When Another Mind Becomes Part of the World

A score earned beside one partner may not travel to another, so the partner belongs inside the measurement.

These four demonstrations follow the chapter's cross-play matrix. The first computes the loss statistic on the chapter's printed means and shows what its population fix did to three reported numbers. The second shows why simply optimizing against the current partner can go in circles. The third is the chapter's return on depth. The fourth shows how the partners you expect to meet decide which policy ranks first, on the laboratory's default, changed and transfer cases.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

Diagonal against off-diagonal

How much of a score is lost when a policy meets a partner from another training run instead of its own, and what did the fix change?

Equation (19.1) fills a table with one entry for every pairing of runs. Equation (19.2) averages the diagonal, averages the off-diagonal, subtracts, and divides by the diagonal average. A loss near zero means strangers score about what matched partners score; a loss near one means the score almost vanishes. The fix view shows the chapter's population method moving the loss and, on the smallest map, moving the two means in opposite directions.

Equation (19.1), written in LaTeX: M_{st} = \mathrm{E}[ u(\pi_1^{(s)}, \pi_2^{(t)})].

Equation (19.2), written in LaTeX: \operatorname{JPC}(M) = \frac{\operatorname{diag}(M) - \operatorname{off}(M)}{\operatorname{diag}(M)}.

Scroll sideways for the whole equation

M_st is the mean return when the first party uses the policy from training run s and the second party uses the policy from run t. diag(M) is the average of the entries with s equal to t (matched partners). off(M) is the average of the entries with s different from t (strangers). JPC(M), the joint policy correlation loss, is their difference divided by diag(M), a proportional loss. E is the mean over many episodes, u is the joint return of the two parties (their returns added together), and pi_1^(s), pi_2^(t) are the policies the first and second party use. Returns here are in the units of the chapter's table. In the fix view, 'highest-level only' and 'full mixed strategy' are the two ways the chapter's population method can be deployed.

Predict first. For the small4 map the diagonal mean is 20.15 and the off-diagonal mean is 5.71. Before selecting it, guess whether the loss is nearer 0.3 or 0.7.

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Diagonal against off-diagonal. Two bars, diagonal 30.44 above off-diagonal 20.03, and a bar chart of the five losses with Laser Tag small2 at 0.342.
Task: Laser Tag small2, View: Means by pairing
Values the chapter prints for three Laser Tag maps and two control tasks (the diagonal and off-diagonal means, and in the fix view its printed losses and reductions), plus recomputations marked as such, arranged here as a constructed teaching example; the loss is computed with the laboratory's cross-play function.

Calculated values

Diagonal mean
30.44
Off-diagonal mean
20.03
Loss from Equation (19.2)
0.342

Loss = (30.44 - 20.03) / 30.44 = 10.41 / 30.44 = 0.342. Strangers lose 34.2 percent of the diagonal mean on average. That compares these pairings only; it is not a universal measure of cooperation. The diagonal says how the developed pairs do together; only the off-diagonal says what happens with a stranger. Figure 19.1 shows the matrix is not symmetric: the entry for run four's first party with run one's second party is 27.3 and the reverse entry is 3.7, so 27.3 / 3.7 = 7.4 (the matrix has more in it than one summary number).

Worked steps

  1. Diagonal mean (own training partner) = 30.44.
  2. Off-diagonal mean (a stranger) = 20.03.
  3. Gap = 30.44 - 20.03 = 10.41.
  4. Loss = 10.41 / 30.44 = 0.342.

Use the idea

When two components of a system were developed together, evaluate them also against independently developed replacements, and report the diagonal, the off-diagonal and the loss, not only the diagonal.

Where the conclusion applies

The statistic is meaningful only when the diagonal mean is positive, and it compares particular pairings under one procedure. It is not a universal measure of cooperation and does not show why a gap appears. A small loss does not show the policies are independent or interchangeable. The fix view uses only losses the chapter prints, plus one subtraction for the small4 full-strategy loss.

Common wrong turn: Reading the loss as a universal measure of cooperation
The chapter says the denominator is meaningful only where the reference score supports the comparison, and the statistic is not a universal measure of cooperation. A loss of 0.003 records a small difference between the evaluated means; it does not show the parties are interchangeable.
What this does not settle

The anchor is a single conference paper reporting its authors' own method on gridworld tasks, and its reductions are their measurements rather than a general rate. The instrument requires multiple independent training runs, which many deployments cannot afford. In asymmetric games the effect varies across parties.

Chapter 19 source: "What this does not settle".

Check your understanding: A matrix has a diagonal mean of 10 and an off-diagonal mean of 8. What is the loss, and what changes if the diagonal mean is 0?
(10 - 8) / 10 = 0.20. With a diagonal mean of 0 the denominator is 0, so the statistic is undefined and the chapter says to report the difference instead.

Chapter 19 source: section "What the instrument found". Demonstration C19-D01.

Demonstration 2 of 4

A best-response cycle

If each party always plays its best answer to the other's current move, does the sequence settle?

Each update computes Equation (19.4) exactly against the option the other party is playing now. Because the other party then moves, the world the first party optimized against is gone, which is the added argument in Equation (19.3). In this game the sequence of options each party plays repeats every six updates. The diamonds show the same option scored against a fixed, uniform mix of the three options (equal weights): its average is 0, so the +1 gains never add up to progress against that set.

Equation (19.3), written in LaTeX: P(x' \mid x, a) \longrightarrow P(x' \mid x, a, \pi_{-i}).

Equation (19.4), written in LaTeX: \operatorname{BR}_i(\pi_{-i}) = \arg\max_{\pi_i} u_i(\pi_i, \pi_{-i}).

Scroll sideways for the whole equation

BR_i(pi_-i) is the policy for party i that maximizes its payoff u_i when the other party's policy pi_-i is held fixed. P(x' | x, a, pi_-i) is the chance of next state x' given state x, action a and the other party's policy. Here each party picks one of three options A, B, C. A beats C, B beats A and C beats B; a win pays 1, a loss pays (-1) and a tie pays 0. An update is one party switching to its best answer; the parties alternate, party one first.

Predict first. Start party two at A and show 6 updates. Does the sequence stop at some pair of options, or return to where it began?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A best-response cycle. Left, two stepped lines of the options each party plays over 6 updates starting with party two at A; party two returns to A after update 6. Right, bars of +1 for each mover's gain and diamonds at 0 for its average against all three options.
Updates shown: 6, Where party two starts: A
Constructed example: the three-option game and the six-update trace described in the chapter's worked cycle, with ties defined by this reader.

Calculated values

Updates shown
6
Party two back at its start
after 6 updates
Mover's payoff after its own update
+1 every time
Mover's average against A, B and C
0.00 every time
Mean payoff to party one
0.00

First update: with party two at A, party one's payoffs for A, B, C are 0, 1, (-1), so Equation (19.4) picks B. Mean payoff to party one over the 6 updates = (1 + (-1) + 1 + (-1) + 1 + (-1)) / 6 = 0.00. Each mover gains 1 against the partner it faces, yet the partner then moves and the gain is gone: option B pays 1 against A but (-1) against C, the move party two then makes (the dependence in Equation (19.3)). Against a fixed evaluation set that is a uniform mix of the three options, any option pays (1 + 0 + (-1)) / 3 = 0.00, so there is no progress there. Party two is back at A after update 6, so the cycle repeats.

Worked steps

  1. Update 1: party two plays A. Payoffs of A, B, C to party one are 0, 1, (-1), so party one plays B.
  2. Update 2: party one plays B. Payoffs of A, B, C to party two are (-1), 0, 1, so party two plays C.
  3. Update 3: party two plays C. Payoffs of A, B, C to party one are 1, (-1), 0, so party one plays A.
  4. Update 4: party one plays A. Payoffs of A, B, C to party two are 0, 1, (-1), so party two plays B.
  5. Update 5: party two plays B. Payoffs of A, B, C to party one are (-1), 0, 1, so party one plays C.
  6. Update 6: party one plays C. Payoffs of A, B, C to party two are 1, (-1), 0, so party two plays A.

Use the idea

When two adapting components keep re-tuning to each other, track a fixed set of evaluation partners rather than only the latest partner, because progress against the latest one can be undone by the next update.

Where the conclusion applies

A constructed zero-sum game with one-step pure best responses, ties paying 0 (a choice made for this reader), and no noise. Some games converge under the same rule, and other update rules behave differently. The cycle shows that convergence is not guaranteed; it does not explain the transfer losses in Demonstration 1.

Common wrong turn: Every correct update must make progress
Each of the six updates is an exact best response and improves its mover's payoff against the opponent it faces, yet the sequence goes nowhere. The chapter's point is that each update can improve against the current opponent without producing progress against a fixed evaluation population.
What this does not settle

Best-response cycling is one distinct difficulty with adapting counterparties. The reported 34.2 percent transfer loss does not identify a cycle or prove that the training procedure has no single destination.

Chapter 19 source: "Best-response cycling is one distinct difficulty with adapting counterparties".

Check your understanding: If party two starts at B instead, which option does party one answer with first, and how many updates until party two is back at B?
Party one's payoffs against B are A: (-1), B: 0, C: 1, so it plays C. Party two answers A, party one B, party two C, party one A, party two B, so party two is back at B after 6 updates.

Chapter 19 source: section "A cycle, worked". Demonstration C19-D02.

Demonstration 3 of 4

Return on depth

How much of the available reduction in transfer loss does each added stretch of the fix buy?

Each point is the loss left after the method is trained to that depth, and each bar of the third view is the extra reduction bought by moving to the next tested depth. Most of the reduction arrives by level three and almost all by level five, so the last doubling of effort buys under one point. The chapter also shows where its own printed figures and recomputed differences disagree (44 against 47.1 at level three), and this view keeps them separate.

Equation (19.2), written in LaTeX: \operatorname{JPC}(M) = \frac{\operatorname{diag}(M) - \operatorname{off}(M)}{\operatorname{diag}(M)}.

Scroll sideways for the whole equation

JPC(M), the joint policy correlation loss, is (diag(M) - off(M)) / diag(M), a proportional loss, shown here in percent of the diagonal mean. diag(M) is the mean return against the policy's own training partner and off(M) the mean return against a stranger. The method trains in levels: level 0 plays uniformly at random and each higher level learns a policy that best responds to a mixture over the levels below. Depth is the number of levels. The values are for the largest map, Laser Tag small4.

Predict first. Level five removes 56.1 points of loss. Will level ten remove clearly more or almost the same?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Return on depth. A line of loss at four depths of the method on small4, falling from 71.7 for independent learners to 24.6, 15.6 and a derived 15.0 (the printed reduction 56.7 subtracted from the baseline), with level three highlighted, and a bar chart of the loss.
Depth of the method: Level three, Right panel: Loss left
Constructed example: the chapter's printed small4 baseline and level three and five losses, and the printed level-ten reduction, used as a teaching example.

Calculated values

Depth
Level three
Loss
24.6 percent
Points removed from the baseline
47.1
Extra points from the last step
47.1

Points removed at level three = 71.7 - 24.6 = 47.1. The step from the previous tested depth added 47.1 - 0.0 = 47.1 points. The chapter prints 44 points for level three, but its displayed losses give 71.7 - 24.6 = 47.1; the chapter keeps the two figures separate and does not establish the cause. Compare the measured points before extrapolating: three tested depths do not establish a diminishing-return law.

Worked steps

  1. Baseline loss = 71.7 percent; level three leaves 24.6 percent.
  2. Points removed = 71.7 - 24.6 = 47.1.
  3. Previous tested depth removed 0.0, so this step added 47.1 - 0.0 = 47.1.
  4. The chapter prints 44 for this reduction; the recomputed 47.1 stays separate.

Use the idea

Before paying for more depth in a training procedure, measure the loss at two or three intermediate depths and compare the points each step buys with its cost.

Where the conclusion applies

One method on one map from one paper. Level ten is the printed reduction 56.7 subtracted from the baseline. The flattening is observed at three tested depths, not explained, and it may not appear on other maps, with other methods, or for other counterparties.

Common wrong turn: Treating the flattening as a law of diminishing returns
Three tested depths show a small additional improvement from level five to level ten in this setting. The chapter says they do not identify the reason for that flattening, establish a universal diminishing-return law, or show that remaining failures lie beyond the method's reach.
What this does not settle

The anchor is a single conference paper reporting its authors' own method on gridworld tasks, and its reductions are their measurements rather than a general rate.

Chapter 19 source: "What this does not settle".

Check your understanding: A different procedure loses 60 points at baseline, 25 at depth 4 and 20 at depth 8. How many extra points does the step from depth 4 to depth 8 buy?
Points removed at depth 4: 60 - 25 = 35. At depth 8: 60 - 20 = 40. The step adds 40 - 35 = 5 points, much less than the first 35.

Chapter 19 source: section "How much of the method is doing the work". Demonstration C19-D03.

Demonstration 4 of 4

Which partners will it meet?

If the diagonal favors one policy, can the partners you expect to meet make another policy the better choice?

The matrix holds fixed numbers. The weights decide how much each column counts, so changing the mix changes which row averages higher while the matrix, and the loss computed from it, stay the same. In the three-policy matrix, one policy succeeds with probability 0.6 against every partner and two specialists succeed only against their own, so the mix decides whether specialization or breadth wins.

Equation (19.1), written in LaTeX: M_{st} = \mathrm{E}[ u(\pi_1^{(s)}, \pi_2^{(t)})].

Equation (19.2), written in LaTeX: \operatorname{JPC}(M) = \frac{\operatorname{diag}(M) - \operatorname{off}(M)}{\operatorname{diag}(M)}.

Scroll sideways for the whole equation

M_st is the success probability of policy s (row) against partner t (column). The weight under each column is the probability that the policy meets that partner. A policy's value against the mix is its row averaged with these weights. JPC(M), the joint policy correlation loss, is the loss of Equation (19.2) for the whole matrix. In this demonstration u is a success indicator (1 for success, 0 otherwise), so each entry M_st = E[u] is a success probability. The changed (supervisor) mix is the distribution a supervisor routes to; the laboratory calls it the changed partner or supervisor distribution.

Predict first. With the two-policy matrix and equal weights, policy 0 has the better diagonal entry. Which policy has the better value against the mix?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Which partners will it meet?. A 2 by 2 matrix of success probabilities with the diagonal outlined, and bars comparing each policy's matched-partner value with its value against the mix; Policy 1 ranks first.
Matrix: Two policies (default case), Partner mix: Declared mix
Constructed example: the laboratory's default two-policy matrix with its default and changed partner weights, and its transfer three-policy matrix with the transfer weights, computed with the laboratory's cross-play function.

Calculated values

Policy 0 against the mix
0.575
Policy 1 against the mix
0.650
Ranks first
Policy 1
Diagonal alone names
Policy 0
Loss from Equation (19.2)
0.676

Policy 0 = 0.50 x 0.95 + 0.50 x 0.20 = 0.575. Policy 1 = 0.50 x 0.40 + 0.50 x 0.90 = 0.650. Policy 1 ranks first, although the diagonal alone names Policy 0. The matrix is not symmetric: cell (0, 1) is 0.20 but cell (1, 0) is 0.40, so a row and the matching column answer different questions. The loss is fixed by the matrix, whatever the mix: (0.925 - 0.300) / 0.925 = 0.676.

Worked steps

  1. Weights on partners 0 to 1 are 0.50, 0.50 (the declared partner mix).
  2. Policy 0 = 0.50 x 0.95 + 0.50 x 0.20 = 0.575.
  3. Policy 1 = 0.50 x 0.40 + 0.50 x 0.90 = 0.650.
  4. Policy 1 ranks first, although the diagonal alone names Policy 0.

Use the idea

Before choosing a component, write down the partner distribution you expect in use, and check whether the ranking survives plausible changes to it, instead of ranking by the diagonal.

Where the conclusion applies

Two constructed matrices and declared partner mixes; the weighted average over partners is this laboratory's way of asking the deployment question, not a formula printed in the chapter. It assumes the matrix entries share one success definition, and that partners do not change how they behave once a policy is chosen.

Common wrong turn: Assuming the matrix is symmetric
Readers tend to assume a cross-play matrix is symmetric, and that assumption produces a wrong diagnosis. The cell for one role assignment and the cell for the opposite assignment record different pairings, so read rows (a first-party policy across partners) and columns (a second-party policy across replacements) separately.
What this does not settle

A partner mixture is a declared deployment distribution. The chapter says the off-diagonal measures the tested replacements under the same evaluation conditions; it does not predict every future replacement, and Equation (19.2) cannot supply missing replacements.

Chapter 19 source: "Equation (19.2) summarizes the measured matrix; it cannot supply missing replacements".

Check your understanding: For the two-policy matrix with weight 0.7 on partner 0, which policy ranks first, and what are the two values?
Policy 0: 0.7 x 0.95 + 0.3 x 0.2 = 0.725. Policy 1: 0.7 x 0.4 + 0.3 x 0.9 = 0.55. Policy 0 ranks first, since 0.7 is above the crossing at 0.56.

Chapter 19 source: section "Self-play is the diagonal". Demonstration C19-D04.