The Mathematics of AI Agents, laboratory reader ยท Chapter 14

Building a World Inside

An agent that plans inside a learned model scores well only as far as the model's error stays smaller than what the decision can tolerate.

These four demonstrations follow the chapter. First the reported temperature table, where the score inside a learned model and the score in the actual environment come apart. Then the finite certificate for a learned transition model: one number for how wrong the model is and how that error compounds against the actual error, why an optimizer finds the one place a model is generous and the action gap decides whether a ranking survives, and the number of imagined steps those pieces certify. Apart from the reported table, all worlds, errors and gaps are constructed teaching values.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

The score inside the model against the score in the world

If you pick the setting that scores best inside the learned model, how does it do in the actual environment?

The table has two columns because the researchers had two environments. Over the first four rows the columns run in opposite directions, so the setting that flatters the model most is the least useful in the world. Judging by the model column alone cannot show this; only the second column can. The scores do not supply the transition error epsilon or any action gap, so they do not give a certified horizon.

Equation (14.1), written in LaTeX: \epsilon = \max_{x,a} \tfrac{1}{2}\sum_{x'}|\hat P(x'\mid x,a) - P(x'\mid x,a)|.

Scroll sideways for the whole equation

Temperature is the sampling parameter of the imagined world: low values make it cleaner and more predictable, high values harder. The model score is the controller's score when evaluated in its own learned environment, and the actual score is its score in the real game. The random policy and the best leaderboard entry are the chapter's reference scores. The table does not report the transition error epsilon of Equation (14.1), which is the maximum over states and actions of half the summed difference between the model's and the true next-state probabilities.

Predict first. Rank the five settings by the model column alone. Which wins, and how does it score in the actual environment against the random policy's 210?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The score inside the model against the score in the world. Two bar charts of five sampling temperatures. Left, model scores from 2086 down to 732. Right, actual scores from 193 up to 1092 with the random policy at 210 and the best leaderboard entry at 820. Judged by the dream column the pick is temperature 0.10; the highlighted row is 0.10.
Setting to inspect (temperature): 0.10, Pick the winner by: The score inside the model
Constructed example (reported values in a constructed comparison): the scores are copied from the chapter's temperature table (reported by the cited paper for one game setup); the judging rule and the comparison are defined for this reader.

Calculated values

Temperature
0.10
Score inside the model
2086
Score in the actual environment
193
Model minus actual
1893
Actual minus random policy (210)
-17
Picked by the dream column
0.10
Picked: score in the actual environment
193
Picked: actual minus random policy (210)
-17
Reported spreads (model, actual)
2086 +/- 140, 193 +/- 58

At T = 0.10: model 2086 - actual 193 = 1893. Actual against the random policy: 193 - 210 = -17. Worst to best setting in the actual column: 1092 - 193 = 899 points. Judged only by the model column, the winner is T = 0.10 with 2086. In the actual environment it scores 193, and 193 - 210 = -17 puts it below the random policy: ranked first and confidently wrong. The chapter reports the authors' account of this row: at the lowest temperature the monsters fail to shoot at all (mode collapse), so a policy learned in that imagined world scores near-perfectly and then fails in the actual one. The reported spread 193 +/- 58 overlaps the random policy's 210 +/- 108.

Worked steps

  1. Read the chosen row: model 2086, actual 193.
  2. Model minus actual: 2086 - 193 = 1893.
  3. Actual minus the random policy's 210: 193 - 210 = -17.
  4. Rank the five settings by the dream column alone: the largest is 2086 at T = 0.10.
  5. That pick's actual score is 193: 193 - 210 = -17 against the random policy.
  6. Spreads were reported only for 0.10 and 1.15; the table does not give the transition error of Equation (14.1).

Use the idea

When a system proposes an action, predicts its outcome and scores itself on the prediction, it is reading the left column. Hold out some real outcomes and keep the second column, however small, and compare the two.

Where the conclusion applies

The scores are the chapter's reported numbers for one game setup and one training and evaluation procedure, each with its own variability (spreads were reported for 0.10 and 1.15 here). They show that the two evaluations can come apart, not a general temperature rule and not a measure of epsilon.

Common wrong turn: A higher score inside the model means a better agent
The chapter notes that an evaluation run only in the dream would have reported 2086 at the low setting, ranked it first, and been confidently wrong. Many agent evaluations have one column.
What this does not settle

The anchor is a preprint reporting on its authors' own systems in one game, and the temperature that worked there is not a value to copy.

Chapter 14 source: "What this does not settle".

Check your understanding: At T = 1.30, what is the model score minus the actual score, and which column is higher?
732 - 753 = -21, so the actual score is higher by 21. The columns can also run the other way.

Chapter 14 source: section "The knob the authors added, and what it bought". Demonstration C14-D01.

Demonstration 2 of 4

One step of error, compounded, against the actual error

How large can the value error become when a small one-step error compounds, and how large is it in a constructed model?

The bound multiplies gamma, epsilon and R, then divides by (1 - gamma) squared: one factor of 1/(1-gamma) counts the steps that matter and the second counts that a mistake at each step displaces everything after it. The finite bound grows with H(H-1) instead. Each bound holds under its own assumptions, and the actual error from the two recurrences stays under both. Raising gamma shrinks nothing in the finite bound but inflates the discounted one.

Equation (14.1), written in LaTeX: \epsilon = \max_{x,a} \tfrac{1}{2}\sum_{x'}|\hat P(x'\mid x,a) - P(x'\mid x,a)|.

Equation (14.2), written in LaTeX: |\hat V^{\pi}(x) - V^{\pi}(x)| \le \frac{\gamma \epsilon R}{(1-\gamma)^2}.

Scroll sideways for the whole equation

P(x' | x, a) is the true probability of moving to state x' from state x after action a; the hat marks the model's estimate. Epsilon is the largest row error, half the summed absolute difference. V(x) is the true value of state x under a fixed behaviour and V-hat(x) the value the model computes, gamma the discount factor and R the largest immediate reward. The actual error is the largest gap between the two value recurrences after H rewards, and the finite bound R x epsilon x H(H-1)/2 is the chapter's finite-horizon envelope. The value scale R/(1-gamma) is the most any run could collect.

Predict first. In the default case (error 0.02, R = 1) raise gamma to 0.99. About how large is the discounted bound?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: One step of error, compounded, against the actual error. Default case. Left, on a log axis, the actual value error and the finite bound against the horizon up to 5, with the discounted bound 1.80 as a dashed line. Right, the bound and the largest possible value against the discount factor, marked at gamma 0.90.
Case: Default: error 0.02, horizon 5, Discount factor gamma: 0.9
Constructed example: the laboratory's default, changed and transfer cases (kernels (0.90, 0.10), (0.20, 0.80) and model rows shifted by 0.02, horizons 5 and 20, identical kernels with rewards 1 and 2) and the chapter's 0.01 illustration, all computed with the laboratory's transition-model function.

Calculated values

Epsilon (Equation 14.1)
0.020
Actual error after H = 5
0.0868
Finite bound at H = 5
0.200
Discounted bound (Equation 14.2)
1.800
Tightest valid bound
0.200
Largest possible value R/(1-gamma)
10.0

Epsilon = 1/2 x (|0.88 - 0.90| + |0.12 - 0.10|) = 1/2 x (0.02 + 0.02) = 0.02. Finite bound = 1 x 0.02 x 5 x 4 / 2 = 0.200. Discounted bound = 0.90 x 0.02 x 1 / (1 - 0.90)^2 = 1.800. Check by hand after two rewards (state 2): true 1.00 + 0.90 x (0.20 x 0.00 + 0.80 x 1.00) = 1.720; model 1.00 + 0.90 x (0.22 x 0.00 + 0.78 x 1.00) = 1.702; |1.720 - 1.702| = 0.018. The two recurrences give 0.0868 after 5 (the left plot starts at H = 2 because H = 1 has error 0 and is not drawn). Each bound holds under its own assumptions; here the actual error 0.0868 is below both, 43 percent of the smaller one (finite bound 0.200).

Worked steps

  1. Row error by Equation (14.1): 1/2 x (|0.88 - 0.90| + |0.12 - 0.10|) = 1/2 x (0.02 + 0.02) = 0.02.
  2. Epsilon is the largest row error: 0.02. Rewards are shared; R = 1.
  3. Finite bound after H = 5: R x eps x H(H-1)/2 = 1 x 0.02 x 5 x 4 / 2 = 0.200.
  4. Discounted bound by Equation (14.2): gamma x eps x R / (1-gamma)^2 = 1.800.
  5. Each bound holds under its own assumptions; the smaller of the two numbers is 0.200.
  6. Two rewards by hand (state 2): V = 1.00 + 0.90 x (0.20 x 0.00 + 0.80 x 1.00) = 1.720, model 1.00 + 0.90 x (0.22 x 0.00 + 0.78 x 1.00) = 1.702, difference 0.018.
  7. Actual error after 5 rewards, from the two recurrences: 0.0868.

Use the idea

Before trusting a long imagined rollout, divide the bound by the largest possible value. If the share is near or above 1, the guarantee has nothing to say. If the model cannot be improved, consult the world more often.

Where the conclusion applies

The bound needs a fixed behaviour, the same expected immediate reward in model and world, rewards in [0, R], and error at most epsilon for every state and action that behaviour reaches. The kernels are the laboratory's constructed two-state defaults, so the actual error is that of one small model, not a general rate. A single number can also hide where the model is bad.

Common wrong turn: Reading the bound as the error the model actually makes
The chapter says the bound is worst case and often loose, and that a model with a large error may perform far better than it permits. At horizon 20 the finite bound is 3.8 while the actual error stays small.
What this does not settle

The bound in Equation (14.2) is worst case and often loose, and a model with a large error may perform far better than it permits.

Chapter 14 source: "What this does not settle".

Check your understanding: With gamma = 0.8, epsilon = 0.05 and R = 1, what is the discounted bound, and what is the largest possible value?
Bound = 0.8 x 0.05 / (1 - 0.8)^2 = 0.04 / 0.04 = 1.0. The largest possible value is 1 / 0.2 = 5, so the bound is a fifth of the range.

Chapter 14 source: section "What one step of error becomes". Demonstration C14-D02.

Demonstration 3 of 4

The optimizer finds where the model is generous, and the gap decides

A model can be wrong about every action value and still choose correctly, or be right about eleven tools and choose the twelfth. What decides which?

A shared offset moves every value together, so the ordering is untouched however large the offset is. One favorable error on a tool that is truly poor changes the choice only if it lifts that tool above the best, but when it does, the average error is tiny (the error divided by twelve) while the largest is the whole error: this is why Equation (14.1) is a maximum and not an average. An adversarial error marks the best down and the runner-up up, so the two meet when twice the error equals the gap.

Equation (14.3), written in LaTeX: \Delta(x) = \max_{a}Q^{\pi}(x,a) - \max_{a \neq a^{*}}Q^{\pi}(x,a).

Scroll sideways for the whole equation

Q(x, a) is the true value of taking action a at state x and then following the fixed behaviour; here the twelve options are tools with constructed values. Delta is the gap between the best action's value and the runner-up's (a* is a best action; if several actions tie for the maximum, the gap is defined as zero). The model's value of each action is off by the chosen error, and the chapter's ranking guarantee is that the model keeps the true ranking whenever twice that error is smaller than the gap: 2 delta < Delta. The largest error is the maximum over options, the average is the mean.

Predict first. Use the adversarial pattern with error 0.15, where twice the error equals the gap 0.30. What does the model do with tools 5 and 6?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The optimizer finds where the model is generous, and the gap decides. Left, true values of twelve tools as a navy line with a peak at tool 6 and model values as hollow diamonds; a star marks the optimizer's pick, tool 6. Right, the error of each tool; largest error 0.05, average 0.050. Pattern: the same offset on every tool, size 0.05.
Error pattern: Shared offset on every tool, Size of the error: 0.05
Constructed example: twelve option values defined for this reader (best 0.90, runner-up 0.60), with error patterns that illustrate the chapter's ranking argument and its eleven-right, one-wrong tool example.

Calculated values

True gap (Equation 14.3)
0.30
Largest error
0.050
Average error over 12 tools
0.050
Twice the largest error
0.10
Guarantee 2 x error < gap holds
yes
Model's pick
tool 6
Model's value of the pick
0.95
True value of the pick
0.90

Gap = 0.90 - 0.60 = 0.30. Twice the largest error = 2 x 0.05 = 0.10. Error pattern: the same offset on every tool. Average error = 0.60 / 12 = 0.050, largest error 0.05. Model values: tool 5 0.65, tool 6 0.95, tool 12 0.15. The model still picks tool 6, and the guarantee holds: 2 x 0.05 = 0.10 is below the gap 0.30.

Worked steps

  1. True values: tool 6 is best at 0.90, tool 5 is next at 0.60, tool 12 is lowest at 0.10.
  2. Gap by Equation (14.3): 0.90 - 0.60 = 0.30.
  3. Error pattern: the same offset on every tool, size 0.05.
  4. Model values: tool 5 0.65, tool 6 0.95, tool 12 0.15.
  5. Largest error 0.05; average error 0.60 / 12 = 0.050.
  6. Guarantee test: 2 x 0.05 = 0.10 against the gap 0.30: holds.
  7. The optimizer takes the highest model value: tool 6, true value 0.90.

Use the idea

Judge a planning model against the decision it serves, not against a single accuracy figure. Ask how large the gap is between the best option and the next one, compare twice the value error to it in the same units, and check whether the errors are concentrated on options the optimizer will rank.

Where the conclusion applies

Twelve options with constructed values, a uniform bound on the action-value error, and three hand-made error patterns. The chapter adds that this bound is an extra assumption: Equation (14.2) bounds state values and does not supply it. The test 2 x error < gap is sufficient, not necessary, so a failed test leaves the ranking undecided rather than wrong.

Common wrong turn: Judging a planning model by its average accuracy
The chapter says the useful question is never how accurate the model is, but whether its error is small relative to the gaps in the decision it serves, and that the diagnostic is whether errors are correlated across the options it ranks, not the average.
What this does not settle

This schematic explains the selection mechanism rather than supplying a measured error landscape.

Chapter 14 source: "This schematic explains the selection mechanism rather than supplying a measured error landscape.".

Check your understanding: True values 0.90 and 0.60, adversarial error 0.12. Which action does the model pick, and does the guarantee hold?
Model values: 0.90 - 0.12 = 0.78 and 0.60 + 0.12 = 0.72, so the model still picks tool 6. Twice the error is 0.24, below the gap 0.30, so the guarantee holds.

Chapter 14 source: section "The number that actually matters". Demonstration C14-D03.

Demonstration 4 of 4

How many imagined steps are certified

Given the model error, the reward range and the action gap, how many steps ahead can the certificate vouch for?

Twice the error bound grows with H times H-1, so each extra step costs more than the last. The certificate keeps the largest H for which that quantity is still strictly below the gap. Because H(H-1) grows like H squared, the horizon scales roughly with the square root of gap over R x epsilon: for long horizons, cutting the error by 8 times raises it by about 2.8 times, not 8. A twenty-step plan committed in advance fails the test by a wide margin, which is why the countermeasure is to plan a few steps, look, and replan.

Equation (14.4), written in LaTeX: H^{*} = \max\{ H \in \mathbb{Z}_{>0} : H \leq H_{\max}, R \epsilon H(H-1) < \Delta \}.

Scroll sideways for the whole equation

H is the number of rewards the agent plans over, H-max the declared planning cap (9 here), and H-star the largest H that passes. R is the largest immediate reward, epsilon the uniform one-step error from Equation (14.1), and Delta the lower bound on the action gap from Equation (14.3). The left side R x epsilon x H(H-1) is twice the finite-horizon action-value error bound R x epsilon x H(H-1)/2.

Predict first. With R = 10 and a gap of 2, cut the error by 8 times, from 0.08 to 0.01. Does the certified horizon grow by 8 times, or by far less?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How many imagined steps are certified. Points for horizons 1 to 9 of R x epsilon x H(H-1) against a horizontal action-gap line at 0.5; circles under the line are certified, crosses over it are not. The certified horizon is 5.
Reward range and action gap: R = 1, gap 0.5 (Figure 14.2), One-step error epsilon: 0.02
Constructed example: the chapter's Figure 14.2 gaps (R = 1, epsilon 0.02, gaps 0.2 and 0.5) and its priced twenty-step plan (R = 10, gap 2, epsilon 0.08 and 0.01); the other error values are defined for this reader, and each curve is checked against the laboratory's finite-horizon bound.

Calculated values

Certified horizon H*
5
First horizon that fails
6
R x epsilon x H(H-1) at H* = 5
0.40
Action gap
0.5
Same quantity at a twenty-step plan
7.6

At H = 5: 1 x 0.02 x 5 x 4 = 0.40, below the gap 0.5. At H = 6: 1 x 0.02 x 6 x 5 = 0.60, not below the gap 0.5. The certificate covers 5 rewards. Its failure at 6 is inconclusive: it does not show that the ranking reverses, only that this guarantee stops. A twenty-step plan committed in advance would need 1 x 0.02 x 20 x 19 = 7.6 to be below the gap 0.5, which it is not: plan 5, look, and replan.

Worked steps

  1. Inputs: R = 1, epsilon = 0.02, action gap = 0.5, planning cap 9.
  2. Test each H: R x epsilon x H(H-1) must be strictly below 0.5.
  3. Largest H that passes: H* = 5, because 1 x 0.02 x 5 x 4 = 0.40.
  4. First failure at H = 6: 1 x 0.02 x 6 x 5 = 0.60.
  5. Twenty steps: 1 x 0.02 x 20 x 19 = 7.6, far above the gap.

Use the idea

Use the certified horizon to set how many steps an agent commits to before it observes the world again. Plan that many, execute a few, look, and replan, rather than emitting a long plan against an error nobody measured.

Where the conclusion applies

A finite covered domain, equal expected immediate rewards in model and world, the same initial state and fixed behaviour, uniform error epsilon, zero terminal continuation and a declared gap. The R = 10 case uses the chapter's priced plan, where a logged prediction miss rate is declared a conservative proxy for epsilon, which is not the same thing. A failed test is inconclusive, not proof of a wrong decision, and the inputs are usually not measured.

Common wrong turn: A failed certificate proves the ranking reverses
The chapter says the certificate's failure at six is inconclusive, not evidence that the model must make a wrong decision there. It only says this sufficient guarantee stops.
What this does not settle

The horizon arithmetic is constructed for teaching. How to bound what a system could achieve given a model, rather than what it does achieve, waits for Chapters 24 and 26.

Chapter 14 source: "What this does not settle".

Check your understanding: With R = 1, epsilon = 0.04 and a gap of 0.5, what is the certified horizon?
R x epsilon x H(H-1) = 0.04 x H(H-1). At H = 4 it is 0.04 x 12 = 0.48, below 0.5. At H = 5 it is 0.04 x 20 = 0.80, which is not. H-star = 4.

Chapter 14 source: section "How many steps a model is good for". Demonstration C14-D04.