The Mathematics of AI Agents, laboratory reader ยท Chapter 12

Credit for Consequences

Predictions of future return can be repaired from the final outcome or from the next prediction. Neither repair says which action caused the outcome.

These four demonstrations follow the chapter's progression: update rules pass by pass on four runs, the eight-episode construction on which they disagree, the dial that blends them beside the random-walk experiment, and a tool-call example where the sign of the error depends on predictions rather than on blame, and a learned value can change the next action.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

One run, three update rules, pass by pass

When a run pays only at its last step, which state estimates move, and how far does the news travel on each pass under each rule?

The outcome rule waits for the end and then moves every visited estimate toward the return that followed it. The one-step rule uses the reward plus the next estimate as its target. With every estimate at 0.5 and zero reward before the last step, the first three errors are exactly zero, so only the last state moves on the first pass; on later passes the raised estimate reaches one state further back each time. The trace rule sits between the two. In the truncated run the final estimate is supplied, so it anchors the targets instead of the terminal value 0.

Equation (12.1), written in LaTeX: \hat V_{t+1}(x_t) = \hat V_t(x_t) + \alpha[\operatorname{Ret}_t - \hat V_t(x_t)].

Equation (12.2), written in LaTeX: \delta_t = r_t + \gamma \hat V_t(x_{t+1}) - \hat V_t(x_t).

Equation (12.3), written in LaTeX: \hat V_{t+1}(x_t) = \hat V_t(x_t) + \alpha \delta_t.

Scroll sideways for the whole equation

V-hat is the current estimate of a state's future return. Ret is the return that actually followed the state. alpha is the step size, the fraction of the gap that is kept. r is the immediate reward, gamma is the discount factor (1 except in the a, b, c, d run, where it is 0.9), x_t is the state visited at step t, and delta is the one-step error: the reward plus the next state's estimate minus the state's own estimate. The end of a terminated run has value 0; a truncated run keeps a supplied estimate. The trace rule uses lambda to carry later errors back to earlier states. A pass is one sequential run through the same trajectory.

Predict first. In the four-action run (step size 0.1, estimates 0.5, final reward 1), after the first pass, how many of the four estimates have moved under the outcome rule, and how many under the one-step rule?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: One run, three update rules, pass by pass. Grouped bars for each state of the four-action run after 1 pass: the change in the estimate under the outcome rule and one-step rule. The outcome rule moved 4 of 4 states and the one-step rule 1.
Run: Book: four actions, final reward 1, Passes through the run: 1
Constructed example: the book's four-action trajectory (estimates 0.5, step size 0.1, final reward 1) and the laboratory's default, changed and transfer cases (draft, review, done; truncated; a, b, c, d), computed with the laboratory's trajectory-credit function and repeated pass by pass.

Calculated values

Outcome rule, states in order
0.550, 0.550, 0.550, 0.550
One-step rule, states in order
0.500, 0.500, 0.500, 0.550
States moved (outcome / one-step)
4 / 1

Pass 1 of the four-action run. Outcome rule: state 1: return 1, 0.5 + 0.1 x (1 - 0.5) = 0.55. One-step rule: state 1: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5; state 2: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5; state 3: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5; state 4: delta = 1 + 1 x 0 - 0.5 = 0.5, so 0.5 + 0.1 x 0.5 = 0.55. After this pass the outcome rule has moved 4 of 4 estimates and the one-step rule 1. Neither movement says which action deserved the result: each is a repair of a prediction of future return.

Worked steps

  1. Outcome rule, state 1: return 1, 0.5 + 0.1 x (1 - 0.5) = 0.55.
  2. Outcome rule, state 2: return 1, 0.5 + 0.1 x (1 - 0.5) = 0.55.
  3. Outcome rule, state 3: return 1, 0.5 + 0.1 x (1 - 0.5) = 0.55.
  4. One-step rule, state 1: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5.
  5. One-step rule, state 2: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5.
  6. One-step rule, state 3: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5.
  7. One-step rule, state 4: delta = 1 + 1 x 0 - 0.5 = 0.5, so 0.5 + 0.1 x 0.5 = 0.55.

Use the idea

When a long run ends in success or failure, an update that touches every visited state is a statement about the returns that followed those states under the policy that was run. It is not an itemised receipt for each action, and the one-step rule's zero errors do not say the early steps were unimportant.

Where the conclusion applies

States visited once each in a fixed order, one sequential pass repeated on the same trajectory, and values and step sizes declared for each run (the notebook's draft, review, done run, its truncated variant and its a, b, c, d run, and the book's four-action run). Different starting estimates, discounts or replay orders change which states move. Repeating one trajectory is a teaching device, not a training set.

Common wrong turn: Equal updates mean equal credit
The chapter says the four identical updates of the outcome rule are predictions of subsequent return, not allocations of causal responsibility, and that the first three TD errors vanish by arithmetic without certifying that those transitions were unimportant.
What this does not settle

The convergence result is for the linear one-step rule with linearly independent state representations under repeated presentation, and arbitrary language-model agent representations and training procedures need not satisfy those conditions.

Chapter 12 source: "What this does not settle".

Check your understanding: In the draft, review, done run (step size 0.5, lambda 0.8, terminal) after one pass, what does the outcome rule give draft, and what does the one-step rule give draft?
The return from draft is 0 + 1 = 1, so the outcome rule gives 0 + 0.5 x (1 - 0) = 0.5. The one-step error at draft is 0 + 0 - 0 = 0 (review is still 0 when draft is updated), so draft stays at 0.

Chapter 12 source: section "The two rules on one trajectory". Demonstration C12-D01.

Demonstration 2 of 4

Eight episodes, two answers about state A

Why do the two rules agree about state B but disagree about state A, which was seen only once?

B is visited eight times, so both rules average its eight returns. A is visited once, followed by B with no reward. The outcome rule copies that single return. The one-step rule says A is worth whatever B is worth, and so it borrows the evidence from the seven episodes that never contained A.

Equation (12.1), written in LaTeX: \hat V_{t+1}(x_t) = \hat V_t(x_t) + \alpha[\operatorname{Ret}_t - \hat V_t(x_t)].

Equation (12.2), written in LaTeX: \delta_t = r_t + \gamma \hat V_t(x_{t+1}) - \hat V_t(x_t).

Equation (12.3), written in LaTeX: \hat V_{t+1}(x_t) = \hat V_t(x_t) + \alpha \delta_t.

Scroll sideways for the whole equation

A and B are the two states. Each episode ends with a reward when it reaches the end. The outcome rule sets a state to the average return that followed its visits. The one-step rule sets a state to the reward plus the estimate of the state that followed it. No discounting (gamma is 1). x_t is the state at step t and V-hat its estimate. Values are what each rule settles on when it is run in batch form, repeatedly over the same eight episodes with a small step size.

Predict first. With six of the seven B-only episodes paying 1 and the A episode ending at 0, what does each rule say A is worth?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Eight episodes, two answers about state A. Two groups of bars. For state A, seen once, the outcome rule gives 0.00 and the one-step rule 0.75. For state B, seen eight times, both rules give 0.75.
B-only episodes that pay 1 (out of 7): 6, Reward at the end of the one A episode: 0
Constructed example: the book's eight-episode construction (six ones among the seven B-only episodes, the A episode ending at 0), with those counts varied.

Calculated values

Value of B (both rules)
0.750
Value of A, outcome rule
0.000
Value of A, one-step rule
0.750
Disagreement about A
0.750

B's returns are 6 + 0 ones in 8 visits, so B = (6 + 0) / 8 = 0.750 under either rule. The outcome rule gives A the one return it saw: 0. The one-step rule gives A = 0 + B = 0.750. Disagreement about A = |0.750 - 0.000| = 0.750. The outcome rule fits the single A episode exactly (error 0) but is 0.750 away from B's value; the one-step rule gives A the value of its successor B, using the seven episodes that never contained A.

Worked steps

  1. B is visited 8 times; its returns hold 6 + 0 ones, so B = (6 + 0) / 8 = 0.750.
  2. A is visited once and the return that followed was 0, so the outcome rule sets A = 0.
  3. Every A transition goes to B with reward 0, so the one-step rule sets A = 0 + B = 0.750.
  4. Disagreement = |0.750 - 0.000| = 0.750.

Use the idea

When a state is rare but its successor is common, ask whether the stated process really sends the rare state to that successor. If so, the successor-based answer pools more evidence; if not, it imports a mistake.

Where the conclusion applies

The process is stipulated: A always goes to B with no reward, and B pays at random. Eight observations do not prove that. The one-step answer is better only if that process is real and the successor's value is reliable. The outcome rule is not wrong about the data; it fits them exactly.

Common wrong turn: The rule that fits the data perfectly is the right one
The chapter says the outcome rule is not wrong about the data: it fits the single A episode with zero error, which is exactly what its guarantee promises. The guarantee names a criterion, fit to a sample, and what anybody wanted is accuracy on cases not yet seen.
Check your understanding: Suppose three of the seven B-only episodes pay 1 and the A episode ends at 0. What is B, and how far is the outcome rule's value for A from the one-step rule's?
B = (3 + 0) / 8 = 0.375. The outcome rule gives A = 0, the one-step rule gives A = 0.375, so they differ by 0.375.

Chapter 12 source: section "Eight episodes, two answers". Demonstration C12-D02.

Demonstration 3 of 4

One dial between the two rules

How does the blending parameter lambda divide weight between short lookahead targets and the complete return, and what does it do to prediction on a random walk?

Each extra step of lookahead is weighted lambda times the one before it, so small lambda piles weight on the one-step target and large lambda spreads it out. Whatever weight the lookahead targets leave over goes to the complete return, so the two opening rules are the two ends of the dial. The right panel scores the dial on the chapter's random walk. When each training set is presented repeatedly, the outcome rule (lambda 1) has the lowest error on the training sample and the highest error against the truth. Presented once, the best lambda is small but is not clearly the one-step end, and the chapter attributes that to the one-step rule being slow at propagating information back.

Equation (12.4), written in LaTeX: \operatorname{Ret}^{\lambda}_t = (1-\lambda)\sum_{k=1}^{N-t-1}\lambda^{ k-1}[\textstyle\sum_{j=0}^{k-1}\gamma^{ j} r_{t+j} + \gamma^{ k}\hat V_t(x_{t+k})] + \lambda^{N-t-1}\operatorname{Ret}_t.

Scroll sideways for the whole equation

lambda is the blending parameter between 0 and 1. A k-step target adds the next k rewards to the estimate k steps ahead. N is the number of transitions in the whole episode and t the current step, so n = N - t is the number of transitions left (8 here, as in Figure 12.4); j counts the rewards inside a target, and gamma is the discount factor (the weights do not depend on it). With n transitions left, the first n - 1 targets get weight (1 - lambda) x lambda^(k-1) and the complete return gets the remaining weight, lambda^(n-1). The weights always add to 1. It is not the step size and not the discount factor. In the right panel, the true value of each interior state B to F of the seven-state random walk is k/6 for k = 1 to 5, and the error is the root mean square miss over those five states.

Predict first. With 8 transitions remaining and lambda = 0.9, does the complete return or the one-step target get more weight?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: One dial between the two rules. Left: eight bars of weight for the k-step targets at lambda 0.9; the first is 0.100 and the last, the full return, is 0.478. Right: root mean squared error of the random-walk predictions at seven lambda values, from 0.130 at lambda 0 to 0.173 at lambda 1; the selected lambda is circled at 0.145.
Blending parameter lambda: 0.9, How each training set is presented: Repeatedly, until the weights stop changing
Constructed example: the weights of Equation (12.4) for the book's eight-transition case (Figure 12.4) and the chapter's random-walk chain (Figure 12.2), with seeded walks and lambda values defined for this reader.

Calculated values

Weight on the one-step target
0.1000
Weight on the full return
0.4783
Weight on the other lookahead targets
0.4217
Weights add to
1.0000
Random-walk error at this lambda
0.145
Random-walk error at lambda 0 and 1
0.130 and 0.173
Training-sample error at lambda 0 and 1
0.427 and 0.407

First weight = 1 - 0.9 = 0.1000. Full-return weight = 0.9^7 = 0.4783. Check: lookahead weights add to 1 - 0.9^7 = 0.5217, and 0.5217 + 0.4783 = 1.0000. Raising lambda moves weight from the short targets toward the complete return, but the weights always add to one. In the constructed random walk, each of 100 seeded sets of 10 walks was presented repeatedly until the weights stopped changing. The error against the true values k/6 is 0.130 at lambda 0 and 0.173 at lambda 1 (standard error about 0.009), while the error on the training sample runs the other way, 0.427 to 0.407: the outcome rule fits the sample best and the truth worst.

Worked steps

  1. First weight = 1 - lambda = 1 - 0.9 = 0.1000.
  2. Weight of the k-step target = (1 - lambda) x lambda^(k-1), for k = 1 to 7.
  3. Full-return weight = lambda^7 = 0.9^7 = 0.4783.
  4. Lookahead weights add to 1 - 0.4783 = 0.5217.
  5. Total = 0.5217 + 0.4783 = 1.0000.

Use the idea

When a method is described as temporal-difference (TD) learning with a lambda setting, read it as a choice of target mixture: how much to trust the agent's own estimates relative to what actually happened. For a fixed policy and a prediction target, it changes the estimator, not the objective.

Where the conclusion applies

The weights are for a single episode that terminates after n transitions, with terminal value zero. The right panel is a constructed experiment: 100 training sets of 10 walks drawn with a fixed seed, a table of five values, and, for the single presentation, a step size chosen from a grid at each lambda. It reproduces the shape of the chapter's experiment, not Sutton's published numbers, and it does not say which lambda is best for any other problem.

Common wrong turn: Lambda is the discount factor or the learning rate
The chapter says lambda is a statement about the estimator, how much the agent trusts its own estimates relative to what happened, while the discount factor is a statement about the objective. It is also not the step size, which controls how far each update moves, and not exploration.
What this does not settle

The random walk is five numbers; the results are a demonstration rather than a general performance claim.

Chapter 12 source: "What this does not settle".

Check your understanding: With 8 transitions remaining and lambda = 0.3, what weight does the first target get, and what does the complete return get?
First target: 1 - 0.3 = 0.7. Complete return: 0.3^7 = 0.0002187, which the figure labels as less than 0.001. The weights still add to 1.

Chapter 12 source: section "The family between the two rules". Demonstration C12-D03.

Demonstration 4 of 4

A positive error for a useless search, and a value that switches the action

Can the one-step error be positive after a search that cost something and returned nothing useful, and when does an updated value change the next action?

The error compares the old forecast with the observed reward plus the new forecast. A cost of 0.02 makes the reward slightly negative, but the sign is set by the two estimates. The right panel adds the bridge to action: training moves the continuation estimate of y from 0.5 toward 1, so inspect rises from 0.4 to -0.1 plus the new estimate, and it overtakes finish (0.6) once that estimate passes 0.7, which with these numbers needs a step size above 0.4. The update needed visits to y; a controller that always finishes never reaches y.

Equation (12.2), written in LaTeX: \delta_t = r_t + \gamma \hat V_t(x_{t+1}) - \hat V_t(x_t).

Equation (12.3), written in LaTeX: \hat V_{t+1}(x_t) = \hat V_t(x_t) + \alpha \delta_t.

Scroll sideways for the whole equation

r is the reward for the call, here a cost of 0.02 (reward -0.02). The discount factor gamma is 1. The prior estimate is the agent's forecast of future return before the search (0.20); the successor estimate is its forecast after the search. delta is the reward plus the successor estimate minus the prior estimate; it is the temporal-difference (TD) error. Right panel: at state x, finish ends with reward 0.6, and inspect costs 0.1 and reaches state y, where a fixed continuation policy takes the final action. The estimate of y starts at 0.5 and is trained with step size alpha on a run that reaches y and pays 1.

Predict first. With the prior estimate at 0.20, which successor estimates give a positive error: 0.10, 0.50 or 0.80?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A positive error for a useless search, and a value that switches the action. Left: four bars for the first search, reward -0.02, successor estimate 0.50, minus own estimate -0.20, and the error 0.28. Right: three bars, finish 0.6, inspect before training 0.4 and inspect after training 0.80; the action chosen after training is inspect.
Estimate after the search: 0.5, Step size for training the estimate of y: 0.8
Constructed example: the book's first search (reward -0.02, estimates 0.20 before and 0.50 or 0.10 after) and its finish-or-inspect example (finish 0.6, inspect cost 0.1, estimate 0.5, step size 0.8); other successor estimates and step sizes are defined for this reader; both updates are checked against the laboratory's trajectory-credit function.

Calculated values

TD error delta
0.28
Sign
positive
Successor estimate that gives zero
0.22
Continuation estimate after training
0.900
Inspect after training
0.800
Action chosen after training
inspect

delta = (-0.02) + 0.50 - 0.20 = 0.28. The error is zero at a successor estimate of 0.20 - (-0.02) = 0.22; above that it is positive, below it negative. The book's first search was unhelpful and cost 0.02, yet the error is positive, because the estimate after it is higher than the estimate before it by more than that cost, so the forecast for the earlier state is raised. Training on a run that reaches y and pays 1: delta = 1 + 0 - 0.5 = 0.5, so the continuation estimate becomes 0.5 + 0.8 x 0.5 = 0.900. Lookahead: inspect now scores (-0.1) + 0.900 = 0.800, above finish's 0.6, so the chosen action switches to inspect. Before training, inspect scored -0.1 + 0.5 = 0.4 against 0.6, so the controller finished; this update repairs predictions only and does not say which action caused any outcome.

Worked steps

  1. Search error: delta = (-0.02) + 0.50 - 0.20 = 0.28 (positive).
  2. Zero error at a successor estimate of 0.20 + 0.02 = 0.22.
  3. Before training: inspect = (-0.1) + 0.5 = 0.4 against finish 0.6, so finish.
  4. Training run reaches y and pays 1: delta = 1 + 0 - 0.5 = 0.5.
  5. Update: 0.5 + 0.8 x 0.5 = 0.900.
  6. After training: inspect = (-0.1) + 0.900 = 0.800 against finish 0.6: inspect.

Use the idea

When reading a log of TD errors from an agent, do not treat a positive error as praise for the call just made, or a negative one as blame. Credit for a particular call needs a comparison with what would have happened without it, which these updates do not provide.

Where the conclusion applies

One transition with the reward and estimates set by hand, as in the chapter's example, and a model at x that is known while the continuation value is learned. Estimates here are declared numbers, not fitted values; one reward of 1 may overstate the continuation policy's mean, so the switch is an improvement in the estimated comparison, and its real return still needs evaluation.

Common wrong turn: A positive error praises the call that was just made
The chapter says the TD error is not a verdict on the call: with a reward of -0.02, a prior estimate of 0.20 and a successor estimate of 0.50 the error is 0.28, positive despite the search's cost and lack of useful results, and an optimistic estimate can fall after a useful action.
What this does not settle

Chapter 14 takes up what happens when the dynamics are learned explicitly rather than implied.

Chapter 12 source: "What this does not settle".

Check your understanding: At step size 0.5, what is the estimate of y after training, and which action does the controller now choose?
delta = 1 + 0 - 0.5 = 0.5, so the estimate becomes 0.5 + 0.5 x 0.5 = 0.75. Inspect scores -0.1 + 0.75 = 0.65, above finish's 0.6, so the controller now inspects.

Chapter 12 source: section "When a learned value changes the next action". Demonstration C12-D04.