Demonstration 1 of 4
One run, three update rules, pass by pass
When a run pays only at its last step, which state estimates move, and how far does the news travel on each pass under each rule?
The outcome rule waits for the end and then moves every visited estimate toward the return that followed it. The one-step rule uses the reward plus the next estimate as its target. With every estimate at 0.5 and zero reward before the last step, the first three errors are exactly zero, so only the last state moves on the first pass; on later passes the raised estimate reaches one state further back each time. The trace rule sits between the two. In the truncated run the final estimate is supplied, so it anchors the targets instead of the terminal value 0.
Scroll sideways for the whole equation
V-hat is the current estimate of a state's future return. Ret is the return that actually followed the state. alpha is the step size, the fraction of the gap that is kept. r is the immediate reward, gamma is the discount factor (1 except in the a, b, c, d run, where it is 0.9), x_t is the state visited at step t, and delta is the one-step error: the reward plus the next state's estimate minus the state's own estimate. The end of a terminated run has value 0; a truncated run keeps a supplied estimate. The trace rule uses lambda to carry later errors back to earlier states. A pass is one sequential run through the same trajectory.
Predict first. In the four-action run (step size 0.1, estimates 0.5, final reward 1), after the first pass, how many of the four estimates have moved under the outcome rule, and how many under the one-step rule?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's four-action trajectory (estimates 0.5, step size 0.1, final reward 1) and the laboratory's default, changed and transfer cases (draft, review, done; truncated; a, b, c, d), computed with the laboratory's trajectory-credit function and repeated pass by pass.
Calculated values
- Outcome rule, states in order
- 0.550, 0.550, 0.550, 0.550
- One-step rule, states in order
- 0.500, 0.500, 0.500, 0.550
- States moved (outcome / one-step)
- 4 / 1
Pass 1 of the four-action run. Outcome rule: state 1: return 1, 0.5 + 0.1 x (1 - 0.5) = 0.55. One-step rule: state 1: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5; state 2: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5; state 3: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5; state 4: delta = 1 + 1 x 0 - 0.5 = 0.5, so 0.5 + 0.1 x 0.5 = 0.55. After this pass the outcome rule has moved 4 of 4 estimates and the one-step rule 1. Neither movement says which action deserved the result: each is a repair of a prediction of future return.
Worked steps
- Outcome rule, state 1: return 1, 0.5 + 0.1 x (1 - 0.5) = 0.55.
- Outcome rule, state 2: return 1, 0.5 + 0.1 x (1 - 0.5) = 0.55.
- Outcome rule, state 3: return 1, 0.5 + 0.1 x (1 - 0.5) = 0.55.
- One-step rule, state 1: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5.
- One-step rule, state 2: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5.
- One-step rule, state 3: delta = 0 + 1 x 0.5 - 0.5 = 0, so 0.5 + 0.1 x 0 = 0.5.
- One-step rule, state 4: delta = 1 + 1 x 0 - 0.5 = 0.5, so 0.5 + 0.1 x 0.5 = 0.55.
Use the idea
When a long run ends in success or failure, an update that touches every visited state is a statement about the returns that followed those states under the policy that was run. It is not an itemised receipt for each action, and the one-step rule's zero errors do not say the early steps were unimportant.
Where the conclusion applies
States visited once each in a fixed order, one sequential pass repeated on the same trajectory, and values and step sizes declared for each run (the notebook's draft, review, done run, its truncated variant and its a, b, c, d run, and the book's four-action run). Different starting estimates, discounts or replay orders change which states move. Repeating one trajectory is a teaching device, not a training set.
Common wrong turn: Equal updates mean equal credit
What this does not settle
The convergence result is for the linear one-step rule with linearly independent state representations under repeated presentation, and arbitrary language-model agent representations and training procedures need not satisfy those conditions.
Chapter 12 source: "What this does not settle".
Check your understanding: In the draft, review, done run (step size 0.5, lambda 0.8, terminal) after one pass, what does the outcome rule give draft, and what does the one-step rule give draft?
Chapter 12 source: section "The two rules on one trajectory". Demonstration C12-D01.
Demonstration 2 of 4
Eight episodes, two answers about state A
Why do the two rules agree about state B but disagree about state A, which was seen only once?
B is visited eight times, so both rules average its eight returns. A is visited once, followed by B with no reward. The outcome rule copies that single return. The one-step rule says A is worth whatever B is worth, and so it borrows the evidence from the seven episodes that never contained A.
Scroll sideways for the whole equation
A and B are the two states. Each episode ends with a reward when it reaches the end. The outcome rule sets a state to the average return that followed its visits. The one-step rule sets a state to the reward plus the estimate of the state that followed it. No discounting (gamma is 1). x_t is the state at step t and V-hat its estimate. Values are what each rule settles on when it is run in batch form, repeatedly over the same eight episodes with a small step size.
Predict first. With six of the seven B-only episodes paying 1 and the A episode ending at 0, what does each rule say A is worth?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's eight-episode construction (six ones among the seven B-only episodes, the A episode ending at 0), with those counts varied.
Calculated values
- Value of B (both rules)
- 0.750
- Value of A, outcome rule
- 0.000
- Value of A, one-step rule
- 0.750
- Disagreement about A
- 0.750
B's returns are 6 + 0 ones in 8 visits, so B = (6 + 0) / 8 = 0.750 under either rule. The outcome rule gives A the one return it saw: 0. The one-step rule gives A = 0 + B = 0.750. Disagreement about A = |0.750 - 0.000| = 0.750. The outcome rule fits the single A episode exactly (error 0) but is 0.750 away from B's value; the one-step rule gives A the value of its successor B, using the seven episodes that never contained A.
Worked steps
- B is visited 8 times; its returns hold 6 + 0 ones, so B = (6 + 0) / 8 = 0.750.
- A is visited once and the return that followed was 0, so the outcome rule sets A = 0.
- Every A transition goes to B with reward 0, so the one-step rule sets A = 0 + B = 0.750.
- Disagreement = |0.750 - 0.000| = 0.750.
Use the idea
When a state is rare but its successor is common, ask whether the stated process really sends the rare state to that successor. If so, the successor-based answer pools more evidence; if not, it imports a mistake.
Where the conclusion applies
The process is stipulated: A always goes to B with no reward, and B pays at random. Eight observations do not prove that. The one-step answer is better only if that process is real and the successor's value is reliable. The outcome rule is not wrong about the data; it fits them exactly.
Common wrong turn: The rule that fits the data perfectly is the right one
Check your understanding: Suppose three of the seven B-only episodes pay 1 and the A episode ends at 0. What is B, and how far is the outcome rule's value for A from the one-step rule's?
Chapter 12 source: section "Eight episodes, two answers". Demonstration C12-D02.
Demonstration 3 of 4
One dial between the two rules
How does the blending parameter lambda divide weight between short lookahead targets and the complete return, and what does it do to prediction on a random walk?
Each extra step of lookahead is weighted lambda times the one before it, so small lambda piles weight on the one-step target and large lambda spreads it out. Whatever weight the lookahead targets leave over goes to the complete return, so the two opening rules are the two ends of the dial. The right panel scores the dial on the chapter's random walk. When each training set is presented repeatedly, the outcome rule (lambda 1) has the lowest error on the training sample and the highest error against the truth. Presented once, the best lambda is small but is not clearly the one-step end, and the chapter attributes that to the one-step rule being slow at propagating information back.
Scroll sideways for the whole equation
lambda is the blending parameter between 0 and 1. A k-step target adds the next k rewards to the estimate k steps ahead. N is the number of transitions in the whole episode and t the current step, so n = N - t is the number of transitions left (8 here, as in Figure 12.4); j counts the rewards inside a target, and gamma is the discount factor (the weights do not depend on it). With n transitions left, the first n - 1 targets get weight (1 - lambda) x lambda^(k-1) and the complete return gets the remaining weight, lambda^(n-1). The weights always add to 1. It is not the step size and not the discount factor. In the right panel, the true value of each interior state B to F of the seven-state random walk is k/6 for k = 1 to 5, and the error is the root mean square miss over those five states.
Predict first. With 8 transitions remaining and lambda = 0.9, does the complete return or the one-step target get more weight?
Choose an example
Scroll sideways for the whole figure
Constructed example: the weights of Equation (12.4) for the book's eight-transition case (Figure 12.4) and the chapter's random-walk chain (Figure 12.2), with seeded walks and lambda values defined for this reader.
Calculated values
- Weight on the one-step target
- 0.1000
- Weight on the full return
- 0.4783
- Weight on the other lookahead targets
- 0.4217
- Weights add to
- 1.0000
- Random-walk error at this lambda
- 0.145
- Random-walk error at lambda 0 and 1
- 0.130 and 0.173
- Training-sample error at lambda 0 and 1
- 0.427 and 0.407
First weight = 1 - 0.9 = 0.1000. Full-return weight = 0.9^7 = 0.4783. Check: lookahead weights add to 1 - 0.9^7 = 0.5217, and 0.5217 + 0.4783 = 1.0000. Raising lambda moves weight from the short targets toward the complete return, but the weights always add to one. In the constructed random walk, each of 100 seeded sets of 10 walks was presented repeatedly until the weights stopped changing. The error against the true values k/6 is 0.130 at lambda 0 and 0.173 at lambda 1 (standard error about 0.009), while the error on the training sample runs the other way, 0.427 to 0.407: the outcome rule fits the sample best and the truth worst.
Worked steps
- First weight = 1 - lambda = 1 - 0.9 = 0.1000.
- Weight of the k-step target = (1 - lambda) x lambda^(k-1), for k = 1 to 7.
- Full-return weight = lambda^7 = 0.9^7 = 0.4783.
- Lookahead weights add to 1 - 0.4783 = 0.5217.
- Total = 0.5217 + 0.4783 = 1.0000.
Use the idea
When a method is described as temporal-difference (TD) learning with a lambda setting, read it as a choice of target mixture: how much to trust the agent's own estimates relative to what actually happened. For a fixed policy and a prediction target, it changes the estimator, not the objective.
Where the conclusion applies
The weights are for a single episode that terminates after n transitions, with terminal value zero. The right panel is a constructed experiment: 100 training sets of 10 walks drawn with a fixed seed, a table of five values, and, for the single presentation, a step size chosen from a grid at each lambda. It reproduces the shape of the chapter's experiment, not Sutton's published numbers, and it does not say which lambda is best for any other problem.
Common wrong turn: Lambda is the discount factor or the learning rate
What this does not settle
The random walk is five numbers; the results are a demonstration rather than a general performance claim.
Chapter 12 source: "What this does not settle".
Check your understanding: With 8 transitions remaining and lambda = 0.3, what weight does the first target get, and what does the complete return get?
Chapter 12 source: section "The family between the two rules". Demonstration C12-D03.
Demonstration 4 of 4
A positive error for a useless search, and a value that switches the action
Can the one-step error be positive after a search that cost something and returned nothing useful, and when does an updated value change the next action?
The error compares the old forecast with the observed reward plus the new forecast. A cost of 0.02 makes the reward slightly negative, but the sign is set by the two estimates. The right panel adds the bridge to action: training moves the continuation estimate of y from 0.5 toward 1, so inspect rises from 0.4 to -0.1 plus the new estimate, and it overtakes finish (0.6) once that estimate passes 0.7, which with these numbers needs a step size above 0.4. The update needed visits to y; a controller that always finishes never reaches y.
Scroll sideways for the whole equation
r is the reward for the call, here a cost of 0.02 (reward -0.02). The discount factor gamma is 1. The prior estimate is the agent's forecast of future return before the search (0.20); the successor estimate is its forecast after the search. delta is the reward plus the successor estimate minus the prior estimate; it is the temporal-difference (TD) error. Right panel: at state x, finish ends with reward 0.6, and inspect costs 0.1 and reaches state y, where a fixed continuation policy takes the final action. The estimate of y starts at 0.5 and is trained with step size alpha on a run that reaches y and pays 1.
Predict first. With the prior estimate at 0.20, which successor estimates give a positive error: 0.10, 0.50 or 0.80?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's first search (reward -0.02, estimates 0.20 before and 0.50 or 0.10 after) and its finish-or-inspect example (finish 0.6, inspect cost 0.1, estimate 0.5, step size 0.8); other successor estimates and step sizes are defined for this reader; both updates are checked against the laboratory's trajectory-credit function.
Calculated values
- TD error delta
- 0.28
- Sign
- positive
- Successor estimate that gives zero
- 0.22
- Continuation estimate after training
- 0.900
- Inspect after training
- 0.800
- Action chosen after training
- inspect
delta = (-0.02) + 0.50 - 0.20 = 0.28. The error is zero at a successor estimate of 0.20 - (-0.02) = 0.22; above that it is positive, below it negative. The book's first search was unhelpful and cost 0.02, yet the error is positive, because the estimate after it is higher than the estimate before it by more than that cost, so the forecast for the earlier state is raised. Training on a run that reaches y and pays 1: delta = 1 + 0 - 0.5 = 0.5, so the continuation estimate becomes 0.5 + 0.8 x 0.5 = 0.900. Lookahead: inspect now scores (-0.1) + 0.900 = 0.800, above finish's 0.6, so the chosen action switches to inspect. Before training, inspect scored -0.1 + 0.5 = 0.4 against 0.6, so the controller finished; this update repairs predictions only and does not say which action caused any outcome.
Worked steps
- Search error: delta = (-0.02) + 0.50 - 0.20 = 0.28 (positive).
- Zero error at a successor estimate of 0.20 + 0.02 = 0.22.
- Before training: inspect = (-0.1) + 0.5 = 0.4 against finish 0.6, so finish.
- Training run reaches y and pays 1: delta = 1 + 0 - 0.5 = 0.5.
- Update: 0.5 + 0.8 x 0.5 = 0.900.
- After training: inspect = (-0.1) + 0.900 = 0.800 against finish 0.6: inspect.
Use the idea
When reading a log of TD errors from an agent, do not treat a positive error as praise for the call just made, or a negative one as blame. Credit for a particular call needs a comparison with what would have happened without it, which these updates do not provide.
Where the conclusion applies
One transition with the reward and estimates set by hand, as in the chapter's example, and a model at x that is known while the continuation value is learned. Estimates here are declared numbers, not fitted values; one reward of 1 may overstate the continuation policy's mean, so the switch is an improvement in the estimated comparison, and its real return still needs evaluation.
Common wrong turn: A positive error praises the call that was just made
What this does not settle
Chapter 14 takes up what happens when the dynamics are learned explicitly rather than implied.
Chapter 12 source: "What this does not settle".
Check your understanding: At step size 0.5, what is the estimate of y after training, and which action does the controller now choose?
Chapter 12 source: section "When a learned value changes the next action". Demonstration C12-D04.