Chapter 12: separate solutions

Question 1

Compute default trace update to draft.

Updating draft uses learning rate 0.5, a temporal-difference error of 1 at review (reward 1, no bootstrap) and trace weight discount times lambda, 1(0.8)=0.8: 0.5*1*0.8=0.4.

Question 2

What are changed return targets?

[3,3], because final bootstrap 2 adds to final reward 1.

Question 3

What should lambda = 0 match?

TD(0) under the same sequential update convention.