Compute default trace update to draft.
Updating draft uses learning rate 0.5, a temporal-difference error of 1 at review (reward 1, no bootstrap) and trace weight discount times lambda, 1(0.8)=0.8: 0.5*1*0.8=0.4.
What are changed return targets?
[3,3], because final bootstrap 2 adds to final reward 1.
What should lambda = 0 match?
TD(0) under the same sequential update convention.