Demonstration 1 of 4
What one episode records: return-to-go and the action mask
What return is each step credited with, and which tokens of an episode receive training loss?
Equation (13.1) lists what an episode contains and then defines G_t by adding the rewards from step t on, each multiplied by gamma raised to the number of steps since t. A late reward therefore counts less in an early step's return when gamma is below 1, and the last step's return is always just its own reward. Observations enter as context and carry mask 0, so only the policy's own sampled action gets loss: the chapter's loss is -[(1)(-0.4) + (0)(-1.2)] = 0.4.
Scroll sideways for the whole equation
tau is one whole episode: observations o, actions a and rewards r, ending at the terminal observation o_T. r_t is the reward at step t. G_t is the return-to-go: the discounted sum of rewards from step t to the end. gamma is the discount, a number from 0 to 1 that shrinks later rewards. T is the number of reward steps. The mask is a zero-or-one weight on each token's loss: the sampled retrieve action has log probability -0.4 and the observed tool reply text has log probability -1.2.
Predict first. In the exercise stream (rewards 1 then 2), what is G_0 when gamma is 0.5? Then switch the tool-text mask to 1 and watch the loss.
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's own exercise (rewards 1 and 2), the chapter's mask unit test (log probabilities -0.4 and -1.2) and a three-step release task with the book's retrieval cost 0.05, defined for this reader.
Calculated values
- G_0
- 0.950
- G_1
- 1.000
- G_2
- 1.000
- Sum of rewards, no discount
- 0.950
- Weight on the last reward in G_0
- 1.000
- Selected-token loss
- 0.4
G_0 = (-0.05) + 1.0 x 0.0 + 1.0 x 1.0 = 0.950. G_1 = 0.0 + 1.0 x 1.0 = 1.000. G_2 = 1.0 = 1.000. With gamma = 1 nothing is discounted, so G_0 equals the plain sum of the rewards. G_t only adds rewards from step t onward. Loss = -[(1)(-0.4)+(0)(-1.2)] = 0.4: only the policy's own retrieve action receives token loss.
Worked steps
- Rewards are r_0 = (-0.05), r_1 = 0.0, r_2 = 1.0.
- The last step has nothing after it: G_2 = r_2 = 1.000.
- G_1 = r_1 + gamma x G_2 = 0.0 + 1.0 x 1.0 = 1.000.
- G_0 = r_0 + gamma x G_1 = (-0.05) + 1.0 x 1.0 = 0.950.
- Masked loss: -[(1)(-0.4) + (0)(-1.2)] = 0.4.
Use the idea
When a training log reports one number per episode, ask which return it is, and when a loss is reported, ask which tokens it covers. A comparison needs the same definition on both sides.
Where the conclusion applies
A short episode with fixed rewards and no randomness, and one gradient computation for the mask. Real trajectories include many tokens, and only the policy's own actions receive credit. The release task values are constructed for this reader; the log probabilities -0.4 and -1.2 are the chapter's.
Common wrong turn: A stable gradient means the right tokens are being trained
What this does not settle
The masked-token check confirms one gradient computation. It does not show that a completed policy update improved true success.
Chapter 13 source: "What this does not settle".
Check your understanding: For rewards 1 then 2 with gamma = 0.9, what return is credited to the first action?
Chapter 13 source: section "What one training episode contains". Demonstration C13-D01.
Demonstration 2 of 4
Retrieve or answer: the gradient of one choice and what a baseline changes
If a policy retrieves with probability p = sigma(phi), what is the expected learning signal, and what does a baseline do to its noise?
Expected utility is 0.55 plus 0.20 times p. Because p = sigma(phi) flattens near 0 and 1, the slope is 0.20 p (1 - p), largest at p = 0.5. Equation (13.2) says each sampled update is the score times the return; its average is that slope. Subtracting a baseline that depends only on the context leaves the average alone because the scores average to zero, but it shrinks the swings between samples. A baseline that uses the sampled action breaks that argument.
Scroll sideways for the whole equation
phi is the trainable policy parameter and p = sigma(phi) the probability of retrieving before answering. J is the expected utility: 0.55 for answering directly and 0.80 - 0.05 = 0.75 for retrieving first, mixed by p. sigma is the logistic function, sigma(z) = 1 / (1 + e^(-z)). In Equation (13.2), pi_phi(a_t | o_t) is the probability the policy gives action a_t after observation o_t, tau is an episode generated by following that policy, E is the average over such episodes, G_t is the return of the sampled action, and the log-probability term is its score: 1 - p for retrieving and -p for answering. Here T = 1 and gamma^t = 1, because there is one decision. b is a baseline subtracted from G, and the update signal g is the score times (G - b).
Predict first. At phi = 0, switch the baseline from none to the expected return b. What happens to the mean update signal and to its spread?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's retrieve-or-answer values 0.55, 0.80 and the cost 0.05; the baselines are defined for this reader from the chapter's Figure 13.2 discussion.
Calculated values
- Retrieve probability p
- 0.5000
- Expected utility J
- 0.6500
- True gradient dJ/dphi
- 0.0500
- Mean of the update signal
- 0.0500
- Spread (standard deviation)
- 0.3957
- Baseline value
- 0
p = 1 / (1 + e^(0.0)) = 0.5000; true gradient = 0.20 x 0.5000 x 0.5000 = 0.0500. Mean of g = 0.2250 x 0.0 + 0.2750 x (-0.5) + 0.1000 x (-0.025) + 0.4000 x 0.475 = 0.0500. E[g^2] = 0.2250 x 0.0000 + 0.2750 x 0.2500 + 0.1000 x 0.0006 + 0.4000 x 0.2256 = 0.1591, so spread = sqrt(0.1591 - 0.0025) = 0.3957. This is Equation (13.2) with T = 1 and no baseline: the mean equals the true gradient 0.0500, and the spread 0.3957 is the noise in each sampled update.
Worked steps
- Retrieve probability p = sigma(0.0) = 0.5000.
- Scores: retrieve 1 - p = 0.5000, answer -p = (-0.5). Check: p x (1 - p) + (1 - p) x (-p) = 0.
- Returns G: answer fails 0, answer wins 1, retrieve fails -0.05, retrieve wins 0.95.
- Baseline b is 0; update signal g = score x (G - b).
- g by outcome: answer fails 0.0, answer wins (-0.5), retrieve fails (-0.025), retrieve wins 0.475.
- Mean = sum of probability x g = 0.0500; true gradient 0.0500.
- Spread = sqrt(E[g^2] - mean^2) = sqrt(0.1591 - 0.0025) = 0.3957.
Use the idea
With a logistic parameterization, a policy that is already nearly sure of its choice receives almost no learning signal, even if the other choice is better. A training log should show action frequencies, not only an average reward.
Where the conclusion applies
One decision with two actions, a fixed environment and exact expectations over the four outcomes. The expectation-zero argument requires the stated conditioning and action independence. The success values 0.55 and 0.80 and the cost 0.05 are the book's constructed numbers.
Common wrong turn: A learned critic has found which earlier action caused the outcome
What this does not settle
Equation (13.2) supplies a direction in expectation, not a guarantee that any single sampled update is beneficial. The retrieve-or-answer numbers are stipulated for teaching.
Chapter 13 source: "What this does not settle".
Check your understanding: With retrieve success 0.70 (cost still 0.05) and p = 0.5, what is the true gradient?
Chapter 13 source: section "Retrieve or answer". Demonstration C13-D02.
Demonstration 3 of 4
A verifier can teach the wrong success
If the verifier rewards the action that the task does not need, does more training help or hurt the task?
The update raises the probability of whichever action earns more reward than average. If that action has the lower success, the policy drifts toward failure while its reward curve rises. Swapping the two rewards makes the same update raise success. In the transfer case the terminal potential 2 lifts action 0's trained-on reward above action 1's, and in the story case the shortcut tool is accepted every time. The diamonds show a separately sampled evaluation: a single sampled rate is a noisy view of the exact number.
Scroll sideways for the whole equation
The policy chooses among actions, each with a reward the learner trains on and a true-success probability (what the task needs). The default case has rewards 1 and 2 with success 0.9 and 0.2; the changed case swaps the rewards; the transfer case has rewards 0 and 1, success 0 and 1 and terminal potentials 2 and 0; the story case uses the chapter's expected verifier scores 0.63, 0.73 and 0.95 for answering directly, retrieving and a shortcut tool with true success 0.55, 0.80 and 0. Training makes a sampled update in the form of Equation (13.2), with a baseline equal to the policy's expected reward and softmax logits, so only the trained-on reward drives the change. In that equation pi_phi(a_t | o_t) is the policy's probability of action a_t, tau an episode, E an average over episodes, G_t the return, and gamma^t = 1 because there is one decision (T = 1). Two random seeds show the spread; the diamond is a separately sampled evaluation.
Predict first. In the default case (the verifier prefers action 1), will the longest training leave true success above or below the starting 0.55?
Choose an example
Scroll sideways for the whole figure
Constructed example: the laboratory's default, changed and transfer training cases, and the chapter's noisy-verifier numbers (0.9, 0.3, 0.55, 0.80, 0.05, shortcut 0.95) as fixed expected scores, run by the laboratory's training function.
Calculated values
- Start: reward trained on
- 1.500
- Start: true success
- 0.550
- Seed 3: reward trained on
- 1.943
- Seed 3: true success
- 0.240
- Seed 11: true success
- 0.257
- Seed 3: sampled evaluation, n = 200
- 0.245
- Learner's best action
- action 1
- Task's best action
- action 0
Default case. Start: 0.500 x 1.0 + 0.500 x 2.0 = 1.500 reward and 0.500 x 0.9 + 0.500 x 0.2 = 0.550 true success. After 120 episodes seed 3 has action 0 0.0566, action 1 0.9434: reward = 0.0566 x 1.0 + 0.9434 x 2.0 = 1.9434; true success = 0.0566 x 0.9 + 0.9434 x 0.2 = 0.2396. A sample of 200 from exact 0.240 has standard error sqrt(0.240 x 0.760 / 200) = 0.030, so the sampled evaluation 0.245 is a noisy view of the exact number. The verifier pays most for the action with the lowest task success, so the learner moves toward it. The reward curve climbs while the task result falls, and a plot of reward alone would call this progress.
Worked steps
- Trained-on rewards (reward + gamma x terminal potential): action 0 1.0 + 1 x 0.0 = 1.0, action 1 2.0 + 1 x 0.0 = 2.0.
- Start with equal odds; reward = 0.500 x 1.0 + 0.500 x 2.0 = 1.500, true success = 0.500 x 0.9 + 0.500 x 0.2 = 0.550.
- Each episode samples an action and raises its odds when its reward beats the policy's average reward.
- After 120 episodes (seed 3): action 0 0.0566, action 1 0.9434.
- Reward trained on = 0.0566 x 1.0 + 0.9434 x 2.0 = 1.9434.
- True success = 0.0566 x 0.9 + 0.9434 x 0.2 = 0.2396.
- The learner's best action is action 1; the task's best is action 0 (a conflict).
Use the idea
Before calling a rising reward curve an improvement, compute what the policy would score under a check that the reward model does not share, such as an outside audit of the outcome, and keep the evaluation sample separate from training.
Where the conclusion applies
One-step choice, a fixed learning rate (0.08, or 0.1 in the transfer case), and two seeds. Rewards are fixed numbers, so the story case uses the chapter's expected scores rather than noisy draws. Real tasks have many steps. The seeds show that a run is a sample; they are not an estimate of how often it happens elsewhere.
Common wrong turn: A rising verifier score is a rising true-success curve
What this does not settle
A policy update that raises measured reward while reward-model acceptance diverges from true success is not thereby shown safe to ship. The noisy-verifier numbers are stipulated for teaching.
Chapter 13 source: "What this does not settle".
Check your understanding: In the transfer case, if training ended at probability 0.90 for action 0, what is the true success?
Chapter 13 source: section "A verifier can teach the wrong success". Demonstration C13-D03.
Demonstration 4 of 4
Reward shaping: what cancels and what does not
Which parts of a shaped return survive the cancellation, and when can an end potential change which route ranks first?
Each shaped reward adds the next state's potential and subtracts the current one, so a potential added when entering a state is subtracted when leaving it. Only the start and end potentials remain, as Equation (13.4) says. An end potential that only one route reaches does not cancel across routes, and can change the ranking.
Scroll sideways for the whole equation
r_t is the original reward and r'_t the shaped reward. Pot(x) is a potential assigned to state x: 0 at the start, a chosen value at the evidence state, and chosen end values on the retrieve route and on the direct route. gamma is 1 (no discount). G_0 and G'_0 are the original and shaped returns of the whole episode, and T the number of steps. Both routes start at x_0 with potential 0.
Predict first. Give the direct route's end the potential 0.3 (and the retrieve route's end 0). Which route ranks first?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's release construction (potential 0 at start and end, 0.4 after retrieval) with the book's 0.55, 0.80 and 0.05 and its 0.30 bonus; other potentials defined for this reader.
Calculated values
- Retrieve route G_0 (success case)
- 0.95
- Retrieve route G'_0 (success case)
- 0.95
- Direct route G'_0
- 0.55
- Retrieve route expected G'_0 (0.80 success)
- 0.75
- Ranked first after shaping
- Retrieve then answer
r'_0 = (-0.05) + 1 x 0.4 - 0 = 0.35. r'_1 = 1 + 1 x 0.0 - 0.4 = 0.60. Their sum is 0.95, which equals G_0 - 0 + 1 x 0.0 = 0.95 from Equation (13.4): the 0.4 added at t=0 is taken back at t=1, and only the end potential survives. Routes: direct = 0.55 + 1 x 0.0 = 0.55, retrieve = 0.80 - 0.05 + 1 x 0.0 = 0.75. Retrieving ranks first. Both end potentials are 0, so the shaped returns equal the original ones and nothing has changed.
Worked steps
- Potentials: start 0, evidence state 0.4, retrieve-route end 0.0, direct-route end 0.0.
- r'_0 = r_0 + Pot(evidence) - Pot(start) = (-0.05) + 0.4 - 0 = 0.35.
- r'_1 = r_1 + Pot(end) - Pot(evidence) = 1 + 0.0 - 0.4 = 0.60.
- Sum G'_0 = 0.35 + 0.60 = 0.95; Equation (13.4) gives 0.95 - 0 + 0.0 = 0.95.
- Direct route: 0.55 + 0.0 = 0.55.
- Retrieve route: 0.75 + 0.0 = 0.75.
- Ranked first: Retrieve then answer.
Use the idea
If a team adds a bonus for visiting a useful intermediate state, check where the potential ends up. A bonus that vanishes at the end is bookkeeping; a bonus that remains at some end states is a new objective.
Where the conclusion applies
No discount, a fixed starting state and exact expected returns. The potential is a function of state, as Equation (13.3) requires. The values 0.55, 0.80, 0.05 and 0.4 are the book's constructed numbers; the other potentials are defined for this reader.
Common wrong turn: An extra reward discovers an efficient policy
What this does not settle
The derivation is a finite constructed one. A process scorer that emits arbitrary positive scores does not automatically take the potential-difference form, and fitted scores need separate evidence about what they mean.
Chapter 13 source: "This is a finite constructed derivation, not a blanket guarantee about learned reward models.".
Check your understanding: If the direct route's end state had potential 0.1 and the retrieve route's end 0, which route would rank first, and by how much?
Chapter 13 source: section "When the reward arrives in pieces". Demonstration C13-D04.