The Mathematics of AI Agents, laboratory reader ยท Chapter 13

Learning to Choose: Training an Agent from Trajectories

Learning changes the chooser itself, and a reward that rises is not the same as a task that improves.

These four demonstrations follow the chapter's retrieve-or-answer agent. They show what a trajectory records and which tokens receive loss, how a policy-gradient update moves the chance of retrieving and what a baseline does to its noise, how a verifier that is wrong about success can steer training (the laboratory's default, changed and transfer cases and the chapter's shortcut story), and what reward shaping does and does not change. Every number is a constructed teaching value.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

What one episode records: return-to-go and the action mask

What return is each step credited with, and which tokens of an episode receive training loss?

Equation (13.1) lists what an episode contains and then defines G_t by adding the rewards from step t on, each multiplied by gamma raised to the number of steps since t. A late reward therefore counts less in an early step's return when gamma is below 1, and the last step's return is always just its own reward. Observations enter as context and carry mask 0, so only the policy's own sampled action gets loss: the chapter's loss is -[(1)(-0.4) + (0)(-1.2)] = 0.4.

Equation (13.1), written in LaTeX: \tau=(o_0,a_0,r_0,o_1,a_1,r_1,\ldots,o_T), G_t=\sum_{k=t}^{T-1}\gamma^{k-t}r_k.

Equation, written in LaTeX: -[(1)(-0.4)+(0)(-1.2)]=0.4

Scroll sideways for the whole equation

tau is one whole episode: observations o, actions a and rewards r, ending at the terminal observation o_T. r_t is the reward at step t. G_t is the return-to-go: the discounted sum of rewards from step t to the end. gamma is the discount, a number from 0 to 1 that shrinks later rewards. T is the number of reward steps. The mask is a zero-or-one weight on each token's loss: the sampled retrieve action has log probability -0.4 and the observed tool reply text has log probability -1.2.

Predict first. In the exercise stream (rewards 1 then 2), what is G_0 when gamma is 0.5? Then switch the tool-text mask to 1 and watch the loss.

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: What one episode records: return-to-go and the action mask. Left, bars of reward and return-to-go for the release task at gamma 1.0, G_0 = 0.950. Right, the two token loss terms, 0.4 for the retrieve action and 0.0 for the tool text with mask 0, total loss 0.4.
Reward stream: Release task: -0.05, 0, 1, Discount gamma: 1, Mask on the observed tool text: Mask 0 (excluded from the loss)
Constructed example: the chapter's own exercise (rewards 1 and 2), the chapter's mask unit test (log probabilities -0.4 and -1.2) and a three-step release task with the book's retrieval cost 0.05, defined for this reader.

Calculated values

G_0
0.950
G_1
1.000
G_2
1.000
Sum of rewards, no discount
0.950
Weight on the last reward in G_0
1.000
Selected-token loss
0.4

G_0 = (-0.05) + 1.0 x 0.0 + 1.0 x 1.0 = 0.950. G_1 = 0.0 + 1.0 x 1.0 = 1.000. G_2 = 1.0 = 1.000. With gamma = 1 nothing is discounted, so G_0 equals the plain sum of the rewards. G_t only adds rewards from step t onward. Loss = -[(1)(-0.4)+(0)(-1.2)] = 0.4: only the policy's own retrieve action receives token loss.

Worked steps

  1. Rewards are r_0 = (-0.05), r_1 = 0.0, r_2 = 1.0.
  2. The last step has nothing after it: G_2 = r_2 = 1.000.
  3. G_1 = r_1 + gamma x G_2 = 0.0 + 1.0 x 1.0 = 1.000.
  4. G_0 = r_0 + gamma x G_1 = (-0.05) + 1.0 x 1.0 = 0.950.
  5. Masked loss: -[(1)(-0.4) + (0)(-1.2)] = 0.4.

Use the idea

When a training log reports one number per episode, ask which return it is, and when a loss is reported, ask which tokens it covers. A comparison needs the same definition on both sides.

Where the conclusion applies

A short episode with fixed rewards and no randomness, and one gradient computation for the mask. Real trajectories include many tokens, and only the policy's own actions receive credit. The release task values are constructed for this reader; the log probabilities -0.4 and -1.2 are the chapter's.

Common wrong turn: A stable gradient means the right tokens are being trained
The chapter warns that if the tool-text mask becomes one, the implementation optimizes the probability of a response the policy did not generate, and the gradient can look stable while training on the wrong distribution. Set the mask to 1: the loss rises from 0.4 to 1.6.
What this does not settle

The masked-token check confirms one gradient computation. It does not show that a completed policy update improved true success.

Chapter 13 source: "What this does not settle".

Check your understanding: For rewards 1 then 2 with gamma = 0.9, what return is credited to the first action?
G_0 = 1 + 0.9 x 2 = 2.8. The second step's return stays G_1 = 2.

Chapter 13 source: section "What one training episode contains". Demonstration C13-D01.

Demonstration 2 of 4

Retrieve or answer: the gradient of one choice and what a baseline changes

If a policy retrieves with probability p = sigma(phi), what is the expected learning signal, and what does a baseline do to its noise?

Expected utility is 0.55 plus 0.20 times p. Because p = sigma(phi) flattens near 0 and 1, the slope is 0.20 p (1 - p), largest at p = 0.5. Equation (13.2) says each sampled update is the score times the return; its average is that slope. Subtracting a baseline that depends only on the context leaves the average alone because the scores average to zero, but it shrinks the swings between samples. A baseline that uses the sampled action breaks that argument.

Equation (13.2), written in LaTeX: \nabla_\phi J(\phi)=\operatorname E_{\tau\sim\pi_\phi}[\sum_{t=0}^{T-1}\gamma^t\nabla_\phi\log \pi_\phi(a_t\mid o_t) G_t].

Equation, written in LaTeX: \sum_a\pi_\phi(a\mid o)\nabla_\phi\log \pi_\phi(a\mid o)=0

Scroll sideways for the whole equation

phi is the trainable policy parameter and p = sigma(phi) the probability of retrieving before answering. J is the expected utility: 0.55 for answering directly and 0.80 - 0.05 = 0.75 for retrieving first, mixed by p. sigma is the logistic function, sigma(z) = 1 / (1 + e^(-z)). In Equation (13.2), pi_phi(a_t | o_t) is the probability the policy gives action a_t after observation o_t, tau is an episode generated by following that policy, E is the average over such episodes, G_t is the return of the sampled action, and the log-probability term is its score: 1 - p for retrieving and -p for answering. Here T = 1 and gamma^t = 1, because there is one decision. b is a baseline subtracted from G, and the update signal g is the score times (G - b).

Predict first. At phi = 0, switch the baseline from none to the expected return b. What happens to the mean update signal and to its spread?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Retrieve or answer: the gradient of one choice and what a baseline changes. Left, expected utility against phi with a dashed tangent at phi 0.0, utility 0.650. Right, four bars for the update signal of each sampled outcome with no baseline; the dashed mean line is 0.050 and the spread is 0.396.
Policy parameter phi: 0, Baseline subtracted from the return: None
Constructed example: the chapter's retrieve-or-answer values 0.55, 0.80 and the cost 0.05; the baselines are defined for this reader from the chapter's Figure 13.2 discussion.

Calculated values

Retrieve probability p
0.5000
Expected utility J
0.6500
True gradient dJ/dphi
0.0500
Mean of the update signal
0.0500
Spread (standard deviation)
0.3957
Baseline value
0

p = 1 / (1 + e^(0.0)) = 0.5000; true gradient = 0.20 x 0.5000 x 0.5000 = 0.0500. Mean of g = 0.2250 x 0.0 + 0.2750 x (-0.5) + 0.1000 x (-0.025) + 0.4000 x 0.475 = 0.0500. E[g^2] = 0.2250 x 0.0000 + 0.2750 x 0.2500 + 0.1000 x 0.0006 + 0.4000 x 0.2256 = 0.1591, so spread = sqrt(0.1591 - 0.0025) = 0.3957. This is Equation (13.2) with T = 1 and no baseline: the mean equals the true gradient 0.0500, and the spread 0.3957 is the noise in each sampled update.

Worked steps

  1. Retrieve probability p = sigma(0.0) = 0.5000.
  2. Scores: retrieve 1 - p = 0.5000, answer -p = (-0.5). Check: p x (1 - p) + (1 - p) x (-p) = 0.
  3. Returns G: answer fails 0, answer wins 1, retrieve fails -0.05, retrieve wins 0.95.
  4. Baseline b is 0; update signal g = score x (G - b).
  5. g by outcome: answer fails 0.0, answer wins (-0.5), retrieve fails (-0.025), retrieve wins 0.475.
  6. Mean = sum of probability x g = 0.0500; true gradient 0.0500.
  7. Spread = sqrt(E[g^2] - mean^2) = sqrt(0.1591 - 0.0025) = 0.3957.

Use the idea

With a logistic parameterization, a policy that is already nearly sure of its choice receives almost no learning signal, even if the other choice is better. A training log should show action frequencies, not only an average reward.

Where the conclusion applies

One decision with two actions, a fixed environment and exact expectations over the four outcomes. The expectation-zero argument requires the stated conditioning and action independence. The success values 0.55 and 0.80 and the cost 0.05 are the book's constructed numbers.

Common wrong turn: A learned critic has found which earlier action caused the outcome
The chapter says an actor-critic pairs the policy with a learned value function that serves as the baseline, and that it has not thereby identified which earlier action caused the outcome. The baseline only reduces noise.
What this does not settle

Equation (13.2) supplies a direction in expectation, not a guarantee that any single sampled update is beneficial. The retrieve-or-answer numbers are stipulated for teaching.

Chapter 13 source: "What this does not settle".

Check your understanding: With retrieve success 0.70 (cost still 0.05) and p = 0.5, what is the true gradient?
Net value of retrieving = 0.70 - 0.05 = 0.65, gain = 0.65 - 0.55 = 0.10, so the slope is 0.10 x 0.5 x 0.5 = 0.025.

Chapter 13 source: section "Retrieve or answer". Demonstration C13-D02.

Demonstration 3 of 4

A verifier can teach the wrong success

If the verifier rewards the action that the task does not need, does more training help or hurt the task?

The update raises the probability of whichever action earns more reward than average. If that action has the lower success, the policy drifts toward failure while its reward curve rises. Swapping the two rewards makes the same update raise success. In the transfer case the terminal potential 2 lifts action 0's trained-on reward above action 1's, and in the story case the shortcut tool is accepted every time. The diamonds show a separately sampled evaluation: a single sampled rate is a noisy view of the exact number.

Equation (13.2), written in LaTeX: \nabla_\phi J(\phi)=\operatorname E_{\tau\sim\pi_\phi}[\sum_{t=0}^{T-1}\gamma^t\nabla_\phi\log \pi_\phi(a_t\mid o_t) G_t].

Scroll sideways for the whole equation

The policy chooses among actions, each with a reward the learner trains on and a true-success probability (what the task needs). The default case has rewards 1 and 2 with success 0.9 and 0.2; the changed case swaps the rewards; the transfer case has rewards 0 and 1, success 0 and 1 and terminal potentials 2 and 0; the story case uses the chapter's expected verifier scores 0.63, 0.73 and 0.95 for answering directly, retrieving and a shortcut tool with true success 0.55, 0.80 and 0. Training makes a sampled update in the form of Equation (13.2), with a baseline equal to the policy's expected reward and softmax logits, so only the trained-on reward drives the change. In that equation pi_phi(a_t | o_t) is the policy's probability of action a_t, tau an episode, E an average over episodes, G_t the return, and gamma^t = 1 because there is one decision (T = 1). Two random seeds show the spread; the diamond is a separately sampled evaluation.

Predict first. In the default case (the verifier prefers action 1), will the longest training leave true success above or below the starting 0.55?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A verifier can teach the wrong success. Default case. Left, two training curves of the reward the learner trains on, rising from 1.50. Right, true success bars: start 0.550, seed 3 0.240, seed 11 0.257, with sampled evaluation diamonds.
Case: Default: rewards 1, 2; success 0.9, 0.2, Training length: The case's own length (120; transfer 100)
Constructed example: the laboratory's default, changed and transfer training cases, and the chapter's noisy-verifier numbers (0.9, 0.3, 0.55, 0.80, 0.05, shortcut 0.95) as fixed expected scores, run by the laboratory's training function.

Calculated values

Start: reward trained on
1.500
Start: true success
0.550
Seed 3: reward trained on
1.943
Seed 3: true success
0.240
Seed 11: true success
0.257
Seed 3: sampled evaluation, n = 200
0.245
Learner's best action
action 1
Task's best action
action 0

Default case. Start: 0.500 x 1.0 + 0.500 x 2.0 = 1.500 reward and 0.500 x 0.9 + 0.500 x 0.2 = 0.550 true success. After 120 episodes seed 3 has action 0 0.0566, action 1 0.9434: reward = 0.0566 x 1.0 + 0.9434 x 2.0 = 1.9434; true success = 0.0566 x 0.9 + 0.9434 x 0.2 = 0.2396. A sample of 200 from exact 0.240 has standard error sqrt(0.240 x 0.760 / 200) = 0.030, so the sampled evaluation 0.245 is a noisy view of the exact number. The verifier pays most for the action with the lowest task success, so the learner moves toward it. The reward curve climbs while the task result falls, and a plot of reward alone would call this progress.

Worked steps

  1. Trained-on rewards (reward + gamma x terminal potential): action 0 1.0 + 1 x 0.0 = 1.0, action 1 2.0 + 1 x 0.0 = 2.0.
  2. Start with equal odds; reward = 0.500 x 1.0 + 0.500 x 2.0 = 1.500, true success = 0.500 x 0.9 + 0.500 x 0.2 = 0.550.
  3. Each episode samples an action and raises its odds when its reward beats the policy's average reward.
  4. After 120 episodes (seed 3): action 0 0.0566, action 1 0.9434.
  5. Reward trained on = 0.0566 x 1.0 + 0.9434 x 2.0 = 1.9434.
  6. True success = 0.0566 x 0.9 + 0.9434 x 0.2 = 0.2396.
  7. The learner's best action is action 1; the task's best is action 0 (a conflict).

Use the idea

Before calling a rising reward curve an improvement, compute what the policy would score under a check that the reward model does not share, such as an outside audit of the outcome, and keep the evaluation sample separate from training.

Where the conclusion applies

One-step choice, a fixed learning rate (0.08, or 0.1 in the transfer case), and two seeds. Rewards are fixed numbers, so the story case uses the chapter's expected scores rather than noisy draws. Real tasks have many steps. The seeds show that a run is a sample; they are not an estimate of how often it happens elsewhere.

Common wrong turn: A rising verifier score is a rising true-success curve
The chapter says a training curve measured by the verifier cannot be renamed a true-success curve. In the default case the left curve climbs while the right bars fall.
What this does not settle

A policy update that raises measured reward while reward-model acceptance diverges from true success is not thereby shown safe to ship. The noisy-verifier numbers are stipulated for teaching.

Chapter 13 source: "What this does not settle".

Check your understanding: In the transfer case, if training ended at probability 0.90 for action 0, what is the true success?
0.90 x 0 + 0.10 x 1 = 0.10, far below the start of 0.5 x 0 + 0.5 x 1 = 0.50, even though the trained-on reward is 0.90 x 2 + 0.10 x 1 = 1.90.

Chapter 13 source: section "A verifier can teach the wrong success". Demonstration C13-D03.

Demonstration 4 of 4

Reward shaping: what cancels and what does not

Which parts of a shaped return survive the cancellation, and when can an end potential change which route ranks first?

Each shaped reward adds the next state's potential and subtracts the current one, so a potential added when entering a state is subtracted when leaving it. Only the start and end potentials remain, as Equation (13.4) says. An end potential that only one route reaches does not cancel across routes, and can change the ranking.

Equation (13.3), written in LaTeX: r'_t=r_t+\gamma\operatorname{Pot}(x_{t+1})-\operatorname{Pot}(x_t).

Equation (13.4), written in LaTeX: G'_0=G_0-\operatorname{Pot}(x_0)+\gamma^T\operatorname{Pot}(x_T).

Scroll sideways for the whole equation

r_t is the original reward and r'_t the shaped reward. Pot(x) is a potential assigned to state x: 0 at the start, a chosen value at the evidence state, and chosen end values on the retrieve route and on the direct route. gamma is 1 (no discount). G_0 and G'_0 are the original and shaped returns of the whole episode, and T the number of steps. Both routes start at x_0 with potential 0.

Predict first. Give the direct route's end the potential 0.3 (and the retrieve route's end 0). Which route ranks first?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Reward shaping: what cancels and what does not. Left, original and shaped rewards of the retrieve route with evidence potential 0.4; the shaped sum is 0.95. Right, original and shaped returns of the two routes: direct 0.55, retrieve 0.75; ranked first: Retrieve then answer.
Potential of the evidence state: 0.4, Potential at the end of the retrieve route: 0, Potential at the end of the direct route: 0
Constructed example: the chapter's release construction (potential 0 at start and end, 0.4 after retrieval) with the book's 0.55, 0.80 and 0.05 and its 0.30 bonus; other potentials defined for this reader.

Calculated values

Retrieve route G_0 (success case)
0.95
Retrieve route G'_0 (success case)
0.95
Direct route G'_0
0.55
Retrieve route expected G'_0 (0.80 success)
0.75
Ranked first after shaping
Retrieve then answer

r'_0 = (-0.05) + 1 x 0.4 - 0 = 0.35. r'_1 = 1 + 1 x 0.0 - 0.4 = 0.60. Their sum is 0.95, which equals G_0 - 0 + 1 x 0.0 = 0.95 from Equation (13.4): the 0.4 added at t=0 is taken back at t=1, and only the end potential survives. Routes: direct = 0.55 + 1 x 0.0 = 0.55, retrieve = 0.80 - 0.05 + 1 x 0.0 = 0.75. Retrieving ranks first. Both end potentials are 0, so the shaped returns equal the original ones and nothing has changed.

Worked steps

  1. Potentials: start 0, evidence state 0.4, retrieve-route end 0.0, direct-route end 0.0.
  2. r'_0 = r_0 + Pot(evidence) - Pot(start) = (-0.05) + 0.4 - 0 = 0.35.
  3. r'_1 = r_1 + Pot(end) - Pot(evidence) = 1 + 0.0 - 0.4 = 0.60.
  4. Sum G'_0 = 0.35 + 0.60 = 0.95; Equation (13.4) gives 0.95 - 0 + 0.0 = 0.95.
  5. Direct route: 0.55 + 0.0 = 0.55.
  6. Retrieve route: 0.75 + 0.0 = 0.75.
  7. Ranked first: Retrieve then answer.

Use the idea

If a team adds a bonus for visiting a useful intermediate state, check where the potential ends up. A bonus that vanishes at the end is bookkeeping; a bonus that remains at some end states is a new objective.

Where the conclusion applies

No discount, a fixed starting state and exact expected returns. The potential is a function of state, as Equation (13.3) requires. The values 0.55, 0.80, 0.05 and 0.4 are the book's constructed numbers; the other potentials are defined for this reader.

Common wrong turn: An extra reward discovers an efficient policy
The chapter says the additional score did not discover an efficient policy: it encoded a different trade-off, and the numerical reward cannot be presented as proof that the policy learned better evidence gathering. A 0.30 bonus for avoiding retrieval raises the direct route to 0.85 while its true completion stays 0.55.
What this does not settle

The derivation is a finite constructed one. A process scorer that emits arbitrary positive scores does not automatically take the potential-difference form, and fitted scores need separate evidence about what they mean.

Chapter 13 source: "This is a finite constructed derivation, not a blanket guarantee about learned reward models.".

Check your understanding: If the direct route's end state had potential 0.1 and the retrieve route's end 0, which route would rank first, and by how much?
Direct = 0.55 + 0.1 = 0.65, retrieve = 0.80 - 0.05 = 0.75. Retrieving still ranks first, by 0.10.

Chapter 13 source: section "When the reward arrives in pieces". Demonstration C13-D04.