The Mathematics of AI Agents, laboratory reader ยท Chapter 7

The Equation That Looks Ahead

A choice made now must carry the value of what it makes possible later, and one recursion computes that value without anyone folding it in by hand.

These four demonstrations follow the chapter's document controller. It can release a report now or spend a step verifying first. Each demonstration changes one declared input (a discount, a cost, the time left, a candidate answer) and shows how the value of the future enters the present comparison. Two of them step through a calculation one state or one improvement at a time, and the chapter's workbench and transfer cases are selectable.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

What a discount does to a distant reward

How much is a reward of 100 worth now if it arrives many steps from now, and what sets the planning scale?

Each step of delay multiplies the reward by gamma, so the worth falls geometrically without ever reaching zero. A steady reward of 1 per step adds up to 1/(1 - gamma), and that same number sets the scale on which delayed rewards fade to roughly a third. The factor is a declared valuation, not a measurement of how long a run lasts.

Equation (7.1), written in LaTeX: \operatorname{Ret}_t=\sum_{k=0}^{\infty}\gamma^{k}r_{t+k}

Scroll sideways for the whole equation

Ret_t is the return from time t: the sum of all future rewards, each multiplied by the discount factor gamma raised to how many steps away it is. r_(t+k) is the reward received k steps after t. gamma is a number between 0 and 1 that the designer declares. Here a single reward of 100 arrives after the chosen number of steps, so only one term of the sum is nonzero.

Predict first. With gamma = 0.9, will a reward of 100 that is 50 steps away be worth more or less than 1 today?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: What a discount does to a distant reward. Five decay curves of a reward of 100 against steps away. The curve for discount factor 0.90 is highlighted; at 10 steps it is worth 34.9, and a diamond marks the planning scale of 10 steps.
Discount factor gamma: 0.9, Steps until the reward arrives: 10
Constructed example: the chapter's discount arithmetic for a reward of 100 (factors 0.9, 0.95 and 0.99 at ten, fifty and a hundred steps), with gamma 0.5 from the chapter's horizon discussion.

Calculated values

Worth now
34.9
Share of its value kept
34.87 out of 100
Planning scale 1/(1 - gamma)
10 steps
Worth at that scale
34.9
Return of 1 every step, forever
10.0

Only the reward at step 10 is counted here, so Equation (7.1) reduces to one term: 100 x 0.90^10 = 34.9. The scale 1/(1 - 0.90) = 1 / 0.10 = 10.0 steps is where a reward has fallen to roughly a third: 100 x 0.90^10 = 34.9. A steady reward of 1 per step adds up to 1 / (1 - 0.90) = 10.0. The curve never reaches zero, so a large enough delayed reward can still outweigh an immediate cost.

Worked steps

  1. A reward R = 100 arrives k = 10 steps from now, so only the k = 10 term of Equation (7.1) is nonzero.
  2. Worth now = R x gamma^k = 100 x 0.90^10 = 34.9.
  3. Planning scale = 1 / (1 - gamma) = 1 / 0.10 = 10.0 steps.
  4. At that scale the reward is worth 100 x 0.90^10 = 34.9, about a third of 100.
  5. A steady reward of 1 per step is worth 1 / (1 - 0.90) = 10.0 in total.

Use the idea

Before choosing a discount factor, compute what it does to the delayed benefit you care about. A factor that makes a useful verification or retrieval worth almost nothing will make the controller skip it.

Where the conclusion applies

One delayed reward of fixed size, a constant discount factor, and rewards that arrive on a regular step count. The factors 0.9, 0.95 and 0.99 come from the chapter's text and 0.5 from its discussion of a short horizon. Reading gamma as a survival probability needs a separate assumption of independent continuation; time preference does not.

Common wrong turn: A discount factor is a cutoff
The chapter says the discount factor sets an effective planning scale, not a finite horizon. Every finite delay keeps a positive weight (100 x 0.9^50 = 0.52 is small but above zero), so a large enough delayed reward can still outweigh an immediate cost.
What this does not settle

Discounting is a modeling choice rather than a fact, and a finite horizon is often a better choice than a discount.

Chapter 7 source: "What this does not settle".

Check your understanding: With gamma = 0.95, about how many steps does it take for a reward to fall to roughly a third of its value, and what is a reward of 100 worth at 20 steps?
The scale is 1 / (1 - 0.95) = 1 / 0.05 = 20 steps, and 100 x 0.95^20 = 35.8, roughly a third.

Chapter 7 source: section "What discounting does to a number". Demonstration C07-D01.

Demonstration 2 of 4

Three states, computed backward

Verifying costs something now. Can it still be the better first move, and at what cost does it stop being so?

Start at the terminal state, whose remaining value is 0. The verified state has one action, so its value is its release payoff plus 0. At the start, releasing at once is worth its payoff plus 0, while verifying is worth minus its cost plus the discounted value of the verified state. The comparison that a one-step ledger misses sits in the second term. The same start state carries two values, one for each policy, which is what the superscript in Equation (7.2) records.

Equation (7.2), written in LaTeX: V^{\pi}(x)=\operatorname{E}_{\pi}[\operatorname{Ret}_t \mid x_t=x]

Equation (7.3), written in LaTeX: \begin{aligned}V^{\pi}(x)&=\sum_{a}\pi(a\mid x)\sum_{x'}P(x'\mid x,a)\\& \cdot[r(x,a,x')+\gamma V^{\pi}(x')].\end{aligned}

Scroll sideways for the whole equation

V^pi(x) is the value of state x when the agent follows policy pi: the expected return from x onward. pi(a | x) is the chance the policy takes action a in x, P(x' | x, a) the chance of landing in state x', r(x, a, x') the reward on that move, and gamma the discount factor. The states are start, verified and terminal. Releasing pays its expected payoff (85 at the start and 97 once verified in the chapter's controller). Release now, verified release, verification cost and gamma are shown with each scenario.

Predict first. With the chapter's controller (verification costs 5), which start policy is worth more once you reach the start state?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Three states, computed backward. Bar chart with one row filled: the terminal state with value 0. The other three rows say they are not computed yet.
Scenario: Chapter controller: cost 5, Backward step: 1. Terminal state
Constructed example: the book's three-state controller (release values 85 and 97, verification cost 5; the break-even cost 12 and the break-even discount 0.9278 are derived from it as 97 - 85 and 90/97, not stated in the chapter) and the original workbench problems II.1 and II.3 (80, 95, cost 6, discount 0.90 and 0.75), computed with the laboratory's backward induction.

Calculated values

Value of the terminal state
0.0
Value of the verified state
not computed yet
Start value, release at once
not computed yet
Start value, verify then release
not computed yet
Gain from verifying
not computed yet
Chosen at the start
not computed yet

Begin where nothing more can be collected. The stop action pays 0 and leads back to the terminal state, so its value is 0 + 1.00 x 0 = 0.0. The release reward is paid on the transition into this state, so counting it again here would count it twice. No other value can be written until a successor value exists, which is why the order runs backward.

Worked steps

  1. The terminal state has one action, stop, with reward 0 and the terminal state as its successor.
  2. Its value is 0 + 1.00 x 0 = 0.0.
  3. The release payoff arrives on the transition into this state, so it is not counted a second time here.
  4. Step to the next state: the verified state needs only this number.

Use the idea

When a check, a retrieval or a clarifying question looks like pure cost, write the value of the state it leads to and compare. The break-even cost and the break-even discount tell you how expensive the check can be, and how heavily the future may be discounted, before skipping it is right.

Where the conclusion applies

Three states, one verification opportunity, deterministic movement into the verified state, and expected payoffs standing in for the release gamble. Real verification can fail or be unavailable. Equation (7.3) only prices a policy it is handed; choosing the better one needs the extra comparison of Equation (7.4). The workbench scenarios put the release probabilities 0.80 and 0.95 on a payoff of 100.

Common wrong turn: Judging verification by its own reward
On a one-step ledger verification scores (-5) against 85 for releasing, so the controller never asks. The chapter's point is that the verified state's value comes back through the successor term: (-5) + 97 = 92, a gain of 7.
What this does not settle

Backward induction requires a state space small enough to enumerate, which agent state spaces are not, and Chapter 9 addresses that directly. The two-step example is constructed for teaching.

Chapter 7 source: "What this does not settle".

Check your understanding: If a verified release were worth 90 (support probability 0.90), release at once were worth 85 and verifying cost 8, what would verifying be worth at the start with gamma = 1, and does the controller verify?
Verifying is worth (-8) + 90 = 82 and releasing at once is worth 85, so the controller releases. The break-even cost is 90 - 85 = 5.

Chapter 7 source: section "Three states, computed backward". Demonstration C07-D02.

Demonstration 3 of 4

How many decisions are left can change the answer

When does it pay to spend a step investing in a later payoff instead of taking a smaller payoff now?

With one decision left at the start, investing uses it and the next state has none left, so it is worth 0 and investing scores its cost. With two or more left the next state is worth its payoff, so investing scores its cost plus gamma times that payoff. Equation (7.4) compares that with the stopping action. A discount below one cuts the later payoff until stopping wins, and on the last call stopping is simply the best action: exploration has nothing left to improve.

Equation (7.4), written in LaTeX: V^{\star}(x)=\max_{a}\sum_{x'}P(x'\mid x,a)[r(x,a,x')+\gamma V^{\star}(x')]

Scroll sideways for the whole equation

V*(x) is the best achievable value from state x. For each action a, the bracket is the immediate reward r plus gamma times the value of the next state x', averaged over where the action leads with probability P(x' | x, a); that bracketed value of one action is its action value Q. The max keeps the largest. The start state offers a stopping action (cash 2, or now 1) or an investing action (prepare -1 then a release of 6, or later 0 then a collection of 4). The horizon is the number of decisions remaining.

Predict first. With gamma = 1 in the cash or prepare model, does preparing beat cash when two decisions remain? And when only one remains?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How many decisions are left can change the answer. Two panels for the cash or prepare model. Left: value of each action with 1, 2 and 3 decisions left at discount 1.00. Right: value against discount factor with 2 decisions left. The chosen action is prepare.
Model: Cash or prepare (notebook default and changed), Decisions remaining: 2, Discount factor gamma: 1
Constructed example: the laboratory notebook's cash, prepare and release values (2, -1, 6) for the default and changed cases, and its transfer values (now 1, later 0, collect 4), computed with the laboratory's backward induction.

Calculated values

Value of the immediate action
2.00
Value of the investing action
5.00
Start value (the larger)
5.00
Chosen
prepare
Break-even discount
0.50

Cash = 2 + 1.00 x 0 = 2.00. Prepare = (-1) + 1.00 x 6 = 5.00, because the next state is worth 6 with 1 decision left. Prepare wins by 3.00: the cost now is repaid by the payoff it enables.

Worked steps

  1. With 2 decisions left, the next state is worth 6 (one more decision collects 6).
  2. Action value of cash = 2 + 1.00 x 0 = 2.00.
  3. Action value of prepare = (-1) + 1.00 x 6 = 5.00.
  4. Equation (7.4) keeps the larger: start value = 5.00, chosen: prepare.
  5. Break-even discount = (2 - (-1)) / 6 = 0.50.

Use the idea

A policy computed for three remaining steps can be wrong once steps have been spent (in the cash or prepare model at discount 1, preparing is right with two or three decisions left but wrong with one). Store the policy for each remaining horizon, and expect information-gathering moves to disappear near the end of a budget.

Where the conclusion applies

A fully observed three-state model with certain transitions, rewards fixed in advance, and a stopping state that pays zero. The optimum holds inside this model only; an outage or a missing permission that the model omits can reverse the ranking. Ties are reported as ties.

Common wrong turn: One policy for every remaining horizon
The chapter says the same context with three calls left and with one call left are different states with different values. At discount 1 preparing wins with two or three decisions left and loses with one, so a policy stored for three steps cannot be reused unchanged on the last call.
What this does not settle

The recursion assumes a Markov state. Direct model-based computation and state-indexed action selection add requirements for a specified model and access to that state, respectively.

Chapter 7 source: "What this does not settle".

Check your understanding: In the now or later model with two decisions left and gamma = 0.2, which action is worth more, and by how much?
Later = 0 + 0.2 x 4 = 0.8 and now = 1 + 0.2 x 0 = 1.0, so now wins by 0.2. The break-even discount is (1 - 0) / 4 = 0.25.

Chapter 7 source: section "Horizons, and why agents have short ones". Demonstration C07-D03.

Demonstration 4 of 4

A residual bounds the error of a value guess

If you only check how far a candidate value function is from its own one-step lookahead, what do you learn about its distance from the best values?

The residual measures the largest one-step inconsistency in the guess. Dividing by 1 - gamma, the contraction margin, turns that into a ceiling on the true error, so the ceiling grows as gamma approaches 1 and does not exist at 1. Stepping through the stages shows the residual shrink as the guess improves: releasing at once is the value of one policy, the improvement step picks the better start action, and a guess that already satisfies the recursion has residual 0 and so has error 0.

Equation (7.5), written in LaTeX: \|V-V^\star\|_\infty\leq \frac{\|TV-V\|_\infty}{1-\gamma}

Equation (7.4), written in LaTeX: V^{\star}(x)=\max_{a}\sum_{x'}P(x'\mid x,a)[r(x,a,x')+\gamma V^{\star}(x')]

Scroll sideways for the whole equation

V is a candidate value for each state. T maps V to the one-step lookahead values, the right side of Equation (7.4). V* is the best achievable value. The double bars with a small infinity mean the largest absolute difference over the three states. ||TV - V|| is the residual, ||V - V*|| the actual error, and gamma the discount factor, which must be below 1 for the bound. The stages are the all-zeros guess, the value of releasing at once, and the values after one improvement step.

Predict first. At gamma = 0.99, will the bound from the all-zero guess be close to the actual error, or far above it?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A residual bounds the error of a value guess. Left: grouped bars of the candidate V, its lookahead TV and V* for the start, verified and terminal states. Right: bars for the actual error 97.0, the residual 97.0 and the bound 1940.0.
Discount factor gamma: 0.95, Candidate value function: 1. All zeros
Constructed example: the book's three-state controller values with discount factors defined for the reader; best values computed with the laboratory's backward induction and each lookahead step computed with the laboratory's one-step backward update.

Calculated values

Residual ||TV - V||
97.00
Actual error ||V - V*||
97.00
Bound from Equation (7.5)
1940.0
Best start action under V*
verify

Candidate: guess: all zeros. At the start, TV = max(85 + 0.95 x 0, (-5) + 0.95 x 0) = max(85.0, (-5.00)) = 85.00, against V = 0.00. The verified state gives TV = 97 + 0.95 x 0 = 97.0 against V = 0. The residual is the largest gap, 97.00. Bound = 97.00 / (1 - 0.95) = 97.00 / 0.05 = 1940.0. The true error is 97.00, which is inside the bound, though the bound can be loose (here 20.0 times the error). It bounds error for this declared model only.

Worked steps

  1. Candidate V (guess: all zeros): start 0.00, verified 0.0, terminal 0.0.
  2. Lookahead at the start: max(85 + 0.95 x 0, (-5) + 0.95 x 0) = 85.00.
  3. Lookahead at the verified state: 97 + 0.95 x 0 = 97.0.
  4. Residual = largest gap between TV and V = 97.00.
  5. Bound = 97.00 / (1 - 0.95) = 1940.0.
  6. Actual error (needs V*) = 97.00; it never exceeds the bound.

Use the idea

When values are learned or approximated, the residual can be computed without knowing V*. A small residual with a discount well below 1 is a certificate about the declared model, not about the world the model describes.

Where the conclusion applies

A finite model with bounded rewards and 0 <= gamma < 1. The three-state controller here (verification cost 5, verified release 97, unverified release 85) is constructed, and a discount below 1 changes which start action is best compared with the chapter's undiscounted case. A small residual against a wrong model still guides a bad action.

Common wrong turn: A small residual means a good action
The chapter says model error is not solution error: a small residual against a wrong transition or reward law can still guide a bad action. The bound certifies the values only for the declared model.
What this does not settle

These results describe the supplied rewards, transitions, state, and discount convention. They cannot validate those inputs.

Chapter 7 source: "Summary".

Check your understanding: For the all-zeros guess at gamma = 0.9, the residual is 97. What does Equation (7.5) allow the error to be, at most?
97 / (1 - 0.9) = 97 / 0.1 = 970. The true error is 97, well inside that ceiling.

Chapter 7 source: section "Where the reward comes from". Demonstration C07-D04.