Demonstration 1 of 4
What a discount does to a distant reward
How much is a reward of 100 worth now if it arrives many steps from now, and what sets the planning scale?
Each step of delay multiplies the reward by gamma, so the worth falls geometrically without ever reaching zero. A steady reward of 1 per step adds up to 1/(1 - gamma), and that same number sets the scale on which delayed rewards fade to roughly a third. The factor is a declared valuation, not a measurement of how long a run lasts.
Scroll sideways for the whole equation
Ret_t is the return from time t: the sum of all future rewards, each multiplied by the discount factor gamma raised to how many steps away it is. r_(t+k) is the reward received k steps after t. gamma is a number between 0 and 1 that the designer declares. Here a single reward of 100 arrives after the chosen number of steps, so only one term of the sum is nonzero.
Predict first. With gamma = 0.9, will a reward of 100 that is 50 steps away be worth more or less than 1 today?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's discount arithmetic for a reward of 100 (factors 0.9, 0.95 and 0.99 at ten, fifty and a hundred steps), with gamma 0.5 from the chapter's horizon discussion.
Calculated values
- Worth now
- 34.9
- Share of its value kept
- 34.87 out of 100
- Planning scale 1/(1 - gamma)
- 10 steps
- Worth at that scale
- 34.9
- Return of 1 every step, forever
- 10.0
Only the reward at step 10 is counted here, so Equation (7.1) reduces to one term: 100 x 0.90^10 = 34.9. The scale 1/(1 - 0.90) = 1 / 0.10 = 10.0 steps is where a reward has fallen to roughly a third: 100 x 0.90^10 = 34.9. A steady reward of 1 per step adds up to 1 / (1 - 0.90) = 10.0. The curve never reaches zero, so a large enough delayed reward can still outweigh an immediate cost.
Worked steps
- A reward R = 100 arrives k = 10 steps from now, so only the k = 10 term of Equation (7.1) is nonzero.
- Worth now = R x gamma^k = 100 x 0.90^10 = 34.9.
- Planning scale = 1 / (1 - gamma) = 1 / 0.10 = 10.0 steps.
- At that scale the reward is worth 100 x 0.90^10 = 34.9, about a third of 100.
- A steady reward of 1 per step is worth 1 / (1 - 0.90) = 10.0 in total.
Use the idea
Before choosing a discount factor, compute what it does to the delayed benefit you care about. A factor that makes a useful verification or retrieval worth almost nothing will make the controller skip it.
Where the conclusion applies
One delayed reward of fixed size, a constant discount factor, and rewards that arrive on a regular step count. The factors 0.9, 0.95 and 0.99 come from the chapter's text and 0.5 from its discussion of a short horizon. Reading gamma as a survival probability needs a separate assumption of independent continuation; time preference does not.
Common wrong turn: A discount factor is a cutoff
What this does not settle
Discounting is a modeling choice rather than a fact, and a finite horizon is often a better choice than a discount.
Chapter 7 source: "What this does not settle".
Check your understanding: With gamma = 0.95, about how many steps does it take for a reward to fall to roughly a third of its value, and what is a reward of 100 worth at 20 steps?
Chapter 7 source: section "What discounting does to a number". Demonstration C07-D01.
Demonstration 2 of 4
Three states, computed backward
Verifying costs something now. Can it still be the better first move, and at what cost does it stop being so?
Start at the terminal state, whose remaining value is 0. The verified state has one action, so its value is its release payoff plus 0. At the start, releasing at once is worth its payoff plus 0, while verifying is worth minus its cost plus the discounted value of the verified state. The comparison that a one-step ledger misses sits in the second term. The same start state carries two values, one for each policy, which is what the superscript in Equation (7.2) records.
Scroll sideways for the whole equation
V^pi(x) is the value of state x when the agent follows policy pi: the expected return from x onward. pi(a | x) is the chance the policy takes action a in x, P(x' | x, a) the chance of landing in state x', r(x, a, x') the reward on that move, and gamma the discount factor. The states are start, verified and terminal. Releasing pays its expected payoff (85 at the start and 97 once verified in the chapter's controller). Release now, verified release, verification cost and gamma are shown with each scenario.
Predict first. With the chapter's controller (verification costs 5), which start policy is worth more once you reach the start state?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's three-state controller (release values 85 and 97, verification cost 5; the break-even cost 12 and the break-even discount 0.9278 are derived from it as 97 - 85 and 90/97, not stated in the chapter) and the original workbench problems II.1 and II.3 (80, 95, cost 6, discount 0.90 and 0.75), computed with the laboratory's backward induction.
Calculated values
- Value of the terminal state
- 0.0
- Value of the verified state
- not computed yet
- Start value, release at once
- not computed yet
- Start value, verify then release
- not computed yet
- Gain from verifying
- not computed yet
- Chosen at the start
- not computed yet
Begin where nothing more can be collected. The stop action pays 0 and leads back to the terminal state, so its value is 0 + 1.00 x 0 = 0.0. The release reward is paid on the transition into this state, so counting it again here would count it twice. No other value can be written until a successor value exists, which is why the order runs backward.
Worked steps
- The terminal state has one action, stop, with reward 0 and the terminal state as its successor.
- Its value is 0 + 1.00 x 0 = 0.0.
- The release payoff arrives on the transition into this state, so it is not counted a second time here.
- Step to the next state: the verified state needs only this number.
Use the idea
When a check, a retrieval or a clarifying question looks like pure cost, write the value of the state it leads to and compare. The break-even cost and the break-even discount tell you how expensive the check can be, and how heavily the future may be discounted, before skipping it is right.
Where the conclusion applies
Three states, one verification opportunity, deterministic movement into the verified state, and expected payoffs standing in for the release gamble. Real verification can fail or be unavailable. Equation (7.3) only prices a policy it is handed; choosing the better one needs the extra comparison of Equation (7.4). The workbench scenarios put the release probabilities 0.80 and 0.95 on a payoff of 100.
Common wrong turn: Judging verification by its own reward
What this does not settle
Backward induction requires a state space small enough to enumerate, which agent state spaces are not, and Chapter 9 addresses that directly. The two-step example is constructed for teaching.
Chapter 7 source: "What this does not settle".
Check your understanding: If a verified release were worth 90 (support probability 0.90), release at once were worth 85 and verifying cost 8, what would verifying be worth at the start with gamma = 1, and does the controller verify?
Chapter 7 source: section "Three states, computed backward". Demonstration C07-D02.
Demonstration 3 of 4
How many decisions are left can change the answer
When does it pay to spend a step investing in a later payoff instead of taking a smaller payoff now?
With one decision left at the start, investing uses it and the next state has none left, so it is worth 0 and investing scores its cost. With two or more left the next state is worth its payoff, so investing scores its cost plus gamma times that payoff. Equation (7.4) compares that with the stopping action. A discount below one cuts the later payoff until stopping wins, and on the last call stopping is simply the best action: exploration has nothing left to improve.
Scroll sideways for the whole equation
V*(x) is the best achievable value from state x. For each action a, the bracket is the immediate reward r plus gamma times the value of the next state x', averaged over where the action leads with probability P(x' | x, a); that bracketed value of one action is its action value Q. The max keeps the largest. The start state offers a stopping action (cash 2, or now 1) or an investing action (prepare -1 then a release of 6, or later 0 then a collection of 4). The horizon is the number of decisions remaining.
Predict first. With gamma = 1 in the cash or prepare model, does preparing beat cash when two decisions remain? And when only one remains?
Choose an example
Scroll sideways for the whole figure
Constructed example: the laboratory notebook's cash, prepare and release values (2, -1, 6) for the default and changed cases, and its transfer values (now 1, later 0, collect 4), computed with the laboratory's backward induction.
Calculated values
- Value of the immediate action
- 2.00
- Value of the investing action
- 5.00
- Start value (the larger)
- 5.00
- Chosen
- prepare
- Break-even discount
- 0.50
Cash = 2 + 1.00 x 0 = 2.00. Prepare = (-1) + 1.00 x 6 = 5.00, because the next state is worth 6 with 1 decision left. Prepare wins by 3.00: the cost now is repaid by the payoff it enables.
Worked steps
- With 2 decisions left, the next state is worth 6 (one more decision collects 6).
- Action value of cash = 2 + 1.00 x 0 = 2.00.
- Action value of prepare = (-1) + 1.00 x 6 = 5.00.
- Equation (7.4) keeps the larger: start value = 5.00, chosen: prepare.
- Break-even discount = (2 - (-1)) / 6 = 0.50.
Use the idea
A policy computed for three remaining steps can be wrong once steps have been spent (in the cash or prepare model at discount 1, preparing is right with two or three decisions left but wrong with one). Store the policy for each remaining horizon, and expect information-gathering moves to disappear near the end of a budget.
Where the conclusion applies
A fully observed three-state model with certain transitions, rewards fixed in advance, and a stopping state that pays zero. The optimum holds inside this model only; an outage or a missing permission that the model omits can reverse the ranking. Ties are reported as ties.
Common wrong turn: One policy for every remaining horizon
What this does not settle
The recursion assumes a Markov state. Direct model-based computation and state-indexed action selection add requirements for a specified model and access to that state, respectively.
Chapter 7 source: "What this does not settle".
Check your understanding: In the now or later model with two decisions left and gamma = 0.2, which action is worth more, and by how much?
Chapter 7 source: section "Horizons, and why agents have short ones". Demonstration C07-D03.
Demonstration 4 of 4
A residual bounds the error of a value guess
If you only check how far a candidate value function is from its own one-step lookahead, what do you learn about its distance from the best values?
The residual measures the largest one-step inconsistency in the guess. Dividing by 1 - gamma, the contraction margin, turns that into a ceiling on the true error, so the ceiling grows as gamma approaches 1 and does not exist at 1. Stepping through the stages shows the residual shrink as the guess improves: releasing at once is the value of one policy, the improvement step picks the better start action, and a guess that already satisfies the recursion has residual 0 and so has error 0.
Scroll sideways for the whole equation
V is a candidate value for each state. T maps V to the one-step lookahead values, the right side of Equation (7.4). V* is the best achievable value. The double bars with a small infinity mean the largest absolute difference over the three states. ||TV - V|| is the residual, ||V - V*|| the actual error, and gamma the discount factor, which must be below 1 for the bound. The stages are the all-zeros guess, the value of releasing at once, and the values after one improvement step.
Predict first. At gamma = 0.99, will the bound from the all-zero guess be close to the actual error, or far above it?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's three-state controller values with discount factors defined for the reader; best values computed with the laboratory's backward induction and each lookahead step computed with the laboratory's one-step backward update.
Calculated values
- Residual ||TV - V||
- 97.00
- Actual error ||V - V*||
- 97.00
- Bound from Equation (7.5)
- 1940.0
- Best start action under V*
- verify
Candidate: guess: all zeros. At the start, TV = max(85 + 0.95 x 0, (-5) + 0.95 x 0) = max(85.0, (-5.00)) = 85.00, against V = 0.00. The verified state gives TV = 97 + 0.95 x 0 = 97.0 against V = 0. The residual is the largest gap, 97.00. Bound = 97.00 / (1 - 0.95) = 97.00 / 0.05 = 1940.0. The true error is 97.00, which is inside the bound, though the bound can be loose (here 20.0 times the error). It bounds error for this declared model only.
Worked steps
- Candidate V (guess: all zeros): start 0.00, verified 0.0, terminal 0.0.
- Lookahead at the start: max(85 + 0.95 x 0, (-5) + 0.95 x 0) = 85.00.
- Lookahead at the verified state: 97 + 0.95 x 0 = 97.0.
- Residual = largest gap between TV and V = 97.00.
- Bound = 97.00 / (1 - 0.95) = 1940.0.
- Actual error (needs V*) = 97.00; it never exceeds the bound.
Use the idea
When values are learned or approximated, the residual can be computed without knowing V*. A small residual with a discount well below 1 is a certificate about the declared model, not about the world the model describes.
Where the conclusion applies
A finite model with bounded rewards and 0 <= gamma < 1. The three-state controller here (verification cost 5, verified release 97, unverified release 85) is constructed, and a discount below 1 changes which start action is best compared with the chapter's undiscounted case. A small residual against a wrong model still guides a bad action.
Common wrong turn: A small residual means a good action
What this does not settle
These results describe the supplied rewards, transitions, state, and discount convention. They cannot validate those inputs.
Chapter 7 source: "Summary".
Check your understanding: For the all-zeros guess at gamma = 0.9, the residual is 97. What does Equation (7.5) allow the error to be, at most?
Chapter 7 source: section "Where the reward comes from". Demonstration C07-D04.