Demonstration 1 of 4
The score inside the model against the score in the world
If you pick the setting that scores best inside the learned model, how does it do in the actual environment?
The table has two columns because the researchers had two environments. Over the first four rows the columns run in opposite directions, so the setting that flatters the model most is the least useful in the world. Judging by the model column alone cannot show this; only the second column can. The scores do not supply the transition error epsilon or any action gap, so they do not give a certified horizon.
Scroll sideways for the whole equation
Temperature is the sampling parameter of the imagined world: low values make it cleaner and more predictable, high values harder. The model score is the controller's score when evaluated in its own learned environment, and the actual score is its score in the real game. The random policy and the best leaderboard entry are the chapter's reference scores. The table does not report the transition error epsilon of Equation (14.1), which is the maximum over states and actions of half the summed difference between the model's and the true next-state probabilities.
Predict first. Rank the five settings by the model column alone. Which wins, and how does it score in the actual environment against the random policy's 210?
Choose an example
Scroll sideways for the whole figure
Constructed example (reported values in a constructed comparison): the scores are copied from the chapter's temperature table (reported by the cited paper for one game setup); the judging rule and the comparison are defined for this reader.
Calculated values
- Temperature
- 0.10
- Score inside the model
- 2086
- Score in the actual environment
- 193
- Model minus actual
- 1893
- Actual minus random policy (210)
- -17
- Picked by the dream column
- 0.10
- Picked: score in the actual environment
- 193
- Picked: actual minus random policy (210)
- -17
- Reported spreads (model, actual)
- 2086 +/- 140, 193 +/- 58
At T = 0.10: model 2086 - actual 193 = 1893. Actual against the random policy: 193 - 210 = -17. Worst to best setting in the actual column: 1092 - 193 = 899 points. Judged only by the model column, the winner is T = 0.10 with 2086. In the actual environment it scores 193, and 193 - 210 = -17 puts it below the random policy: ranked first and confidently wrong. The chapter reports the authors' account of this row: at the lowest temperature the monsters fail to shoot at all (mode collapse), so a policy learned in that imagined world scores near-perfectly and then fails in the actual one. The reported spread 193 +/- 58 overlaps the random policy's 210 +/- 108.
Worked steps
- Read the chosen row: model 2086, actual 193.
- Model minus actual: 2086 - 193 = 1893.
- Actual minus the random policy's 210: 193 - 210 = -17.
- Rank the five settings by the dream column alone: the largest is 2086 at T = 0.10.
- That pick's actual score is 193: 193 - 210 = -17 against the random policy.
- Spreads were reported only for 0.10 and 1.15; the table does not give the transition error of Equation (14.1).
Use the idea
When a system proposes an action, predicts its outcome and scores itself on the prediction, it is reading the left column. Hold out some real outcomes and keep the second column, however small, and compare the two.
Where the conclusion applies
The scores are the chapter's reported numbers for one game setup and one training and evaluation procedure, each with its own variability (spreads were reported for 0.10 and 1.15 here). They show that the two evaluations can come apart, not a general temperature rule and not a measure of epsilon.
Common wrong turn: A higher score inside the model means a better agent
What this does not settle
The anchor is a preprint reporting on its authors' own systems in one game, and the temperature that worked there is not a value to copy.
Chapter 14 source: "What this does not settle".
Check your understanding: At T = 1.30, what is the model score minus the actual score, and which column is higher?
Chapter 14 source: section "The knob the authors added, and what it bought". Demonstration C14-D01.
Demonstration 2 of 4
One step of error, compounded, against the actual error
How large can the value error become when a small one-step error compounds, and how large is it in a constructed model?
The bound multiplies gamma, epsilon and R, then divides by (1 - gamma) squared: one factor of 1/(1-gamma) counts the steps that matter and the second counts that a mistake at each step displaces everything after it. The finite bound grows with H(H-1) instead. Each bound holds under its own assumptions, and the actual error from the two recurrences stays under both. Raising gamma shrinks nothing in the finite bound but inflates the discounted one.
Scroll sideways for the whole equation
P(x' | x, a) is the true probability of moving to state x' from state x after action a; the hat marks the model's estimate. Epsilon is the largest row error, half the summed absolute difference. V(x) is the true value of state x under a fixed behaviour and V-hat(x) the value the model computes, gamma the discount factor and R the largest immediate reward. The actual error is the largest gap between the two value recurrences after H rewards, and the finite bound R x epsilon x H(H-1)/2 is the chapter's finite-horizon envelope. The value scale R/(1-gamma) is the most any run could collect.
Predict first. In the default case (error 0.02, R = 1) raise gamma to 0.99. About how large is the discounted bound?
Choose an example
Scroll sideways for the whole figure
Constructed example: the laboratory's default, changed and transfer cases (kernels (0.90, 0.10), (0.20, 0.80) and model rows shifted by 0.02, horizons 5 and 20, identical kernels with rewards 1 and 2) and the chapter's 0.01 illustration, all computed with the laboratory's transition-model function.
Calculated values
- Epsilon (Equation 14.1)
- 0.020
- Actual error after H = 5
- 0.0868
- Finite bound at H = 5
- 0.200
- Discounted bound (Equation 14.2)
- 1.800
- Tightest valid bound
- 0.200
- Largest possible value R/(1-gamma)
- 10.0
Epsilon = 1/2 x (|0.88 - 0.90| + |0.12 - 0.10|) = 1/2 x (0.02 + 0.02) = 0.02. Finite bound = 1 x 0.02 x 5 x 4 / 2 = 0.200. Discounted bound = 0.90 x 0.02 x 1 / (1 - 0.90)^2 = 1.800. Check by hand after two rewards (state 2): true 1.00 + 0.90 x (0.20 x 0.00 + 0.80 x 1.00) = 1.720; model 1.00 + 0.90 x (0.22 x 0.00 + 0.78 x 1.00) = 1.702; |1.720 - 1.702| = 0.018. The two recurrences give 0.0868 after 5 (the left plot starts at H = 2 because H = 1 has error 0 and is not drawn). Each bound holds under its own assumptions; here the actual error 0.0868 is below both, 43 percent of the smaller one (finite bound 0.200).
Worked steps
- Row error by Equation (14.1): 1/2 x (|0.88 - 0.90| + |0.12 - 0.10|) = 1/2 x (0.02 + 0.02) = 0.02.
- Epsilon is the largest row error: 0.02. Rewards are shared; R = 1.
- Finite bound after H = 5: R x eps x H(H-1)/2 = 1 x 0.02 x 5 x 4 / 2 = 0.200.
- Discounted bound by Equation (14.2): gamma x eps x R / (1-gamma)^2 = 1.800.
- Each bound holds under its own assumptions; the smaller of the two numbers is 0.200.
- Two rewards by hand (state 2): V = 1.00 + 0.90 x (0.20 x 0.00 + 0.80 x 1.00) = 1.720, model 1.00 + 0.90 x (0.22 x 0.00 + 0.78 x 1.00) = 1.702, difference 0.018.
- Actual error after 5 rewards, from the two recurrences: 0.0868.
Use the idea
Before trusting a long imagined rollout, divide the bound by the largest possible value. If the share is near or above 1, the guarantee has nothing to say. If the model cannot be improved, consult the world more often.
Where the conclusion applies
The bound needs a fixed behaviour, the same expected immediate reward in model and world, rewards in [0, R], and error at most epsilon for every state and action that behaviour reaches. The kernels are the laboratory's constructed two-state defaults, so the actual error is that of one small model, not a general rate. A single number can also hide where the model is bad.
Common wrong turn: Reading the bound as the error the model actually makes
What this does not settle
The bound in Equation (14.2) is worst case and often loose, and a model with a large error may perform far better than it permits.
Chapter 14 source: "What this does not settle".
Check your understanding: With gamma = 0.8, epsilon = 0.05 and R = 1, what is the discounted bound, and what is the largest possible value?
Chapter 14 source: section "What one step of error becomes". Demonstration C14-D02.
Demonstration 3 of 4
The optimizer finds where the model is generous, and the gap decides
A model can be wrong about every action value and still choose correctly, or be right about eleven tools and choose the twelfth. What decides which?
A shared offset moves every value together, so the ordering is untouched however large the offset is. One favorable error on a tool that is truly poor changes the choice only if it lifts that tool above the best, but when it does, the average error is tiny (the error divided by twelve) while the largest is the whole error: this is why Equation (14.1) is a maximum and not an average. An adversarial error marks the best down and the runner-up up, so the two meet when twice the error equals the gap.
Scroll sideways for the whole equation
Q(x, a) is the true value of taking action a at state x and then following the fixed behaviour; here the twelve options are tools with constructed values. Delta is the gap between the best action's value and the runner-up's (a* is a best action; if several actions tie for the maximum, the gap is defined as zero). The model's value of each action is off by the chosen error, and the chapter's ranking guarantee is that the model keeps the true ranking whenever twice that error is smaller than the gap: 2 delta < Delta. The largest error is the maximum over options, the average is the mean.
Predict first. Use the adversarial pattern with error 0.15, where twice the error equals the gap 0.30. What does the model do with tools 5 and 6?
Choose an example
Scroll sideways for the whole figure
Constructed example: twelve option values defined for this reader (best 0.90, runner-up 0.60), with error patterns that illustrate the chapter's ranking argument and its eleven-right, one-wrong tool example.
Calculated values
- True gap (Equation 14.3)
- 0.30
- Largest error
- 0.050
- Average error over 12 tools
- 0.050
- Twice the largest error
- 0.10
- Guarantee 2 x error < gap holds
- yes
- Model's pick
- tool 6
- Model's value of the pick
- 0.95
- True value of the pick
- 0.90
Gap = 0.90 - 0.60 = 0.30. Twice the largest error = 2 x 0.05 = 0.10. Error pattern: the same offset on every tool. Average error = 0.60 / 12 = 0.050, largest error 0.05. Model values: tool 5 0.65, tool 6 0.95, tool 12 0.15. The model still picks tool 6, and the guarantee holds: 2 x 0.05 = 0.10 is below the gap 0.30.
Worked steps
- True values: tool 6 is best at 0.90, tool 5 is next at 0.60, tool 12 is lowest at 0.10.
- Gap by Equation (14.3): 0.90 - 0.60 = 0.30.
- Error pattern: the same offset on every tool, size 0.05.
- Model values: tool 5 0.65, tool 6 0.95, tool 12 0.15.
- Largest error 0.05; average error 0.60 / 12 = 0.050.
- Guarantee test: 2 x 0.05 = 0.10 against the gap 0.30: holds.
- The optimizer takes the highest model value: tool 6, true value 0.90.
Use the idea
Judge a planning model against the decision it serves, not against a single accuracy figure. Ask how large the gap is between the best option and the next one, compare twice the value error to it in the same units, and check whether the errors are concentrated on options the optimizer will rank.
Where the conclusion applies
Twelve options with constructed values, a uniform bound on the action-value error, and three hand-made error patterns. The chapter adds that this bound is an extra assumption: Equation (14.2) bounds state values and does not supply it. The test 2 x error < gap is sufficient, not necessary, so a failed test leaves the ranking undecided rather than wrong.
Common wrong turn: Judging a planning model by its average accuracy
What this does not settle
This schematic explains the selection mechanism rather than supplying a measured error landscape.
Chapter 14 source: "This schematic explains the selection mechanism rather than supplying a measured error landscape.".
Check your understanding: True values 0.90 and 0.60, adversarial error 0.12. Which action does the model pick, and does the guarantee hold?
Chapter 14 source: section "The number that actually matters". Demonstration C14-D03.
Demonstration 4 of 4
How many imagined steps are certified
Given the model error, the reward range and the action gap, how many steps ahead can the certificate vouch for?
Twice the error bound grows with H times H-1, so each extra step costs more than the last. The certificate keeps the largest H for which that quantity is still strictly below the gap. Because H(H-1) grows like H squared, the horizon scales roughly with the square root of gap over R x epsilon: for long horizons, cutting the error by 8 times raises it by about 2.8 times, not 8. A twenty-step plan committed in advance fails the test by a wide margin, which is why the countermeasure is to plan a few steps, look, and replan.
Scroll sideways for the whole equation
H is the number of rewards the agent plans over, H-max the declared planning cap (9 here), and H-star the largest H that passes. R is the largest immediate reward, epsilon the uniform one-step error from Equation (14.1), and Delta the lower bound on the action gap from Equation (14.3). The left side R x epsilon x H(H-1) is twice the finite-horizon action-value error bound R x epsilon x H(H-1)/2.
Predict first. With R = 10 and a gap of 2, cut the error by 8 times, from 0.08 to 0.01. Does the certified horizon grow by 8 times, or by far less?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's Figure 14.2 gaps (R = 1, epsilon 0.02, gaps 0.2 and 0.5) and its priced twenty-step plan (R = 10, gap 2, epsilon 0.08 and 0.01); the other error values are defined for this reader, and each curve is checked against the laboratory's finite-horizon bound.
Calculated values
- Certified horizon H*
- 5
- First horizon that fails
- 6
- R x epsilon x H(H-1) at H* = 5
- 0.40
- Action gap
- 0.5
- Same quantity at a twenty-step plan
- 7.6
At H = 5: 1 x 0.02 x 5 x 4 = 0.40, below the gap 0.5. At H = 6: 1 x 0.02 x 6 x 5 = 0.60, not below the gap 0.5. The certificate covers 5 rewards. Its failure at 6 is inconclusive: it does not show that the ranking reverses, only that this guarantee stops. A twenty-step plan committed in advance would need 1 x 0.02 x 20 x 19 = 7.6 to be below the gap 0.5, which it is not: plan 5, look, and replan.
Worked steps
- Inputs: R = 1, epsilon = 0.02, action gap = 0.5, planning cap 9.
- Test each H: R x epsilon x H(H-1) must be strictly below 0.5.
- Largest H that passes: H* = 5, because 1 x 0.02 x 5 x 4 = 0.40.
- First failure at H = 6: 1 x 0.02 x 6 x 5 = 0.60.
- Twenty steps: 1 x 0.02 x 20 x 19 = 7.6, far above the gap.
Use the idea
Use the certified horizon to set how many steps an agent commits to before it observes the world again. Plan that many, execute a few, look, and replan, rather than emitting a long plan against an error nobody measured.
Where the conclusion applies
A finite covered domain, equal expected immediate rewards in model and world, the same initial state and fixed behaviour, uniform error epsilon, zero terminal continuation and a declared gap. The R = 10 case uses the chapter's priced plan, where a logged prediction miss rate is declared a conservative proxy for epsilon, which is not the same thing. A failed test is inconclusive, not proof of a wrong decision, and the inputs are usually not measured.
Common wrong turn: A failed certificate proves the ranking reverses
What this does not settle
The horizon arithmetic is constructed for teaching. How to bound what a system could achieve given a model, rather than what it does achieve, waits for Chapters 24 and 26.
Chapter 14 source: "What this does not settle".
Check your understanding: With R = 1, epsilon = 0.04 and a gap of 0.5, what is the certified horizon?
Chapter 14 source: section "How many steps a model is good for". Demonstration C14-D04.