Demonstration 1 of 4
Regret: what never looking again costs
How fast does the shortfall against an all-knowing agent grow under the rules that always exploit, always rotate, explore with optimism or sample from a belief?
Taking the worse arm every time forgoes the gap on every task, so regret is the gap times T, a straight line; rotating evenly forgoes the average gap, a straight line half as steep for two arms. The optimistic rule and posterior sampling keep trying a thin-record arm only while it is still uncertain, so their regret flattens relative to the lines. The right panel shows why: regret is the sum of gap x pulls, so it is made only of pulls of arms that are not best. A seeded run is not an average: always taking the current best can look best when its first pulls happen to find the best arm.
Scroll sideways for the whole equation
T is the number of tasks. mu-star is the success probability of the best arm (tool). mu of a_t is the success probability of the arm chosen at task t, and E is the average over luck. Regret is the total successes forgone compared with always using the best arm; for a run it equals the sum, over arms, of the gap to the best mean times the pulls of that arm. Worlds: the book's tools (A 0.78, B 0.90, 1,000 tasks), the notebook's default (0.4 and 0.7, 120 tasks), its swapped case (0.7 and 0.4) and its three-arm case (0.2, 0.5 and 0.8, 90 tasks). Each curve is one seeded run.
Predict first. In the book's world, go from the run of 1,000 tasks to the ten-times-longer run of 10,000. By what factor does the regret of being stuck on the worse tool grow, and does the optimistic rule's seeded regret grow by more or less?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's invented success rates 0.78 and 0.90 and its 1,000-task comparison, and the notebook's default, swapped and three-arm worlds; every curve is one seeded run computed with the laboratory's bandit function.
Calculated values
- Stuck on the runner-up arm
- 120.00
- Rotate evenly
- 60.00
- Always the current best (this run)
- 0.24
- Optimistic rule (this run)
- 29.64
- Posterior sampling (this run)
- 20.88
- Always the current best, seeds above half the stuck line (of 200)
- 45
- Mean regret over 200 seeds, always the current best
- 28.05
- Mean regret over 200 seeds, optimistic rule
- 26.10
- Mean regret over 200 seeds, posterior sampling
- 6.31
- Chance the chapter's controller draws two failures in a row from the best arm
- 0.01
- Success rate a stuck agent would report
- 0.78
- Best success rate available
- 0.90
Regret of a run is the sum of gap x pulls: always the current best: 0.12 x 2 = 0.24; optimistic rule: 0.12 x 247 = 29.64; posterior sampling: 0.12 x 174 = 20.88. Stuck on the runner-up arm for 1,000 tasks: 0.12 x 1,000 = 120.00. Rotating evenly: 1,000 x (0.90 - 0.84) = 60.00. The best arm's first two pulls both fail with chance (1 - 0.90) x (1 - 0.90) = 0.01. A stuck agent would report a success rate of 0.78 and look stable, while 0.90 was available. In this one seeded run the always-the-current-best rule pulled the arms that are not best only 2 times, so it beat the optimistic rule. That is the luck of one seed, not a ranking: the danger the chapter names is that an unlucky first pull can lock the best arm out for good, and the stuck line prices that case. Over seeds 0 to 199 of the same laboratory function, the mean regret is 28.05 for always the current best, 26.10 for the optimistic rule and 6.31 for posterior sampling, and always the current best ends above half of the stuck line in 45 of 200 seeds (the best arm was locked out early), so one seeded run is a thin guide to this rule.
Worked steps
- Gaps: best mean minus each arm's mean, arm A 0.12, arm B 0.00.
- Pulls of each arm in this run: always the current best 2, 998; optimistic rule 247, 753; posterior sampling 174, 826.
- Regret = sum of gap x pulls. Optimistic rule: 0.12 x 247 = 29.64.
- Posterior sampling: 0.12 x 174 = 20.88.
- Stuck line: 0.12 x 1,000 = 120.00.
- The chapter's controller drew two failures in a row from the best arm: 0.10 x 0.10 = 0.01.
- Over seeds 0 to 199, mean regret: current best 28.05, optimistic 26.10, posterior 6.31; current best ends above half the stuck line in 45 of 200 seeds.
Use the idea
When a team reports a success rate for a deployed agent, ask how much of its action set it has sampled recently. Regret itself needs the true means, which a deployed agent does not have, so this is the question that can be asked. An agent that took one action thousands of times and the alternatives twice may have stopped measuring.
Where the conclusion applies
Independent fixed success probabilities and one seeded run per rule, so each curve is one draw of luck: it is not a policy performance estimate and not the expectation in Equation (11.1). Regret here is pseudo-regret, computed from the known means, which a deployed agent does not have. The stuck and rotate lines are closed forms. Real tools change over time, which this model does not cover.
Common wrong turn: A stable, respectable success rate means nothing better was available
What this does not settle
The setting assumes actions do not change the world and payoffs are drawn independently, and neither holds for a deployed agent.
Chapter 11 source: "What this does not settle".
Check your understanding: If tool A succeeds 0.78 and tool B 0.90, how much regret does an agent that rotates evenly build up over 500 tasks?
Chapter 11 source: section "What is actually being lost". Demonstration C11-D01.
Demonstration 2 of 4
Optimism: why a thin record gets another call
Why does a tool with a poor record and few pulls get called again, and what changes that?
The bonus grows slowly with the clock and shrinks as a tool is pulled, so a tool that is left alone gains bonus until it wins a call. In the trapped controller, B's two pulls give it a bonus of 1.731 against A's 0.577, enough to overturn a mean gap of 0.778. A smaller c shrinks every bonus: at c = 0.5 the same controller keeps calling A. If B then succeeds its estimate rises, and if it fails its count rises and its bonus shrinks; either way it is being updated again.
Scroll sideways for the whole equation
t is the number of completed pulls. N_t(a) is how many of those pulls went to tool a. mu-hat of a is tool a's average reward so far. c is the confidence coefficient; the book uses the square root of two, about 1.414, for rewards between 0 and 1. ln is the natural logarithm. The index is the mean plus the bonus c x sqrt(ln t / N_t(a)), and the tool with the larger index is called. Situations: pull 4 or pull 5 of the chapter's ten-pull table; the trapped controller after 20 tasks (A has 18 pulls and 14 successes, B has 2 pulls and none); the workbench problem (A has 80 pulls and mean 0.70, B has 20 pulls and mean 0.60).
Predict first. Pick pull 5 of the table with c = 1.414 (square root of 2). Tool B has a mean of 0.000 and tool A has 0.667. Does the rule call A or B?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's ten-pull table with c equal to the square root of two (invented teaching rewards), the chapter's trapped controller (18 pulls with 14 successes, 2 pulls with none) and the original workbench problem III.1 (80 and 20 pulls); the values of c other than the square root of two are defined for this reader.
Calculated values
- Counts A, B
- 2, 1
- Index A
- 1.548
- Index B
- 1.482
- Selected
- A
- Selected by mean alone
- A
- Reward seen after the pull
- 1
With 3 completed pulls, ln 3 = 1.099. A: bonus = 1.414 x sqrt(1.099 / 2) = 1.048, index = 0.500 + 1.048 = 1.548. B: bonus = 1.414 x sqrt(1.099 / 1) = 1.482, index = 0.000 + 1.482 = 1.482. Tool A has the larger index by 0.066, so Equation (11.2) selects it. Its pull number 4 then returns 1. Ranking by mean alone would also select A, so here the bonus does not change the call. A tool with few pulls keeps a large bonus even when its mean is low.
Worked steps
- Completed pulls t = 3, so ln t = ln 3 = 1.099.
- A: bonus = 1.414 x sqrt(1.099 / 2) = 1.048, index = 0.500 + 1.048 = 1.548.
- B: bonus = 1.414 x sqrt(1.099 / 1) = 1.482, index = 0.000 + 1.482 = 1.482.
- Tool A has the larger index by 0.066, so Equation (11.2) selects it. Its pull number 4 then returns 1. Ranking by mean alone would also select A, so here the bonus does not change the call.
Use the idea
When a controller must choose among components with thin records, an index of this shape gives a written reason for each call, which can be audited, instead of a vague claim that the system is curious.
Where the conclusion applies
The table's rewards are the book's invented alternating sequence (A returns 1, 0, 1, 0, and so on; B returns 0, 1, 0, 1), not a test of the stochastic guarantee. Each tool is pulled once first, ties go to A, and the bonus uses unrounded values. Changing c changes the rule; the book's guarantee is stated for c equal to the square root of two with rewards in [0, 1]. The trapped controller and the workbench problem are single decisions given counts and means, not runs.
Common wrong turn: A low estimate is a reason never to look again
What this does not settle
Equation (11.2) requires a constant whose right value depends on the payoff scale.
Chapter 11 source: "What this does not settle".
Check your understanding: At pull 5 of the table the counts are 3 and 1 after 4 completed pulls. With c = 1.414, what is B's bonus?
Chapter 11 source: section "A proved rule and ten constructed pulls". Demonstration C11-D02.
Demonstration 3 of 4
Information gain is not decision value
Can a check remove uncertainty about a claim and still be worth exactly nothing?
Entropy falls whenever the report is correlated with the claim, so the gain is positive. The value in Equation (11.4) is positive only if some report pushes the release score across the fixed value 92 of the alternative. When the reports leave the release score on the same side of 92 as the no-check score, the same action wins either way and VOI is zero, as at accuracy 0.60 with prior 0.85. Where that line falls depends on the prior: accuracy 0.70 is enough at prior 0.85 but not at prior 0.97, and at prior 0.50 even a strong check is needed to lift the score above 92.
Scroll sideways for the whole equation
Y is the proposition the decision turns on: every claim in the draft is supported. O is the report of a check. H is entropy in bits, a measure of uncertainty (1 bit for a fair coin, 0 when certain). I(Y; O) is the uncertainty about Y that O removes. VOI is the decision value: the average best expected utility with the report, minus the best without it. Releasing scores 100 if supported and 0 if not; requesting evidence scores 92.
Predict first. With prior 0.85 and a check that is right 0.60 of the time, does the check remove uncertainty, and does it change what the agent does?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's release score 85 and evidence score 92 (Table 6.1 values); the checks, their accuracies and the priors are defined for this reader.
Calculated values
- Uncertainty before, H(Y)
- 0.610 bits
- Uncertainty after, H(Y | O)
- 0.401 bits
- Information gain I(Y; O)
- 0.209 bits
- Decision value VOI
- 3.71
- Best action changes
- yes
Best action without the check = max(100 x 0.85, 92) = max(85.00, 92.00) = 92.00. supported report (chance 0.7450, posterior release score 96.98): weighted by that chance, max(100 x 0.7225, 92 x 0.7450) = max(72.25, 68.54) = 72.25; unsupported report (chance 0.2550, posterior release score 50.00): weighted by that chance, max(100 x 0.1275, 92 x 0.2550) = max(12.75, 23.46) = 23.46. VOI = 95.71 - 92.00 = 3.71. The report moves the best action in at least one case, so the check is worth something before its cost.
Worked steps
- Before the check: H(Y) = 0.610 bits; best action now = max(85.00, 92) = 92.00.
- Report supported (chance 0.7450): posterior 0.9698, release scores 96.98 against 92.
- Report unsupported (chance 0.2550): posterior 0.5000, release scores 50.00 against 92.
- After the check: H(Y | O) = 0.401 bits, so the gain is 0.610 - 0.401 = 0.209 bits.
- Value: 95.71 - 92.00 = 3.71.
Use the idea
Before paying for a check, write down which action follows each possible report. If it is the same action every time, the check is decoration, however interesting its output.
Where the conclusion applies
A symmetric check (equally accurate on supported and unsupported claims), a single decision between releasing and requesting evidence, and a belief that the stated prior and accuracy are correct. Accuracy and prior here are constructed; a real check may be biased in one direction, which this model does not cover.
Common wrong turn: Resolving uncertainty is the goal
What this does not settle
Equation (11.4) prices a single observation against a fixed decision, not a sequence of them. All numbers in the opening construction are invented for teaching.
Chapter 11 source: "What this does not settle".
Check your understanding: With prior 0.85 and a check that is right 0.70 of the time, a report of unsupported leaves the release score at 70.8. Does that report change the action?
Chapter 11 source: section "Uncertainty is not the same as value". Demonstration C11-D03.
Demonstration 4 of 4
A worked price for one observation
How much is one call to a classifier worth, and how can the same classifier be worth zero in one decision and something in another?
With no shift, evidence wins under both reports and the value is exactly zero. A shift of 10 pushes the release score above 92 after clear but not after flag, so the best action now depends on the report. From a shift of 10 on, releasing already beats requesting evidence before observing (95 at shift 10), so the call can only rescue the flag branch. At a shift of 15 that branch still scores 90 against 92, so the call adds only 1/3 x (92 - 90) = 0.667. At a call cost of exactly 7/3 the shift-10 call ties with skipping it.
Scroll sideways for the whole equation
A classifier reports flag with probability one third and clear with probability two thirds. Releasing scores 75 after flag and 90 after clear, plus the shift you choose; requesting evidence scores 92 either way. VOI is the average best score with the report minus the best score without it. The call cost is subtracted to give the net value. The book's price is 7/3, about 2.333.
Predict first. Add 10 to the release scores. Does the classifier become worth buying at a call cost of 2?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's worked classifier (flag one third, release scores 75 and 90, evidence 92, shift 10 giving 7/3); the other shifts and costs are defined for this reader.
Calculated values
- Release now, no observation
- 95.000
- Best action without the call
- 95.000
- Best expected utility with the call
- 97.333
- Gross value
- 2.333
- Net value
- 2.333
- Worth buying
- yes
Release scores 85 after flag and 100 after clear, so before observing it scores 1/3 x 85 + 2/3 x 100 = 95.000 against 92 for requesting evidence. With the call the best choices are 92 after flag and 100 after clear, so 1/3 x 92 + 2/3 x 100 = 97.333. Gross value = 97.333 - 95.000 = 2.333; net value = 2.333 - 0.0 = 2.333. Net value is positive, so the call is worth buying for this one decision. The classifier itself never changed between states; only the decision it feeds did.
Worked steps
- Release scores 85 after flag and 100 after clear; requesting evidence scores 92 either way.
- Without the call: 1/3 x 85 + 2/3 x 100 = 95.000, best = 95.000.
- With the call: 1/3 x 92 + 2/3 x 100 = 97.333.
- Gross value = 97.333 - 95.000 = 2.333.
- Net value = 2.333 - 0.0 = 2.333.
Use the idea
Price a verification step against the specific decision it feeds. The same step, with the same accuracy and cost, can be a bargain for one decision and worthless for another. Before spending any budget, also check whether the system already holds the answer in a log it failed to structure: free information should always be used.
Where the conclusion applies
One decision, one observation, a coherent probability model, and a gross value that ignores any later decisions the observation could also inform. The conditional release scores are the book's constructed values; a real classifier's scores would have to be estimated, and a wrong estimate can flip the verdict.
Common wrong turn: A call that returns interesting material was worth making
What this does not settle
Equation (11.4) prices a single observation against a fixed decision, not a sequence of them.
Chapter 11 source: "What this does not settle".
Check your understanding: With a shift of 5, what is the gross value, and is the call worth buying at cost 2?
Chapter 11 source: section "A worked price for one observation". Demonstration C11-D04.