The Mathematics of AI Agents, laboratory reader ยท Chapter 11

The Mathematics of Curiosity

Choosing a tool decides which evidence you get about it. Curiosity becomes a costed allocation rule, and information is worth only what it can change.

These four demonstrations follow the chapter's document controller and its two retrieval tools. First the cost of never looking again is measured as regret for three rules in four worlds, then an optimistic rule is traced pull by pull, then information gain is separated from decision value, and finally one observation is priced.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

Regret: what never looking again costs

How fast does the shortfall against an all-knowing agent grow under the rules that always exploit, always rotate, explore with optimism or sample from a belief?

Taking the worse arm every time forgoes the gap on every task, so regret is the gap times T, a straight line; rotating evenly forgoes the average gap, a straight line half as steep for two arms. The optimistic rule and posterior sampling keep trying a thin-record arm only while it is still uncertain, so their regret flattens relative to the lines. The right panel shows why: regret is the sum of gap x pulls, so it is made only of pulls of arms that are not best. A seeded run is not an average: always taking the current best can look best when its first pulls happen to find the best arm.

Equation (11.1), written in LaTeX: \operatorname{Reg}(T)=T\mu^{\star}-\operatorname{E}[\sum_{t=0}^{T-1}\mu_{a_t}]

Scroll sideways for the whole equation

T is the number of tasks. mu-star is the success probability of the best arm (tool). mu of a_t is the success probability of the arm chosen at task t, and E is the average over luck. Regret is the total successes forgone compared with always using the best arm; for a run it equals the sum, over arms, of the gap to the best mean times the pulls of that arm. Worlds: the book's tools (A 0.78, B 0.90, 1,000 tasks), the notebook's default (0.4 and 0.7, 120 tasks), its swapped case (0.7 and 0.4) and its three-arm case (0.2, 0.5 and 0.8, 90 tasks). Each curve is one seeded run.

Predict first. In the book's world, go from the run of 1,000 tasks to the ten-times-longer run of 10,000. By what factor does the regret of being stuck on the worse tool grow, and does the optimistic rule's seeded regret grow by more or less?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Regret: what never looking again costs. Left: accumulated regret over 1,000 tasks for five lines: a dashed straight line for being stuck on the runner-up arm ending at 120.0, a dotted line for rotating evenly ending at 60.0, and one seeded run each of always the current best (0.2), the optimistic rule (29.6) and posterior sampling (20.9). Right: stacked bars of how many pulls each rule gave each arm.
World: Book tools: A 0.78, B 0.90, Length of the run: The world's own run (book 1,000; notebook 120; three-arm 90 tasks)
Constructed example: the book's invented success rates 0.78 and 0.90 and its 1,000-task comparison, and the notebook's default, swapped and three-arm worlds; every curve is one seeded run computed with the laboratory's bandit function.

Calculated values

Stuck on the runner-up arm
120.00
Rotate evenly
60.00
Always the current best (this run)
0.24
Optimistic rule (this run)
29.64
Posterior sampling (this run)
20.88
Always the current best, seeds above half the stuck line (of 200)
45
Mean regret over 200 seeds, always the current best
28.05
Mean regret over 200 seeds, optimistic rule
26.10
Mean regret over 200 seeds, posterior sampling
6.31
Chance the chapter's controller draws two failures in a row from the best arm
0.01
Success rate a stuck agent would report
0.78
Best success rate available
0.90

Regret of a run is the sum of gap x pulls: always the current best: 0.12 x 2 = 0.24; optimistic rule: 0.12 x 247 = 29.64; posterior sampling: 0.12 x 174 = 20.88. Stuck on the runner-up arm for 1,000 tasks: 0.12 x 1,000 = 120.00. Rotating evenly: 1,000 x (0.90 - 0.84) = 60.00. The best arm's first two pulls both fail with chance (1 - 0.90) x (1 - 0.90) = 0.01. A stuck agent would report a success rate of 0.78 and look stable, while 0.90 was available. In this one seeded run the always-the-current-best rule pulled the arms that are not best only 2 times, so it beat the optimistic rule. That is the luck of one seed, not a ranking: the danger the chapter names is that an unlucky first pull can lock the best arm out for good, and the stuck line prices that case. Over seeds 0 to 199 of the same laboratory function, the mean regret is 28.05 for always the current best, 26.10 for the optimistic rule and 6.31 for posterior sampling, and always the current best ends above half of the stuck line in 45 of 200 seeds (the best arm was locked out early), so one seeded run is a thin guide to this rule.

Worked steps

  1. Gaps: best mean minus each arm's mean, arm A 0.12, arm B 0.00.
  2. Pulls of each arm in this run: always the current best 2, 998; optimistic rule 247, 753; posterior sampling 174, 826.
  3. Regret = sum of gap x pulls. Optimistic rule: 0.12 x 247 = 29.64.
  4. Posterior sampling: 0.12 x 174 = 20.88.
  5. Stuck line: 0.12 x 1,000 = 120.00.
  6. The chapter's controller drew two failures in a row from the best arm: 0.10 x 0.10 = 0.01.
  7. Over seeds 0 to 199, mean regret: current best 28.05, optimistic 26.10, posterior 6.31; current best ends above half the stuck line in 45 of 200 seeds.

Use the idea

When a team reports a success rate for a deployed agent, ask how much of its action set it has sampled recently. Regret itself needs the true means, which a deployed agent does not have, so this is the question that can be asked. An agent that took one action thousands of times and the alternatives twice may have stopped measuring.

Where the conclusion applies

Independent fixed success probabilities and one seeded run per rule, so each curve is one draw of luck: it is not a policy performance estimate and not the expectation in Equation (11.1). Regret here is pseudo-regret, computed from the known means, which a deployed agent does not have. The stuck and rotate lines are closed forms. Real tools change over time, which this model does not cover.

Common wrong turn: A stable, respectable success rate means nothing better was available
The chapter says under-exploration is invisible in an evaluation: a trapped controller reports 78 percent, its logs show consistent behavior, and an evaluation reports what a system did, not what a differently behaving version would have done. The metric 'success rate a stuck agent would report' sits beside 'best success rate available' for that reason.
What this does not settle

The setting assumes actions do not change the world and payoffs are drawn independently, and neither holds for a deployed agent.

Chapter 11 source: "What this does not settle".

Check your understanding: If tool A succeeds 0.78 and tool B 0.90, how much regret does an agent that rotates evenly build up over 500 tasks?
Half the tasks go to A and cost 0.12 each: 250 x 0.12 = 30. Equivalently 500 x 0.90 - (250 x 0.78 + 250 x 0.90) = 450 - 420 = 30.

Chapter 11 source: section "What is actually being lost". Demonstration C11-D01.

Demonstration 2 of 4

Optimism: why a thin record gets another call

Why does a tool with a poor record and few pulls get called again, and what changes that?

The bonus grows slowly with the clock and shrinks as a tool is pulled, so a tool that is left alone gains bonus until it wins a call. In the trapped controller, B's two pulls give it a bonus of 1.731 against A's 0.577, enough to overturn a mean gap of 0.778. A smaller c shrinks every bonus: at c = 0.5 the same controller keeps calling A. If B then succeeds its estimate rises, and if it fails its count rises and its bonus shrinks; either way it is being updated again.

Equation (11.2), written in LaTeX: a_t=\arg\max_{a\in\mathcal{A}}(\hat\mu_a+c\sqrt{\frac{\ln t}{N_t(a)}})

Scroll sideways for the whole equation

t is the number of completed pulls. N_t(a) is how many of those pulls went to tool a. mu-hat of a is tool a's average reward so far. c is the confidence coefficient; the book uses the square root of two, about 1.414, for rewards between 0 and 1. ln is the natural logarithm. The index is the mean plus the bonus c x sqrt(ln t / N_t(a)), and the tool with the larger index is called. Situations: pull 4 or pull 5 of the chapter's ten-pull table; the trapped controller after 20 tasks (A has 18 pulls and 14 successes, B has 2 pulls and none); the workbench problem (A has 80 pulls and mean 0.70, B has 20 pulls and mean 0.60).

Predict first. Pick pull 5 of the table with c = 1.414 (square root of 2). Tool B has a mean of 0.000 and tool A has 0.667. Does the rule call A or B?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Optimism: why a thin record gets another call. Two stacked bars for the ten-pull table at pull 4: mean (solid) plus bonus (hatched) for A and B. Indices 1.548 for A and 1.482 for B; A is selected.
Situation: Ten-pull table, pull 4, Confidence coefficient c: square root of 2 (about 1.414)
Constructed example: the chapter's ten-pull table with c equal to the square root of two (invented teaching rewards), the chapter's trapped controller (18 pulls with 14 successes, 2 pulls with none) and the original workbench problem III.1 (80 and 20 pulls); the values of c other than the square root of two are defined for this reader.

Calculated values

Counts A, B
2, 1
Index A
1.548
Index B
1.482
Selected
A
Selected by mean alone
A
Reward seen after the pull
1

With 3 completed pulls, ln 3 = 1.099. A: bonus = 1.414 x sqrt(1.099 / 2) = 1.048, index = 0.500 + 1.048 = 1.548. B: bonus = 1.414 x sqrt(1.099 / 1) = 1.482, index = 0.000 + 1.482 = 1.482. Tool A has the larger index by 0.066, so Equation (11.2) selects it. Its pull number 4 then returns 1. Ranking by mean alone would also select A, so here the bonus does not change the call. A tool with few pulls keeps a large bonus even when its mean is low.

Worked steps

  1. Completed pulls t = 3, so ln t = ln 3 = 1.099.
  2. A: bonus = 1.414 x sqrt(1.099 / 2) = 1.048, index = 0.500 + 1.048 = 1.548.
  3. B: bonus = 1.414 x sqrt(1.099 / 1) = 1.482, index = 0.000 + 1.482 = 1.482.
  4. Tool A has the larger index by 0.066, so Equation (11.2) selects it. Its pull number 4 then returns 1. Ranking by mean alone would also select A, so here the bonus does not change the call.

Use the idea

When a controller must choose among components with thin records, an index of this shape gives a written reason for each call, which can be audited, instead of a vague claim that the system is curious.

Where the conclusion applies

The table's rewards are the book's invented alternating sequence (A returns 1, 0, 1, 0, and so on; B returns 0, 1, 0, 1), not a test of the stochastic guarantee. Each tool is pulled once first, ties go to A, and the bonus uses unrounded values. Changing c changes the rule; the book's guarantee is stated for c equal to the square root of two with rewards in [0, 1]. The trapped controller and the workbench problem are single decisions given counts and means, not runs.

Common wrong turn: A low estimate is a reason never to look again
The chapter's trapped controller drew two failures from a nine-in-ten process, and its rule contained no reason to look again: once a tool's estimate falls low enough to stop being selected, nothing will update it. The failure is in the loop from estimates to choices to data, not in the arithmetic of the estimate.
What this does not settle

Equation (11.2) requires a constant whose right value depends on the payoff scale.

Chapter 11 source: "What this does not settle".

Check your understanding: At pull 5 of the table the counts are 3 and 1 after 4 completed pulls. With c = 1.414, what is B's bonus?
ln 4 = 1.386, so the bonus is 1.414 x sqrt(1.386 / 1) = 1.414 x 1.177 = 1.665 (rounded).

Chapter 11 source: section "A proved rule and ten constructed pulls". Demonstration C11-D02.

Demonstration 3 of 4

Information gain is not decision value

Can a check remove uncertainty about a claim and still be worth exactly nothing?

Entropy falls whenever the report is correlated with the claim, so the gain is positive. The value in Equation (11.4) is positive only if some report pushes the release score across the fixed value 92 of the alternative. When the reports leave the release score on the same side of 92 as the no-check score, the same action wins either way and VOI is zero, as at accuracy 0.60 with prior 0.85. Where that line falls depends on the prior: accuracy 0.70 is enough at prior 0.85 but not at prior 0.97, and at prior 0.50 even a strong check is needed to lift the score above 92.

Equation (11.3), written in LaTeX: \operatorname{Gain}(O)=I(Y;O)=H(Y)-H(Y\mid O)

Equation (11.4), written in LaTeX: \operatorname{VOI}(O)=\operatorname{E}_{O}[\max_{a}\operatorname{EU}(a\mid O)]-\max_{a}\operatorname{EU}(a)

Scroll sideways for the whole equation

Y is the proposition the decision turns on: every claim in the draft is supported. O is the report of a check. H is entropy in bits, a measure of uncertainty (1 bit for a fair coin, 0 when certain). I(Y; O) is the uncertainty about Y that O removes. VOI is the decision value: the average best expected utility with the report, minus the best without it. Releasing scores 100 if supported and 0 if not; requesting evidence scores 92.

Predict first. With prior 0.85 and a check that is right 0.60 of the time, does the check remove uncertainty, and does it change what the agent does?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Information gain is not decision value. Left: entropy falls from 0.61 bits before the check to 0.40 bits after it. Right: two bars give the release score after each report, 97.0 after supported and 50.0 after unsupported, against a dashed line at 92 for requesting evidence. The best action changes.
How often the check is right: 0.85, Prior chance the claims are supported: 0.85
Constructed example: the book's release score 85 and evidence score 92 (Table 6.1 values); the checks, their accuracies and the priors are defined for this reader.

Calculated values

Uncertainty before, H(Y)
0.610 bits
Uncertainty after, H(Y | O)
0.401 bits
Information gain I(Y; O)
0.209 bits
Decision value VOI
3.71
Best action changes
yes

Best action without the check = max(100 x 0.85, 92) = max(85.00, 92.00) = 92.00. supported report (chance 0.7450, posterior release score 96.98): weighted by that chance, max(100 x 0.7225, 92 x 0.7450) = max(72.25, 68.54) = 72.25; unsupported report (chance 0.2550, posterior release score 50.00): weighted by that chance, max(100 x 0.1275, 92 x 0.2550) = max(12.75, 23.46) = 23.46. VOI = 95.71 - 92.00 = 3.71. The report moves the best action in at least one case, so the check is worth something before its cost.

Worked steps

  1. Before the check: H(Y) = 0.610 bits; best action now = max(85.00, 92) = 92.00.
  2. Report supported (chance 0.7450): posterior 0.9698, release scores 96.98 against 92.
  3. Report unsupported (chance 0.2550): posterior 0.5000, release scores 50.00 against 92.
  4. After the check: H(Y | O) = 0.401 bits, so the gain is 0.610 - 0.401 = 0.209 bits.
  5. Value: 95.71 - 92.00 = 3.71.

Use the idea

Before paying for a check, write down which action follows each possible report. If it is the same action every time, the check is decoration, however interesting its output.

Where the conclusion applies

A symmetric check (equally accurate on supported and unsupported claims), a single decision between releasing and requesting evidence, and a belief that the stated prior and accuracy are correct. Accuracy and prior here are constructed; a real check may be biased in one direction, which this model does not cover.

Common wrong turn: Resolving uncertainty is the goal
The chapter says a rule that seeks uncertainty as such will happily spend a budget resolving the number of words in a passage, which is uncertain, cheaply resolved and irrelevant. What the agent wants is uncertainty reduction about the quantity the decision turns on, and information has value only when it could change what the agent does.
What this does not settle

Equation (11.4) prices a single observation against a fixed decision, not a sequence of them. All numbers in the opening construction are invented for teaching.

Chapter 11 source: "What this does not settle".

Check your understanding: With prior 0.85 and a check that is right 0.70 of the time, a report of unsupported leaves the release score at 70.8. Does that report change the action?
P(supported and report unsupported) = 0.85 x 0.30 = 0.255 and P(report unsupported) = 0.36, so the release score is 100 x 0.255 / 0.36 = 70.8. That is below 92, which is also the action chosen before the check, so this report changes nothing; only the supported report (score 92.97) does.

Chapter 11 source: section "Uncertainty is not the same as value". Demonstration C11-D03.

Demonstration 4 of 4

A worked price for one observation

How much is one call to a classifier worth, and how can the same classifier be worth zero in one decision and something in another?

With no shift, evidence wins under both reports and the value is exactly zero. A shift of 10 pushes the release score above 92 after clear but not after flag, so the best action now depends on the report. From a shift of 10 on, releasing already beats requesting evidence before observing (95 at shift 10), so the call can only rescue the flag branch. At a shift of 15 that branch still scores 90 against 92, so the call adds only 1/3 x (92 - 90) = 0.667. At a call cost of exactly 7/3 the shift-10 call ties with skipping it.

Equation (11.4), written in LaTeX: \operatorname{VOI}(O)=\operatorname{E}_{O}[\max_{a}\operatorname{EU}(a\mid O)]-\max_{a}\operatorname{EU}(a)

Scroll sideways for the whole equation

A classifier reports flag with probability one third and clear with probability two thirds. Releasing scores 75 after flag and 90 after clear, plus the shift you choose; requesting evidence scores 92 either way. VOI is the average best score with the report minus the best score without it. The call cost is subtracted to give the net value. The book's price is 7/3, about 2.333.

Predict first. Add 10 to the release scores. Does the classifier become worth buying at a call cost of 2?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A worked price for one observation. Left: for each of three moments (after flag, after clear, before observing) a release bar and an evidence bar of 92. Right: bars for gross value 2.33 and net value 2.33 at call cost 0.0. Worth buying: yes.
Added to both release scores: 10, Cost of the call: 0
Constructed example: the book's worked classifier (flag one third, release scores 75 and 90, evidence 92, shift 10 giving 7/3); the other shifts and costs are defined for this reader.

Calculated values

Release now, no observation
95.000
Best action without the call
95.000
Best expected utility with the call
97.333
Gross value
2.333
Net value
2.333
Worth buying
yes

Release scores 85 after flag and 100 after clear, so before observing it scores 1/3 x 85 + 2/3 x 100 = 95.000 against 92 for requesting evidence. With the call the best choices are 92 after flag and 100 after clear, so 1/3 x 92 + 2/3 x 100 = 97.333. Gross value = 97.333 - 95.000 = 2.333; net value = 2.333 - 0.0 = 2.333. Net value is positive, so the call is worth buying for this one decision. The classifier itself never changed between states; only the decision it feeds did.

Worked steps

  1. Release scores 85 after flag and 100 after clear; requesting evidence scores 92 either way.
  2. Without the call: 1/3 x 85 + 2/3 x 100 = 95.000, best = 95.000.
  3. With the call: 1/3 x 92 + 2/3 x 100 = 97.333.
  4. Gross value = 97.333 - 95.000 = 2.333.
  5. Net value = 2.333 - 0.0 = 2.333.

Use the idea

Price a verification step against the specific decision it feeds. The same step, with the same accuracy and cost, can be a bargain for one decision and worthless for another. Before spending any budget, also check whether the system already holds the answer in a log it failed to structure: free information should always be used.

Where the conclusion applies

One decision, one observation, a coherent probability model, and a gross value that ignores any later decisions the observation could also inform. The conditional release scores are the book's constructed values; a real classifier's scores would have to be estimated, and a wrong estimate can flip the verdict.

Common wrong turn: A call that returns interesting material was worth making
The chapter says an agent that runs a call whose result cannot change its action is not being thorough: it is failing to distinguish what the call would tell it from what it would do differently. The call has a real cost, a real information gain and zero value.
What this does not settle

Equation (11.4) prices a single observation against a fixed decision, not a sequence of them.

Chapter 11 source: "What this does not settle".

Check your understanding: With a shift of 5, what is the gross value, and is the call worth buying at cost 2?
Release scores 80 after flag and 95 after clear, and 90 before. With the call: 1/3 x 92 + 2/3 x 95 = 94.0, so gross value = 94 - 92 = 2.0. Net = 2 - 2 = 0, a tie.

Chapter 11 source: section "A worked price for one observation". Demonstration C11-D04.