The Mathematics of AI Agents, laboratory reader ยท Chapter 6

The Price of a Choice

Probabilities say what an action is likely to lead to; utilities say what those outcomes are worth. A decision needs both.

These four demonstrations follow the chapter's document controller. It must choose among releasing a report now, requesting more evidence, and escalating to a person. Each demonstration changes one declared value and shows which comparison moves and which stays fixed, including the notebook's own decision cases, a probability that is only known within a range, and a gamble whose mean payoff has no limit.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

Three actions, one rule, and where the order flips

If the probabilities stay fixed, which value change flips the winner or the order, and by how much does each action's score depend on it?

Each action's expected utility multiplies each outcome's worth by its probability and adds the products, then Equation (6.3) takes the largest. The left bars show the two columns meeting: a gain part, a failure part and any delay cost. The probabilities never move in this demonstration; only the value of a failure does, and the right panel shows each action's expected utility as a straight line in that value. Where two lines cross, the order flips, which is the one question a person has to answer. In the notebook's transfer case the two actions tie exactly, so the rule alone does not choose.

Equation (6.2), written in LaTeX: \operatorname{EU}(a)=\sum_{y}p(y\mid a) u(y)

Equation (6.3), written in LaTeX: a^{\star}=\arg\max_{a\in\mathcal{A}}\operatorname{EU}(a)

Scroll sideways for the whole equation

a is an action, y an outcome (a supported release or its failure), p(y | a) the probability of outcome y under action a, and u(y) its declared utility. In Table 6.1 a supported release is worth 100 and delay subtracts the same amount from every outcome, which is the same as subtracting it once from the average because the probabilities add to one. The arg max picks the action with the largest EU(a), the expected utility of action a. In the right panel u is the utility of the failure outcome, the one value that changes along the horizontal axis. In the notebook default, review is worth 2 and abstaining 0 for certain; in the transfer case retry pays 8 or the failure value at cost 1 and escalating pays 1.

Predict first. In the book's table, make an unsupported release worth -400 instead of 0 (day of delay 40). Which action loses the most, and does the winner change?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Three actions, one rule, and where the order flips. Left: for each action a gain bar, a failure bar, any hatched delay cost and a diamond at its expected utility (Table 6.1, failure worth 0). Right: expected utility of each action against the utility of the failure outcome, with the current value marked and each tie point circled.
Decision table: Book Table 6.1: release, evidence, escalate, Value of the failure outcome: As supplied (0, -50 or -12)
Constructed example: the book's Table 6.1 values, and the notebook's default, changed and transfer decision cases, computed with the laboratory's expected-utility function.

Calculated values

Release now
85.0
Request evidence
92.0
Escalate to a person
59.5
Chosen by Equation (6.3)
Request evidence
Utility at which two actions tie
Release now and Escalate to a person at -175.9

EU(Release now) = 0.85 x 100 + 0.15 x 0 = 85.0. EU(Request evidence) = 0.97 x 100 + 0.03 x 0 - 5 = 92.0. EU(Escalate to a person) = 0.995 x 100 + 0.005 x 0 - 40 = 59.5. Request evidence has the highest expected utility, so Equation (6.3) selects it. Order from highest to lowest: Request evidence, Release now, Escalate to a person. Release now and Escalate to a person tie when 85 + 0.15 u = 59.5 + 0.005 u, so u = (-25.5) / 0.145 = -175.9. The probabilities never changed; only the declared values did.

Worked steps

  1. Hour price: indifferent between 0.85 now and 0.90 after an hour, so an hour costs (0.90 - 0.85) x 100 = 5.
  2. Day price: indifferent between 0.50 now and 0.90 after a day, so a day costs (0.90 - 0.50) x 100 = 40.
  3. Evidence against release: (0.97 - 0.85) x 100 - 5 = 12 - 5 = 7 units above.
  4. Escalation against release: (0.995 - 0.85) x 100 - 40 = 14.5 - 40 = (-25.5), 25.5 units below.
  5. Evidence against escalation: 7 - (-25.5) = 32.5, the same as 92.0 - 59.5 = 32.5 from the table.

Use the idea

Before trusting an automated choice, write the table: one row per action, including asking for evidence and handing the decision to a person. Then check whether a change in a value you are unsure of changes the winner.

Where the conclusion applies

One decision, two mutually exclusive outcomes per action (one for review and abstaining), and utilities on one shared scale. The probabilities and utilities are constructed teaching values from Table 6.1 and the notebook, not estimates for any real controller. The 'made worse' failure values are the book's -400 and the notebook's -100; the transfer value -24 is defined for this reader as double the supplied value. Expected utility compares actions; it does not say whose values the numbers represent, and a large negative utility is not the same as a hard permission boundary.

Common wrong turn: Folding belief into worth
The probability column and the value column come from different places. An agent that treats a low-probability outcome as low-value, or a catastrophic outcome as merely unlikely, has folded belief into worth; here the probabilities do not move while the value of a failure alone changes the order.
What this does not settle

All values in Table 6.1 are constructed for teaching. Nothing here says where a utility function should come from, only what follows once one exists.

Chapter 6 source: "What this does not settle".

Check your understanding: In the book's table with an unsupported release worth -400 and a day costing 40, what does releasing now score?
0.85 x 100 + 0.15 x (-400) = 85 - 60 = 25. Requesting evidence scores 92 - 12 = 80, so it still wins, and escalation (57.5) now beats releasing now.

Chapter 6 source: section "The decision rule". Demonstration C06-D01.

Demonstration 2 of 4

Elicit one number: the price of an hour

Between releasing now and requesting evidence, what single number decides the choice, and what if the support probability is itself shaky?

Requesting evidence scores 97 minus the price of an hour; releasing now scores 100 times its support probability. The lines cross at the break-even price. A person does not need to state a whole utility function, only which side of that one threshold they are on. When the support probability is shaky, release now has a range of scores: if the evidence line sits above the whole range, or below it, the decision is robust and the analysis can stop; if it sits inside, the analysis has located the question a person has to settle.

Equation (6.3), written in LaTeX: a^{\star}=\arg\max_{a\in\mathcal{A}}\operatorname{EU}(a)

Equation (6.2), written in LaTeX: \operatorname{EU}(a)=\sum_{y}p(y\mid a) u(y)

Scroll sideways for the whole equation

EU(a) is the expected utility of action a. Releasing now succeeds with the chosen support probability; requesting evidence succeeds with probability 0.97 but costs one hour, priced in utility points. An unsupported release is worth 0 and a supported one 100. 'Range' means the support probability is only known to lie between 0.80 and 0.90, so releasing now has a range of expected utilities, 80 to 90, instead of one number.

Predict first. With release support 0.85, if an hour of delay is worth exactly 12 points, does Equation (6.3) choose releasing now, requesting evidence, or neither because they tie?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Elicit one number: the price of an hour. Expected utility against the price of an hour of delay: a falling line for requesting evidence and a flat line at 85.0 for releasing now; the break-even is 12.0 points and the current price 5 is marked.
Price of an hour of delay: 5, Support probability when releasing now: 0.85
Constructed example: the book's Table 6.1 values, computed with the laboratory's expected-utility function; the hour prices 0 and 15 and the support range are defined for this reader.

Calculated values

Release now
85.0
Request evidence
92.0
Break-even hour price
12.0
Chosen
Request evidence

Break-even price = 0.97 x 100 - 0.85 x 100 = 97.0 - 85.0 = 12.0 points per hour. Release now scores 0.85 x 100 = 85.0. At 5 points, request evidence scores 97.0 - 5 = 92.0. Request evidence wins because 5 is below the break-even price. Only one question needs an answer from a person: is an hour worth more or less than the break-even price?

Worked steps

  1. Request evidence: 0.97 x 100 - 5 = 92.0.
  2. Release now: 0.85 x 100 = 85.0.
  3. Break-even hour price: 97.0 - 85.0 = 12.0.
  4. Request evidence wins because 5 is below the break-even price.

Use the idea

When a decision depends on a value nobody has written down, compute the threshold at which the decision flips and ask the person responsible which side they are on. That question is answerable; asking what an outcome is worth in the abstract usually is not. Then vary each shaky input across the range anyone would defend and report whether the choice changes.

Where the conclusion applies

Two actions, utilities on the Table 6.1 scale, and delay priced additively. The 0.85, 0.90 and 0.97 support probabilities are constructed, and the range 0.80 to 0.90 is an illustrative uncertainty defined for this reader. A tie is reported as a tie, because Equation (6.3) does not break ties. A range of scores is not a probability distribution over the support.

Common wrong turn: Asking people what outcomes are worth
Do not ask people what outcomes are worth. Compute the threshold at which the decision changes, then ask which side of it they are on: the first question is unanswerable and the second is ordinary.
What this does not settle

Two agents can hold the same probability with very different confidence, and Equation (6.2) cannot tell them apart, because a probability enters the average as a weight and carries no record of how it was obtained.

Chapter 6 source: "When the probability itself is shaky".

Check your understanding: If releasing now had support probability 0.90, what hour price would make the two actions tie?
97 - 90 = 7 points. Below 7 requesting evidence wins; above 7 releasing now wins.

Chapter 6 source: section "Eliciting one number". Demonstration C06-D02.

Demonstration 3 of 4

Risk is the shape of the utility curve

How much mean payoff would an agent give up to replace a gamble with a sure amount, and what if the gamble's mean payoff has no limit?

Average the utilities, not the payoffs, then ask which sure payoff has that utility. A curve that bends down puts the certainty equivalent below the mean; a straight line puts it at the mean; a curve that bends up puts it above. The gap is a number in payoff units, not a temperament. All three curves agree on the order of payoffs, so Equation (6.1) cannot tell them apart; they disagree about gambles. The unbounded gamble shows the average must be checked: with the square-root curve the series converges to 1 + sqrt(2) and the certainty equivalent is finite, while for the other curves it does not converge.

Equation (6.4), written in LaTeX: u(\operatorname{CE})=\operatorname{E}[u(Y)]

Equation (6.2), written in LaTeX: \operatorname{EU}(a)=\sum_{y}p(y\mid a) u(y)

Equation (6.1), written in LaTeX: y_1 \succ y_2 \iff u(y_1) > u(y_2)

Scroll sideways for the whole equation

Y is the payoff of a gamble that pays 100 with the chosen probability and 0 otherwise, or, in the last gamble, 2^k with probability 2^(-k) for k = 1, 2, 3 and so on. u is the utility curve, E[u(Y)] its probability-weighted average, and CE the certainty equivalent: the sure payoff whose utility equals that average. The plot rescales utility to run from 0 to 1, which does not move CE. Equation (6.1) says only that a larger payoff gets a larger utility; all three curves do that.

Predict first. For a curve that bends down, is the certainty equivalent above or below the mean payoff of 50?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Risk is the shape of the utility curve. A square root utility curve against payoff with the mean payoff 50.0 and the certainty equivalent 25.0 marked on one dashed level; a dotted diagonal shows a straight line for comparison.
Utility curve: Bends down: square root, Gamble: Pays 100 with probability 0.5
Constructed example: the chapter's worked certainty equivalent (100 or 0 with equal chance), one changed probability and the chapter's unbounded gamble with a square-root utility (1 + sqrt(2), about 2.414, and a certainty equivalent of about 5.83).

Calculated values

Mean payoff
50.0
E[u(Y)] for this curve (not comparable across curves)
5.00
Certainty equivalent
25.00
Mean minus certainty equivalent
25.00
A sure 40 against the gamble
sure payoff preferred

Mean payoff = 0.5 x 100 = 50.0. E[u] = 0.5 x sqrt(100) + 0.5 x sqrt(0) = 5.0, so CE = 5.0 x 5.0 = 25.0. Mean minus certainty equivalent = 50.0 - 25.00 = 25.00. The agent would trade the gamble for a guaranteed 25.0, giving up 25.0 of mean payoff to remove the variance. That gap is the price of avoiding risk under this curve. Every curve here ranks payoffs the same way, more being better (Equation (6.1)), yet they disagree about this gamble: a sure 40 beats it exactly when CE < 40, and CE = 25.0, so the sure 40 is preferred.

Worked steps

  1. Mean payoff = 0.5 x 100 + 0.5 x 0 = 50.0.
  2. E[u] = 0.5 x sqrt(100) + 0.5 x sqrt(0) = 5.0, so CE = 5.0 x 5.0 = 25.0.
  3. Mean minus certainty equivalent = 50.0 - 25.00 = 25.00.
  4. A sure 40 is preferred when CE < 40: CE = 25.0, so the sure payoff wins.

Use the idea

When a team says a system should be cautious, ask for the curve or for one certainty equivalent. A stated gap such as 25 on a 50 mean can be checked and argued about; the word cautious cannot.

Where the conclusion applies

Constructed payoffs and probabilities. The square-root and square curves are illustrations of bending down and bending up, not elicited preferences. One certainty equivalent is consistent with risk aversion but does not prove the whole curve is concave. The certainty equivalent is a guaranteed payoff under the declared utility, not an observed willingness to pay. The sure 40 in the interpretation is only a comparison value.

Common wrong turn: Pricing a gamble by its mean payoff
An unbounded expected payoff does not determine the gamble's utility, so the game is not worth any finite price; and one certainty-equivalent comparison is consistent with risk aversion but does not establish global concavity.
What this does not settle

The representation in Equation (6.1) requires complete and transitive comparisons, and real preferences violate both. A cycle such as speed over accuracy, accuracy over cost and cost over speed cannot be summarized by any ranking.

Chapter 6 source: "What this does not settle".

Check your understanding: With a square-root utility and a gamble paying 100 with probability 0.8, what is the certainty equivalent?
E[u] = 0.8 x 10 = 8, so CE = 8 x 8 = 64, which is 16 below the mean payoff of 80.

Chapter 6 source: section "A worked certainty equivalent". Demonstration C06-D03.

Demonstration 4 of 4

Declining is an action with a price

How does the value placed on declining decide how many cases an agent answers and how often it is wrong?

Answering wins when p x 1 + (1 - p) x w exceeds the value of declining d, which happens when p is above the threshold t = (d - w) / (1 - w). Each threshold picks one point on the risk and coverage curve, so the same agent can be reported at very different error rates depending on the price of declining. Raising the cost of declining raises coverage; lowering it hands more away.

Equation (6.3), written in LaTeX: a^{\star}=\arg\max_{a\in\mathcal{A}}\operatorname{EU}(a)

Equation (6.2), written in LaTeX: \operatorname{EU}(a)=\sum_{y}p(y\mid a) u(y)

Scroll sideways for the whole equation

For each case the agent compares two actions. Answer: utility 1 if correct, w (the chosen wrong-answer utility) if not, with p the probability the answer is correct. Decline: d (the chosen utility of declining) for certain. Coverage is the share of cases answered, and the error rate is measured only over answered cases.

Predict first. Make a wrong answer cost 9 instead of 4. Does coverage go up or down, and what happens to the error rate?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Declining is an action with a price. Left: expected utility of answering and of declining against confidence, with the threshold 0.80 marked. Right: expected error rate against coverage with the chosen point at coverage 0.40.
Utility of a wrong answer: -4, Utility of declining: 0
Constructed example: 100 calibrated cases defined for this reader; each answer or decline choice is computed with the laboratory's expected-utility function.

Calculated values

Answer threshold t
0.80
Cases answered
40 of 100
Coverage
0.40
Expected error rate when answering
0.100
Expected utility per case
0.200

Answer when p x 1 + (1 - p) x (-4) > 0.0, that is when p > t = (0.0 - (-4)) / (1 - (-4)) = 4.0 / 5 = 0.80. The 40 cases from p = 0.8025 to 0.9975 are answered, so coverage = 40/100 = 0.40 and their expected error rate is 1 - (0.8025 + 0.9975) / 2 = 0.100. A more expensive wrong answer raises t and lowers coverage; a more expensive decline lowers t and raises coverage. The value attached to declining picks the point on the curve.

Worked steps

  1. Answer: p x 1 + (1 - p) x (-4); decline: 0.0 for certain.
  2. Threshold t = (0.0 - (-4)) / (1 - (-4)) = 0.80.
  3. 40 cases (p from 0.8025 to 0.9975) are answered: coverage 40/100 = 0.40, error 1 - (0.8025 + 0.9975) / 2 = 0.100.

Use the idea

When a system reports an error rate, ask for its coverage and for the value it assigns to declining. Two error rates at different coverage describe different services, and the price of declining is what chose between them.

Where the conclusion applies

One hundred constructed cases whose confidences are spread evenly from 0.5025 to 0.9975 and assumed to be calibrated, so a case with confidence p is correct with probability p. Real confidence scores need not be calibrated; the error rates shown are expected values under that assumption, not observed rates. Because the confidences are spread evenly, the risk and coverage curve is a straight line, and a real system's curve would take its shape from its real confidences.

Common wrong turn: Reading one accuracy figure as the whole product
A system reported as ninety percent accurate has one point on the risk and coverage curve. The same system might reach ninety-eight percent accuracy at sixty percent coverage, which is a different product with different economics, and the single number conceals which one was built.
What this does not settle

Equation (6.2) takes one utility function, but real deployments have several parties with different stakes, and no mathematics in this chapter combines them. Whatever function a system acts on is somebody's, and the question of whose is answerable by reading the numbers.

Chapter 6 source: "Whose utility".

Check your understanding: With a wrong answer worth -4 and declining worth 0, what is the answer threshold t?
t = (0 - (-4)) / (1 - (-4)) = 4 / 5 = 0.80, so only the 40 cases with p above 0.80 are answered.

Chapter 6 source: section "The cost of declining, measured". Demonstration C06-D04.