Demonstration 1 of 4
Three actions, one rule, and where the order flips
If the probabilities stay fixed, which value change flips the winner or the order, and by how much does each action's score depend on it?
Each action's expected utility multiplies each outcome's worth by its probability and adds the products, then Equation (6.3) takes the largest. The left bars show the two columns meeting: a gain part, a failure part and any delay cost. The probabilities never move in this demonstration; only the value of a failure does, and the right panel shows each action's expected utility as a straight line in that value. Where two lines cross, the order flips, which is the one question a person has to answer. In the notebook's transfer case the two actions tie exactly, so the rule alone does not choose.
Scroll sideways for the whole equation
a is an action, y an outcome (a supported release or its failure), p(y | a) the probability of outcome y under action a, and u(y) its declared utility. In Table 6.1 a supported release is worth 100 and delay subtracts the same amount from every outcome, which is the same as subtracting it once from the average because the probabilities add to one. The arg max picks the action with the largest EU(a), the expected utility of action a. In the right panel u is the utility of the failure outcome, the one value that changes along the horizontal axis. In the notebook default, review is worth 2 and abstaining 0 for certain; in the transfer case retry pays 8 or the failure value at cost 1 and escalating pays 1.
Predict first. In the book's table, make an unsupported release worth -400 instead of 0 (day of delay 40). Which action loses the most, and does the winner change?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's Table 6.1 values, and the notebook's default, changed and transfer decision cases, computed with the laboratory's expected-utility function.
Calculated values
- Release now
- 85.0
- Request evidence
- 92.0
- Escalate to a person
- 59.5
- Chosen by Equation (6.3)
- Request evidence
- Utility at which two actions tie
- Release now and Escalate to a person at -175.9
EU(Release now) = 0.85 x 100 + 0.15 x 0 = 85.0. EU(Request evidence) = 0.97 x 100 + 0.03 x 0 - 5 = 92.0. EU(Escalate to a person) = 0.995 x 100 + 0.005 x 0 - 40 = 59.5. Request evidence has the highest expected utility, so Equation (6.3) selects it. Order from highest to lowest: Request evidence, Release now, Escalate to a person. Release now and Escalate to a person tie when 85 + 0.15 u = 59.5 + 0.005 u, so u = (-25.5) / 0.145 = -175.9. The probabilities never changed; only the declared values did.
Worked steps
- Hour price: indifferent between 0.85 now and 0.90 after an hour, so an hour costs (0.90 - 0.85) x 100 = 5.
- Day price: indifferent between 0.50 now and 0.90 after a day, so a day costs (0.90 - 0.50) x 100 = 40.
- Evidence against release: (0.97 - 0.85) x 100 - 5 = 12 - 5 = 7 units above.
- Escalation against release: (0.995 - 0.85) x 100 - 40 = 14.5 - 40 = (-25.5), 25.5 units below.
- Evidence against escalation: 7 - (-25.5) = 32.5, the same as 92.0 - 59.5 = 32.5 from the table.
Use the idea
Before trusting an automated choice, write the table: one row per action, including asking for evidence and handing the decision to a person. Then check whether a change in a value you are unsure of changes the winner.
Where the conclusion applies
One decision, two mutually exclusive outcomes per action (one for review and abstaining), and utilities on one shared scale. The probabilities and utilities are constructed teaching values from Table 6.1 and the notebook, not estimates for any real controller. The 'made worse' failure values are the book's -400 and the notebook's -100; the transfer value -24 is defined for this reader as double the supplied value. Expected utility compares actions; it does not say whose values the numbers represent, and a large negative utility is not the same as a hard permission boundary.
Common wrong turn: Folding belief into worth
What this does not settle
All values in Table 6.1 are constructed for teaching. Nothing here says where a utility function should come from, only what follows once one exists.
Chapter 6 source: "What this does not settle".
Check your understanding: In the book's table with an unsupported release worth -400 and a day costing 40, what does releasing now score?
Chapter 6 source: section "The decision rule". Demonstration C06-D01.
Demonstration 2 of 4
Elicit one number: the price of an hour
Between releasing now and requesting evidence, what single number decides the choice, and what if the support probability is itself shaky?
Requesting evidence scores 97 minus the price of an hour; releasing now scores 100 times its support probability. The lines cross at the break-even price. A person does not need to state a whole utility function, only which side of that one threshold they are on. When the support probability is shaky, release now has a range of scores: if the evidence line sits above the whole range, or below it, the decision is robust and the analysis can stop; if it sits inside, the analysis has located the question a person has to settle.
Scroll sideways for the whole equation
EU(a) is the expected utility of action a. Releasing now succeeds with the chosen support probability; requesting evidence succeeds with probability 0.97 but costs one hour, priced in utility points. An unsupported release is worth 0 and a supported one 100. 'Range' means the support probability is only known to lie between 0.80 and 0.90, so releasing now has a range of expected utilities, 80 to 90, instead of one number.
Predict first. With release support 0.85, if an hour of delay is worth exactly 12 points, does Equation (6.3) choose releasing now, requesting evidence, or neither because they tie?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's Table 6.1 values, computed with the laboratory's expected-utility function; the hour prices 0 and 15 and the support range are defined for this reader.
Calculated values
- Release now
- 85.0
- Request evidence
- 92.0
- Break-even hour price
- 12.0
- Chosen
- Request evidence
Break-even price = 0.97 x 100 - 0.85 x 100 = 97.0 - 85.0 = 12.0 points per hour. Release now scores 0.85 x 100 = 85.0. At 5 points, request evidence scores 97.0 - 5 = 92.0. Request evidence wins because 5 is below the break-even price. Only one question needs an answer from a person: is an hour worth more or less than the break-even price?
Worked steps
- Request evidence: 0.97 x 100 - 5 = 92.0.
- Release now: 0.85 x 100 = 85.0.
- Break-even hour price: 97.0 - 85.0 = 12.0.
- Request evidence wins because 5 is below the break-even price.
Use the idea
When a decision depends on a value nobody has written down, compute the threshold at which the decision flips and ask the person responsible which side they are on. That question is answerable; asking what an outcome is worth in the abstract usually is not. Then vary each shaky input across the range anyone would defend and report whether the choice changes.
Where the conclusion applies
Two actions, utilities on the Table 6.1 scale, and delay priced additively. The 0.85, 0.90 and 0.97 support probabilities are constructed, and the range 0.80 to 0.90 is an illustrative uncertainty defined for this reader. A tie is reported as a tie, because Equation (6.3) does not break ties. A range of scores is not a probability distribution over the support.
Common wrong turn: Asking people what outcomes are worth
What this does not settle
Two agents can hold the same probability with very different confidence, and Equation (6.2) cannot tell them apart, because a probability enters the average as a weight and carries no record of how it was obtained.
Chapter 6 source: "When the probability itself is shaky".
Check your understanding: If releasing now had support probability 0.90, what hour price would make the two actions tie?
Chapter 6 source: section "Eliciting one number". Demonstration C06-D02.
Demonstration 3 of 4
Risk is the shape of the utility curve
How much mean payoff would an agent give up to replace a gamble with a sure amount, and what if the gamble's mean payoff has no limit?
Average the utilities, not the payoffs, then ask which sure payoff has that utility. A curve that bends down puts the certainty equivalent below the mean; a straight line puts it at the mean; a curve that bends up puts it above. The gap is a number in payoff units, not a temperament. All three curves agree on the order of payoffs, so Equation (6.1) cannot tell them apart; they disagree about gambles. The unbounded gamble shows the average must be checked: with the square-root curve the series converges to 1 + sqrt(2) and the certainty equivalent is finite, while for the other curves it does not converge.
Scroll sideways for the whole equation
Y is the payoff of a gamble that pays 100 with the chosen probability and 0 otherwise, or, in the last gamble, 2^k with probability 2^(-k) for k = 1, 2, 3 and so on. u is the utility curve, E[u(Y)] its probability-weighted average, and CE the certainty equivalent: the sure payoff whose utility equals that average. The plot rescales utility to run from 0 to 1, which does not move CE. Equation (6.1) says only that a larger payoff gets a larger utility; all three curves do that.
Predict first. For a curve that bends down, is the certainty equivalent above or below the mean payoff of 50?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's worked certainty equivalent (100 or 0 with equal chance), one changed probability and the chapter's unbounded gamble with a square-root utility (1 + sqrt(2), about 2.414, and a certainty equivalent of about 5.83).
Calculated values
- Mean payoff
- 50.0
- E[u(Y)] for this curve (not comparable across curves)
- 5.00
- Certainty equivalent
- 25.00
- Mean minus certainty equivalent
- 25.00
- A sure 40 against the gamble
- sure payoff preferred
Mean payoff = 0.5 x 100 = 50.0. E[u] = 0.5 x sqrt(100) + 0.5 x sqrt(0) = 5.0, so CE = 5.0 x 5.0 = 25.0. Mean minus certainty equivalent = 50.0 - 25.00 = 25.00. The agent would trade the gamble for a guaranteed 25.0, giving up 25.0 of mean payoff to remove the variance. That gap is the price of avoiding risk under this curve. Every curve here ranks payoffs the same way, more being better (Equation (6.1)), yet they disagree about this gamble: a sure 40 beats it exactly when CE < 40, and CE = 25.0, so the sure 40 is preferred.
Worked steps
- Mean payoff = 0.5 x 100 + 0.5 x 0 = 50.0.
- E[u] = 0.5 x sqrt(100) + 0.5 x sqrt(0) = 5.0, so CE = 5.0 x 5.0 = 25.0.
- Mean minus certainty equivalent = 50.0 - 25.00 = 25.00.
- A sure 40 is preferred when CE < 40: CE = 25.0, so the sure payoff wins.
Use the idea
When a team says a system should be cautious, ask for the curve or for one certainty equivalent. A stated gap such as 25 on a 50 mean can be checked and argued about; the word cautious cannot.
Where the conclusion applies
Constructed payoffs and probabilities. The square-root and square curves are illustrations of bending down and bending up, not elicited preferences. One certainty equivalent is consistent with risk aversion but does not prove the whole curve is concave. The certainty equivalent is a guaranteed payoff under the declared utility, not an observed willingness to pay. The sure 40 in the interpretation is only a comparison value.
Common wrong turn: Pricing a gamble by its mean payoff
What this does not settle
The representation in Equation (6.1) requires complete and transitive comparisons, and real preferences violate both. A cycle such as speed over accuracy, accuracy over cost and cost over speed cannot be summarized by any ranking.
Chapter 6 source: "What this does not settle".
Check your understanding: With a square-root utility and a gamble paying 100 with probability 0.8, what is the certainty equivalent?
Chapter 6 source: section "A worked certainty equivalent". Demonstration C06-D03.
Demonstration 4 of 4
Declining is an action with a price
How does the value placed on declining decide how many cases an agent answers and how often it is wrong?
Answering wins when p x 1 + (1 - p) x w exceeds the value of declining d, which happens when p is above the threshold t = (d - w) / (1 - w). Each threshold picks one point on the risk and coverage curve, so the same agent can be reported at very different error rates depending on the price of declining. Raising the cost of declining raises coverage; lowering it hands more away.
Scroll sideways for the whole equation
For each case the agent compares two actions. Answer: utility 1 if correct, w (the chosen wrong-answer utility) if not, with p the probability the answer is correct. Decline: d (the chosen utility of declining) for certain. Coverage is the share of cases answered, and the error rate is measured only over answered cases.
Predict first. Make a wrong answer cost 9 instead of 4. Does coverage go up or down, and what happens to the error rate?
Choose an example
Scroll sideways for the whole figure
Constructed example: 100 calibrated cases defined for this reader; each answer or decline choice is computed with the laboratory's expected-utility function.
Calculated values
- Answer threshold t
- 0.80
- Cases answered
- 40 of 100
- Coverage
- 0.40
- Expected error rate when answering
- 0.100
- Expected utility per case
- 0.200
Answer when p x 1 + (1 - p) x (-4) > 0.0, that is when p > t = (0.0 - (-4)) / (1 - (-4)) = 4.0 / 5 = 0.80. The 40 cases from p = 0.8025 to 0.9975 are answered, so coverage = 40/100 = 0.40 and their expected error rate is 1 - (0.8025 + 0.9975) / 2 = 0.100. A more expensive wrong answer raises t and lowers coverage; a more expensive decline lowers t and raises coverage. The value attached to declining picks the point on the curve.
Worked steps
- Answer: p x 1 + (1 - p) x (-4); decline: 0.0 for certain.
- Threshold t = (0.0 - (-4)) / (1 - (-4)) = 0.80.
- 40 cases (p from 0.8025 to 0.9975) are answered: coverage 40/100 = 0.40, error 1 - (0.8025 + 0.9975) / 2 = 0.100.
Use the idea
When a system reports an error rate, ask for its coverage and for the value it assigns to declining. Two error rates at different coverage describe different services, and the price of declining is what chose between them.
Where the conclusion applies
One hundred constructed cases whose confidences are spread evenly from 0.5025 to 0.9975 and assumed to be calibrated, so a case with confidence p is correct with probability p. Real confidence scores need not be calibrated; the error rates shown are expected values under that assumption, not observed rates. Because the confidences are spread evenly, the risk and coverage curve is a straight line, and a real system's curve would take its shape from its real confidences.
Common wrong turn: Reading one accuracy figure as the whole product
What this does not settle
Equation (6.2) takes one utility function, but real deployments have several parties with different stakes, and no mathematics in this chapter combines them. Whatever function a system acts on is somebody's, and the question of whose is answerable by reading the numbers.
Chapter 6 source: "Whose utility".
Check your understanding: With a wrong answer worth -4 and declining worth 0, what is the answer threshold t?
Chapter 6 source: section "The cost of declining, measured". Demonstration C06-D04.