Demonstration 1 of 4
Defer by comparing correctness, then price the deadline
If the model is 0.78 sure, when should the case go to an expert instead, and does the expert's accuracy alone settle it once the answer has to arrive in time?
Equation (27.2) compares two chances of being correct and sends the case to whichever side is higher, with ties going to the expert. The model's confidence only matters as one of the two numbers. It also treats the delay as costless. Once the reply can miss its deadline, delegation's value mixes the timely branch and the fallback and pays the fee on both, so the expert's conditional accuracy cannot answer the question alone; the break-even on-time probability is where the two values meet.
Scroll sideways for the whole equation
x is the case. The classifier's best class has probability max_y p(y | x) of being correct. p_E(x) is the probability that the designated expert E is correct on this case. Del(E, Delta) means delegate to E with maximum wait Delta. The arg max picks the most probable class. For the value comparison, a correct result is worth 10 and a wrong one -10, delegating costs a fee of 1 whether or not a reply arrives, and t is the probability the reply comes on time; a missed deadline triggers an authorized fallback worth 2. Probabilities are between 0 and 1.
Predict first. Pick the workbench case (acting is right 0.80, expert 0.95). With a guaranteed on-time return delegation wins 8 against 6. Now set the on-time probability to 0.60. Which route wins?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's two cases (model 0.78, expert 0.95 or 0.60) and the workbench's timely-return problem (acting 0.80, utilities 10 and -10, fee 1, fallback 2).
Calculated values
- Classifier best class probability
- 0.78
- Expert correctness p_E
- 0.95
- Route by Equation (27.2), delay ignored
- Delegate to E
- Value of acting now
- 5.60
- Value of delegating
- 8.00
- Route with the higher value
- delegate
- Timely-return probability that ties them
- 0.657
Equation (27.2): 0.95 >= 0.78, so it delegates (gap 0.95 - 0.78 = 0.17). Act = 0.78 x 10 + 0.22 x (-10) = 5.60. Delegate = t x gross + (1 - t) x 2 - 1 = 1.00 x 9.00 + 0.00 x 2 - 1 = 8.00, where gross = 0.95 x 10 + 0.05 x (-10) = 9.00. The two readings agree: delegate. Break-even: t = (act - 1) / (gross - 2) = (5.60 - 1) / (9.00 - 2) = 0.657.
Worked steps
- Equation (27.2): 0.95 >= 0.78, so it delegates (gap 0.95 - 0.78 = 0.17).
- Act = 0.78 x 10 + 0.22 x (-10) = 5.60.
- Timely review gross = 0.95 x 10 + 0.05 x (-10) = 9.00.
- Delegate = 1.00 x 9.00 + 0.00 x 2 - 1 = 8.00; the fee is paid on both branches.
- Higher value: delegate (5.60 against 8.00).
Use the idea
A single confidence threshold such as 0.80, below which everything is sent to review, is justified only if the destination's correctness and timeliness are the same for all the cases it covers. Otherwise ask how likely the actual destination is to be right on this kind of case, how likely it is to answer in time, and compare with acting.
Where the conclusion applies
Zero-one loss for Equation (27.2), one expert, and a known expert probability for this case type. The classifier's confidence is assumed calibrated. The utilities, fee and fallback value come from the workbench's constructed problem and are applied here to all four cases for comparison. The expert's accuracy is conditional on a timely reply and is not a permanent rating; if the probability is only a rough average, or the packet omits the evidence the expert needs, the comparison is no better than its inputs.
Common wrong turn: Delegate whenever the model is uncertain
What this does not settle
The chapter does not show how to estimate every action value, reviewer-error rate, queue delay, observation value, or authority boundary in a live institution.
Chapter 27 source: "What this does not settle".
Check your understanding: A model is 0.97 sure and an expert is right with probability 0.95. Which route does Equation (27.2) choose, and by how much?
Chapter 27 source: section "A value comparison, not a confidence threshold". Demonstration C27-D01.
Demonstration 2 of 4
Four routes, but only the authorized ones can run
Which routes are in the authorized set, which of them has the highest net value, and what removes a route?
Equation (27.1) first filters the actions: a route is available only if it is authorized and its preconditions hold. Equation (27.3) then takes the highest-value route among those. Timely review has gross value 8.04, and the hold turns that into 0.95 x 8.04 + 0.05 x (-4) - 1 minus its own cost, so its value falls one point for each point of hold cost. An excluded route is not low-scoring, it is absent: if it would have won, it stays a recommendation.
Scroll sideways for the whole equation
A route is one choice such as release now, review, wait or hold. The four rows are the chapter's act (release now), delegate (review), wait and narrow (the reversible hold); returning control is not given a value here. A_auth(x) is the set of actions in state x with Authorized = 1 (a declared grant covers them) and Pre = 1 (their state-specific preconditions hold, here a reviewer being available). R_auth(x) is the set of authorized routes, and V(route, x) is the route's net value in constructed utility units. Fraud has probability 0.10. A hold makes timely review likely (0.95) and costs the amount you choose; the timeout branch is worth -4. Review itself costs 1.
Predict first. The hold costs 2 by default and wins. Raise its cost to 5 with every route authorized. Which route has the highest value?
Choose an example
Scroll sideways for the whole figure
Constructed example: the book's own route-value table for the pending transfer (values -1, 1.164, 2 and 4.438), with the hold cost and the authority status varied.
Calculated values
- Release now
- -1.000
- Review; release on timeout
- 1.164
- Wait for confirmation; release on timeout
- 2.000
- Reversible hold and review; return on timeout
- 4.438
- Routes in the authorized set
- release now, review, wait for confirmation, reversible hold and review
- Highest among the authorized routes
- Reversible hold and review; return on timeout
- Hold cost at which waiting (2) overtakes it
- 4.438
Hold = 0.95 x 8.04 + 0.05 x (-4) - 2 - 1 = 4.438. Timely review gross value = 0.90 x [0.98 x 10 + 0.02 x (-10)] + 0.10 x [0.94 x 0 + 0.06 x (-100)] = 8.64 + (-0.60) = 8.04. The hold wins: 4.438 is above waiting at 2. Its value before its own cost is 0.95 x 8.04 + 0.05 x (-4) - 1 = 6.438, so its cost can rise to 6.438 - 2 = 4.438 before waiting overtakes it.
Worked steps
- Timely review gross = 8.64 + (-0.60) = 8.04.
- Net values: release -1.000, review 1.164, wait 2.000, hold 4.438.
- Hold = 0.95 x 8.04 + 0.05 x (-4) - 2 - 1 = 4.438.
- Status: All four routes authorized and the reviewer is available.
- Authorized set: release now, review, wait for confirmation, reversible hold and review.
- Highest value inside the set: Reversible hold and review; return on timeout at 4.438.
Use the idea
Write the table for a consequential action: one row per route, each with its payoff, delay cost and fallback, and next to each row the grant that authorizes it and the precondition it needs. Then keep only the rows that are actually permitted. A delegation row also needs a named destination, a deadline and a fallback; without them it is not a route.
Where the conclusion applies
Utilities, fraud probability, review error rates (detect fraud 0.94, wrongly reject 0.02) and response probabilities are the chapter's constructed values, and all four routes are costed on one scale. Real stakes may not fit one scale; a hard boundary may belong in the authorized set rather than in a cost. Changing any stated probability can change the winner. The reviewer accuracy here is held fixed, although in practice it can change with workload.
Common wrong turn: A high-scoring route can be executed
What this does not settle
The chapter does not make human review an oracle, nor does it rank every refusal above every act. Its bounded conclusion is structural: delegation requires a named route and must compete with acting, waiting, narrowing and returning control.
Chapter 27 source: "What this does not settle".
Check your understanding: If the hold cost 4 instead of 2, what is its net value and does it still beat waiting?
Chapter 27 source: section "Queue economics: what timely review is worth". Demonstration C27-D02.
Demonstration 3 of 4
Delegating more can make review slower than the deadline
How does sending a larger share of tasks to one reviewer change the mean time in review, and which tasks can then be released?
For one reviewer in the M/M/1 model, with random (Poisson) arrivals and exponentially distributed service times, the mean time in the system is 1/(mu - f x lambda) as long as the review arrivals stay below capacity. The time grows slowly at first and then steeply as the load approaches mu. At or beyond mu no stationary (long-run) mean exists and the backlog grows. The ledger shows what that does to the release contract: authority and a review packet are separate requirements from timing.
Scroll sideways for the whole equation
lambda is the arrival rate of tasks per hour. f is the delegation fraction, the share sent to review, so the reviewer sees f x lambda per hour. mu is the number of reviews the reviewer finishes per hour. The result is the mean hours a case spends waiting plus being reviewed. A stationary mean is that long-run average. In the ledger a task needs review when the agent is not authorized or its risk is above the agent's risk limit; it is released only with human authority, a received review packet and a mean time inside its deadline.
Predict first. With the default case (reviewer finishes 3 per hour), what happens to the mean time as the delegated share goes from 0.5 to 0.7? And at 0.9?
Choose an example
Scroll sideways for the whole figure
Constructed example: the notebook's default, changed and transfer cases and the chapter's queue example (4 tasks per hour, a reviewer finishing 3 per hour), computed with the laboratory's review-queue function, plus one faster reviewer defined for the reader.
Calculated values
- Review arrivals f x lambda (per hour)
- 2.00
- Reviewer capacity mu (per hour)
- 3
- Mean time in system
- 1.00 hours (60 minutes)
- Tasks released
- 2 of 3
- Task outcomes
- routine: autonomous; release: released after review; sensitive: blocked (no review packet)
- Chance the first delegated case meets its deadline
- 0.865
Review arrivals = 0.50 x 4 = 2.00 per hour, below 3. Mean time in system = 1 / (3 - 2.00) = 1 / 1.00 = 1.00 hours, which is 60 minutes, outside the 20-minute window. The reviewer's accuracy did not change; only the load did. If the M/M/1 first-come-first-served assumption held exactly, a delegated case would meet a deadline of 2.0 hours with chance 1 - exp(-1.00 x 2.0) = 0.865; a mean does not promise any single case meets its deadline. Ledger: routine: autonomous; release: released after review; sensitive: blocked (no review packet).
Worked steps
- Review arrivals = f x lambda = 0.50 x 4 = 2.00 per hour; capacity mu = 3.
- Stable, since 2.00 < 3: mean = 1 / (3 - 2.00) = 1.00 hours.
- Tasks above the agent risk limit 0.10 or without agent authority need review: release, sensitive.
- A delegated task is released only with human authority, a received review packet and a mean inside its deadline.
- Result: 2 of 3 tasks released.
Use the idea
Before promising a review deadline, compare the planned review load with the reviewer's capacity. A rule such as delegate everything doubtful can overload the one route that was supposed to supply safety. Keep a route ledger per destination, with case type, deadline, packet completeness, response state and final action, so a failure can be placed: a bad recommendation, a wrong destination, a late answer or an unowned escalation.
Where the conclusion applies
The M/M/1 model: one reviewer, independent random (Poisson) arrivals at a steady rate, exponentially distributed service times with a steady rate, and a long-run mean. Other service-time shapes give a different formula. Batching, priorities, correlated arrivals or a changing reviewer break the model. A mean does not promise that any single case meets its deadline, and the laboratory's deadline test is a planning diagnostic, not a probability. The faster reviewer is defined for this reader.
Common wrong turn: A queue mean proves the deadline is met or missed
Check your understanding: With mu = 4 per hour and f = 0.75, what is the mean time in review in minutes?
Chapter 27 source: section "Queue economics: what timely review is worth". Demonstration C27-D03.
Demonstration 4 of 4
Waiting needs an observation worth its delay and a fallback everywhere
When is it right to wait for a confirmation instead of acting now, and what if one state reachable by the deadline has no authorized fallback?
If the confirmation arrives, the transfer is handled correctly for a value of 9; if not, the fallback releases it at -1. So the gain over releasing now is 10 x q, a straight line. Equation (27.4) allows waiting only when that gain is strictly above the delay cost and the fallback is permitted in every reachable deadline state. One state without a permitted fallback is enough to fail the test, however valuable the observation.
Scroll sideways for the whole equation
VOI(o) is the value of the observation o: here the extra expected value of waiting for the signed confirmation compared with releasing now. DelayCost is what the wait costs (2 units). q is the probability that the confirmation arrives in time. The fallback is what the controller does at the deadline, and it must be an authorized action in every state reachable by then: the observation arrives in time, arrives late, the case changes, or the source fails.
Predict first. At q = 0.20, with a fallback authorized in every state, does the test allow waiting?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's wait route (value 9 on arrival, -1 at the fallback, delay cost 2) with the arrival probability and the missing-fallback state varied.
Calculated values
- Value of information (VOI)
- 5.00
- Delay cost
- 2.00
- VOI minus delay cost
- 3.00
- Fallback authorized in every reachable state
- yes
- Passes Equation (27.4)
- Yes (waiting allowed)
- Break-even arrival probability
- 0.20
VOI = 0.5 x 9 + 0.5 x (-1) - (-1) = 5.00, and VOI - delay cost = 5.00 - 2 = 3.00. Both conditions hold: the observation is worth more than the delay and a fallback is permitted in all four reachable states, so waiting passes. The equation is a necessary test (the word 'only if'): passing it does not prove waiting is the best route, which Equation (27.3) still decides.
Worked steps
- Gain from waiting over releasing now: VOI = 0.5 x 9 + 0.5 x (-1) - (-1) = 5.00.
- Condition 1: VOI 5.00 against delay cost 2, needing strictly more: holds.
- Reachable states by the deadline: in time, late, the case changes, the source fails.
- Condition 2: the fallback must be authorized in every one: all four are.
- Result: Yes (waiting allowed).
Use the idea
When an agent says it will wait, ask four things: which fact would change the action, who supplies it, by when, and what happens if it never comes. If any answer is missing, there is no waiting action, only delay. If the observation source has become unavailable, move to the declared fallback instead of waiting under a plan that assumes its arrival.
Where the conclusion applies
One observation and the chapter's constructed values (9, -1, delay cost 2). VOI is read here as the gain over releasing now, a simple special case of the Chapter 8 idea. The four reachable states are the ones the chapter lists; a real case can have more. Equation (27.4) is a necessary test: passing it does not make waiting the best route.
Common wrong turn: Waiting is just not acting
What this does not settle
The chapter does not prove that a learned rejector yields fair, lawful, safe, or corrigible deployment, and it does not show how to estimate every observation value or authority boundary in a live institution.
Chapter 27 source: "What this does not settle".
Check your understanding: If the delay cost rose to 3 and the confirmation still arrived with probability 0.5, does the test allow waiting (fallback authorized everywhere)?
Chapter 27 source: section "Waiting is not hiding". Demonstration C27-D04.