The Mathematics of AI Agents, laboratory reader ยท Chapter 27

The Mathematics of Delegation

A handoff is an action with a destination, a deadline and a fallback, and it has to beat acting, waiting and narrowing on value.

These four demonstrations follow the chapter's pending transfer. A confident agent could send the money, but a named reviewer may know something it does not. Each demonstration changes one declared value and shows what moves the choice: who is more likely to be right and whether the answer arrives in time, which routes are authorized, how long a queue takes and who can be released through it, and whether waiting is worth its delay.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

Defer by comparing correctness, then price the deadline

If the model is 0.78 sure, when should the case go to an expert instead, and does the expert's accuracy alone settle it once the answer has to arrive in time?

Equation (27.2) compares two chances of being correct and sends the case to whichever side is higher, with ties going to the expert. The model's confidence only matters as one of the two numbers. It also treats the delay as costless. Once the reply can miss its deadline, delegation's value mixes the timely branch and the fallback and pays the fee on both, so the expert's conditional accuracy cannot answer the question alone; the break-even on-time probability is where the two values meet.

Equation (27.2), written in LaTeX: \operatorname{route}^{\star}(x)=\begin{cases}\operatorname{Del}(E,\Delta), & p_E(x)\geq \max_y p(y\mid x),\\\arg\max_y p(y\mid x), & p_E(x)<\max_y p(y\mid x).\end{cases}

Scroll sideways for the whole equation

x is the case. The classifier's best class has probability max_y p(y | x) of being correct. p_E(x) is the probability that the designated expert E is correct on this case. Del(E, Delta) means delegate to E with maximum wait Delta. The arg max picks the most probable class. For the value comparison, a correct result is worth 10 and a wrong one -10, delegating costs a fee of 1 whether or not a reply arrives, and t is the probability the reply comes on time; a missed deadline triggers an authorized fallback worth 2. Probabilities are between 0 and 1.

Predict first. Pick the workbench case (acting is right 0.80, expert 0.95). With a guaranteed on-time return delegation wins 8 against 6. Now set the on-time probability to 0.60. Which route wins?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Defer by comparing correctness, then price the deadline. Left, two bars: expert correctness 0.95 and classifier best class 0.78, with the route chosen by Equation (27.2) marked. Right, net expected utility of acting now, 5.60, against delegating with an on-time probability of 1.00, 8.00.
Case: Chapter: model 0.78, expert 0.95, Probability the reply arrives on time (t): 1.00 (guaranteed)
Constructed example: the chapter's two cases (model 0.78, expert 0.95 or 0.60) and the workbench's timely-return problem (acting 0.80, utilities 10 and -10, fee 1, fallback 2).

Calculated values

Classifier best class probability
0.78
Expert correctness p_E
0.95
Route by Equation (27.2), delay ignored
Delegate to E
Value of acting now
5.60
Value of delegating
8.00
Route with the higher value
delegate
Timely-return probability that ties them
0.657

Equation (27.2): 0.95 >= 0.78, so it delegates (gap 0.95 - 0.78 = 0.17). Act = 0.78 x 10 + 0.22 x (-10) = 5.60. Delegate = t x gross + (1 - t) x 2 - 1 = 1.00 x 9.00 + 0.00 x 2 - 1 = 8.00, where gross = 0.95 x 10 + 0.05 x (-10) = 9.00. The two readings agree: delegate. Break-even: t = (act - 1) / (gross - 2) = (5.60 - 1) / (9.00 - 2) = 0.657.

Worked steps

  1. Equation (27.2): 0.95 >= 0.78, so it delegates (gap 0.95 - 0.78 = 0.17).
  2. Act = 0.78 x 10 + 0.22 x (-10) = 5.60.
  3. Timely review gross = 0.95 x 10 + 0.05 x (-10) = 9.00.
  4. Delegate = 1.00 x 9.00 + 0.00 x 2 - 1 = 8.00; the fee is paid on both branches.
  5. Higher value: delegate (5.60 against 8.00).

Use the idea

A single confidence threshold such as 0.80, below which everything is sent to review, is justified only if the destination's correctness and timeliness are the same for all the cases it covers. Otherwise ask how likely the actual destination is to be right on this kind of case, how likely it is to answer in time, and compare with acting.

Where the conclusion applies

Zero-one loss for Equation (27.2), one expert, and a known expert probability for this case type. The classifier's confidence is assumed calibrated. The utilities, fee and fallback value come from the workbench's constructed problem and are applied here to all four cases for comparison. The expert's accuracy is conditional on a timely reply and is not a permanent rating; if the probability is only a rough average, or the packet omits the evidence the expert needs, the comparison is no better than its inputs.

Common wrong turn: Delegate whenever the model is uncertain
The chapter calls this an intuitive but mistaken rule. Uncertainty is only one side of a comparison: an uncertain model may still be better than the available destination, and a confident model may be worse than a destination with decisive side information.
What this does not settle

The chapter does not show how to estimate every action value, reviewer-error rate, queue delay, observation value, or authority boundary in a live institution.

Chapter 27 source: "What this does not settle".

Check your understanding: A model is 0.97 sure and an expert is right with probability 0.95. Which route does Equation (27.2) choose, and by how much?
0.95 < 0.97, so the route is the most probable class. The classifier leads by 0.97 - 0.95 = 0.02.

Chapter 27 source: section "A value comparison, not a confidence threshold". Demonstration C27-D01.

Demonstration 2 of 4

Four routes, but only the authorized ones can run

Which routes are in the authorized set, which of them has the highest net value, and what removes a route?

Equation (27.1) first filters the actions: a route is available only if it is authorized and its preconditions hold. Equation (27.3) then takes the highest-value route among those. Timely review has gross value 8.04, and the hold turns that into 0.95 x 8.04 + 0.05 x (-4) - 1 minus its own cost, so its value falls one point for each point of hold cost. An excluded route is not low-scoring, it is absent: if it would have won, it stays a recommendation.

Equation (27.1), written in LaTeX: \mathcal A_{\mathrm{auth}}(x)=\{a\in\mathcal A:\operatorname{Authorized}_{\mathcal G}(a,x)=1, \operatorname{Pre}(a,x)=1\}.

Equation (27.3), written in LaTeX: \mathcal R_{\mathrm{auth}}(x)=\{(\mathrm{act},a),(\mathrm{delegate},d,\Delta),(\mathrm{wait},o,\Delta),(\mathrm{narrow},a'),(\mathrm{return})\ \text{that are authorized at }x\}, \operatorname{choose}^{\star}(x)\in\arg\max_{\varrho\in\mathcal R_{\mathrm{auth}}(x)}V(\varrho,x).

Scroll sideways for the whole equation

A route is one choice such as release now, review, wait or hold. The four rows are the chapter's act (release now), delegate (review), wait and narrow (the reversible hold); returning control is not given a value here. A_auth(x) is the set of actions in state x with Authorized = 1 (a declared grant covers them) and Pre = 1 (their state-specific preconditions hold, here a reviewer being available). R_auth(x) is the set of authorized routes, and V(route, x) is the route's net value in constructed utility units. Fraud has probability 0.10. A hold makes timely review likely (0.95) and costs the amount you choose; the timeout branch is worth -4. Review itself costs 1.

Predict first. The hold costs 2 by default and wins. Raise its cost to 5 with every route authorized. Which route has the highest value?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Four routes, but only the authorized ones can run. Horizontal bars for four routes with net values release -1.00, review 1.16, wait 2.00 and hold 4.44. Each carries its Authorized and Pre flags; no route is excluded. The highest authorized route is reversible hold and review; return on timeout.
Cost of the reversible hold: 2, Authority and preconditions: All authorized, reviewer available
Constructed example: the book's own route-value table for the pending transfer (values -1, 1.164, 2 and 4.438), with the hold cost and the authority status varied.

Calculated values

Release now
-1.000
Review; release on timeout
1.164
Wait for confirmation; release on timeout
2.000
Reversible hold and review; return on timeout
4.438
Routes in the authorized set
release now, review, wait for confirmation, reversible hold and review
Highest among the authorized routes
Reversible hold and review; return on timeout
Hold cost at which waiting (2) overtakes it
4.438

Hold = 0.95 x 8.04 + 0.05 x (-4) - 2 - 1 = 4.438. Timely review gross value = 0.90 x [0.98 x 10 + 0.02 x (-10)] + 0.10 x [0.94 x 0 + 0.06 x (-100)] = 8.64 + (-0.60) = 8.04. The hold wins: 4.438 is above waiting at 2. Its value before its own cost is 0.95 x 8.04 + 0.05 x (-4) - 1 = 6.438, so its cost can rise to 6.438 - 2 = 4.438 before waiting overtakes it.

Worked steps

  1. Timely review gross = 8.64 + (-0.60) = 8.04.
  2. Net values: release -1.000, review 1.164, wait 2.000, hold 4.438.
  3. Hold = 0.95 x 8.04 + 0.05 x (-4) - 2 - 1 = 4.438.
  4. Status: All four routes authorized and the reviewer is available.
  5. Authorized set: release now, review, wait for confirmation, reversible hold and review.
  6. Highest value inside the set: Reversible hold and review; return on timeout at 4.438.

Use the idea

Write the table for a consequential action: one row per route, each with its payoff, delay cost and fallback, and next to each row the grant that authorizes it and the precondition it needs. Then keep only the rows that are actually permitted. A delegation row also needs a named destination, a deadline and a fallback; without them it is not a route.

Where the conclusion applies

Utilities, fraud probability, review error rates (detect fraud 0.94, wrongly reject 0.02) and response probabilities are the chapter's constructed values, and all four routes are costed on one scale. Real stakes may not fit one scale; a hard boundary may belong in the authorized set rather than in a cost. Changing any stated probability can change the winner. The reviewer accuracy here is held fixed, although in practice it can change with workload.

Common wrong turn: A high-scoring route can be executed
The chapter says a high-scoring action absent from the authorized set remains a recommendation, not an executable act, and that the system can recommend an excluded action but cannot execute it merely because it ranked highly. Reversibility is not free either: a hold has its own cost.
What this does not settle

The chapter does not make human review an oracle, nor does it rank every refusal above every act. Its bounded conclusion is structural: delegation requires a named route and must compete with acting, waiting, narrowing and returning control.

Chapter 27 source: "What this does not settle".

Check your understanding: If the hold cost 4 instead of 2, what is its net value and does it still beat waiting?
0.95 x 8.04 + 0.05 x (-4) - 4 - 1 = 7.638 - 0.2 - 5 = 2.438, which is above waiting at 2, so it still wins narrowly.

Chapter 27 source: section "Queue economics: what timely review is worth". Demonstration C27-D02.

Demonstration 3 of 4

Delegating more can make review slower than the deadline

How does sending a larger share of tasks to one reviewer change the mean time in review, and which tasks can then be released?

For one reviewer in the M/M/1 model, with random (Poisson) arrivals and exponentially distributed service times, the mean time in the system is 1/(mu - f x lambda) as long as the review arrivals stay below capacity. The time grows slowly at first and then steeply as the load approaches mu. At or beyond mu no stationary (long-run) mean exists and the backlog grows. The ledger shows what that does to the release contract: authority and a review packet are separate requirements from timing.

Equation, written in LaTeX: 1/(\mu-f\lambda)

Scroll sideways for the whole equation

lambda is the arrival rate of tasks per hour. f is the delegation fraction, the share sent to review, so the reviewer sees f x lambda per hour. mu is the number of reviews the reviewer finishes per hour. The result is the mean hours a case spends waiting plus being reviewed. A stationary mean is that long-run average. In the ledger a task needs review when the agent is not authorized or its risk is above the agent's risk limit; it is released only with human authority, a received review packet and a mean time inside its deadline.

Predict first. With the default case (reviewer finishes 3 per hour), what happens to the mean time as the delegated share goes from 0.5 to 0.7? And at 0.9?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Delegating more can make review slower than the deadline. Left, mean time in the review queue against the delegation fraction, with the deadline lines and a marker at f = 0.50 showing 1.00 hours. Right, a ledger grid for 3 tasks with 2 released.
Case: Default: 4 per hour in, 3 per hour out, three tasks, Delegation fraction (f): 0.5
Constructed example: the notebook's default, changed and transfer cases and the chapter's queue example (4 tasks per hour, a reviewer finishing 3 per hour), computed with the laboratory's review-queue function, plus one faster reviewer defined for the reader.

Calculated values

Review arrivals f x lambda (per hour)
2.00
Reviewer capacity mu (per hour)
3
Mean time in system
1.00 hours (60 minutes)
Tasks released
2 of 3
Task outcomes
routine: autonomous; release: released after review; sensitive: blocked (no review packet)
Chance the first delegated case meets its deadline
0.865

Review arrivals = 0.50 x 4 = 2.00 per hour, below 3. Mean time in system = 1 / (3 - 2.00) = 1 / 1.00 = 1.00 hours, which is 60 minutes, outside the 20-minute window. The reviewer's accuracy did not change; only the load did. If the M/M/1 first-come-first-served assumption held exactly, a delegated case would meet a deadline of 2.0 hours with chance 1 - exp(-1.00 x 2.0) = 0.865; a mean does not promise any single case meets its deadline. Ledger: routine: autonomous; release: released after review; sensitive: blocked (no review packet).

Worked steps

  1. Review arrivals = f x lambda = 0.50 x 4 = 2.00 per hour; capacity mu = 3.
  2. Stable, since 2.00 < 3: mean = 1 / (3 - 2.00) = 1.00 hours.
  3. Tasks above the agent risk limit 0.10 or without agent authority need review: release, sensitive.
  4. A delegated task is released only with human authority, a received review packet and a mean inside its deadline.
  5. Result: 2 of 3 tasks released.

Use the idea

Before promising a review deadline, compare the planned review load with the reviewer's capacity. A rule such as delegate everything doubtful can overload the one route that was supposed to supply safety. Keep a route ledger per destination, with case type, deadline, packet completeness, response state and final action, so a failure can be placed: a bad recommendation, a wrong destination, a late answer or an unowned escalation.

Where the conclusion applies

The M/M/1 model: one reviewer, independent random (Poisson) arrivals at a steady rate, exponentially distributed service times with a steady rate, and a long-run mean. Other service-time shapes give a different formula. Batching, priorities, correlated arrivals or a changing reviewer break the model. A mean does not promise that any single case meets its deadline, and the laboratory's deadline test is a planning diagnostic, not a probability. The faster reviewer is defined for this reader.

Common wrong turn: A queue mean proves the deadline is met or missed
The chapter says a mean does not promise that any one case meets its deadline, and the notebook's output is a planning diagnostic. Individual timing needs observed timestamps; the queue formula is a capacity-design check.
Check your understanding: With mu = 4 per hour and f = 0.75, what is the mean time in review in minutes?
Review arrivals = 0.75 x 4 = 3 per hour. Mean = 1 / (4 - 3) = 1 hour, which is 60 minutes.

Chapter 27 source: section "Queue economics: what timely review is worth". Demonstration C27-D03.

Demonstration 4 of 4

Waiting needs an observation worth its delay and a fallback everywhere

When is it right to wait for a confirmation instead of acting now, and what if one state reachable by the deadline has no authorized fallback?

If the confirmation arrives, the transfer is handled correctly for a value of 9; if not, the fallback releases it at -1. So the gain over releasing now is 10 x q, a straight line. Equation (27.4) allows waiting only when that gain is strictly above the delay cost and the fallback is permitted in every reachable deadline state. One state without a permitted fallback is enough to fail the test, however valuable the observation.

Equation (27.4), written in LaTeX: \operatorname{wait}^{\star}(o,\Delta,x) \text{only if} \operatorname{VOI}(o)>\operatorname{DelayCost}(\Delta,x) \text{and} \operatorname{fallback}(\Delta,x_\Delta)\in\mathcal A_{\mathrm{auth}}(x_\Delta)\ \text{for every reachable }x_\Delta.

Scroll sideways for the whole equation

VOI(o) is the value of the observation o: here the extra expected value of waiting for the signed confirmation compared with releasing now. DelayCost is what the wait costs (2 units). q is the probability that the confirmation arrives in time. The fallback is what the controller does at the deadline, and it must be an authorized action in every state reachable by then: the observation arrives in time, arrives late, the case changes, or the source fails.

Predict first. At q = 0.20, with a fallback authorized in every state, does the test allow waiting?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Waiting needs an observation worth its delay and a fallback everywhere. Left, a line of value of information against arrival probability with the delay cost at 2 and a marker at q = 0.5. Right, four boxes for the states reachable by the deadline, all marked fallback authorized.
Probability the confirmation arrives in time (q): 0.5, State with no authorized fallback: None: authorized in all four
Constructed example: the chapter's wait route (value 9 on arrival, -1 at the fallback, delay cost 2) with the arrival probability and the missing-fallback state varied.

Calculated values

Value of information (VOI)
5.00
Delay cost
2.00
VOI minus delay cost
3.00
Fallback authorized in every reachable state
yes
Passes Equation (27.4)
Yes (waiting allowed)
Break-even arrival probability
0.20

VOI = 0.5 x 9 + 0.5 x (-1) - (-1) = 5.00, and VOI - delay cost = 5.00 - 2 = 3.00. Both conditions hold: the observation is worth more than the delay and a fallback is permitted in all four reachable states, so waiting passes. The equation is a necessary test (the word 'only if'): passing it does not prove waiting is the best route, which Equation (27.3) still decides.

Worked steps

  1. Gain from waiting over releasing now: VOI = 0.5 x 9 + 0.5 x (-1) - (-1) = 5.00.
  2. Condition 1: VOI 5.00 against delay cost 2, needing strictly more: holds.
  3. Reachable states by the deadline: in time, late, the case changes, the source fails.
  4. Condition 2: the fallback must be authorized in every one: all four are.
  5. Result: Yes (waiting allowed).

Use the idea

When an agent says it will wait, ask four things: which fact would change the action, who supplies it, by when, and what happens if it never comes. If any answer is missing, there is no waiting action, only delay. If the observation source has become unavailable, move to the declared fallback instead of waiting under a plan that assumes its arrival.

Where the conclusion applies

One observation and the chapter's constructed values (9, -1, delay cost 2). VOI is read here as the gain over releasing now, a simple special case of the Chapter 8 idea. The four reachable states are the ones the chapter lists; a real case can have more. Equation (27.4) is a necessary test: passing it does not make waiting the best route.

Common wrong turn: Waiting is just not acting
The chapter says waiting with no identified observation, no deadline and no fallback is not an information action; it is delay without a model. An agent that cannot say what missing fact would change the action, who supplies it, by when and what happens otherwise has named uncertainty without resolving authority.
What this does not settle

The chapter does not prove that a learned rejector yields fair, lawful, safe, or corrigible deployment, and it does not show how to estimate every observation value or authority boundary in a live institution.

Chapter 27 source: "What this does not settle".

Check your understanding: If the delay cost rose to 3 and the confirmation still arrived with probability 0.5, does the test allow waiting (fallback authorized everywhere)?
VOI = 0.5 x 9 + 0.5 x (-1) - (-1) = 5, and 5 > 3, so waiting passes. It would fail at or below q = 3 / 10 = 0.30, because equality fails the strict test.

Chapter 27 source: section "Waiting is not hiding". Demonstration C27-D04.