The Mathematics of AI Agents, laboratory reader ยท Chapter 22

Safe Enough to Act

A high reward does not make an action permitted: keep task return, measured cost and authority as separate records.

These four demonstrations follow the chapter from the choice between a fixed penalty and a declared limit, through what a one-update bound and a margin do and do not promise, to a tail measure, a spending budget and a gate that decides which actions may be considered at all. Each one changes a declared value and shows which rule reacts and which does not.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

A penalty hides an exchange rate

With the same two policies, when does a fixed penalty choose a policy that a cost limit would remove?

Each policy's score falls with the multiplier at the rate of its own cost, so the policy with more cost loses ground faster. The lines cross at the reward gap divided by the cost gap. The constraint rule never uses the multiplier: it draws a boundary in the reward-cost plane and only policies on the allowed side are eligible, as in Figure 22.2. In the transfer case the multiplier would have to reach 2100 before the penalty rule gave up act, although act already breaks a limit of 0.

Equation (22.2), written in LaTeX: \operatorname{Score}_{\lambda_{\mathrm{safe}}}(\pi)=\operatorname{Ret}^{\pi}-\lambda_{\mathrm{safe}}\operatorname{Cost}_1^{\pi}.

Equation (22.1), written in LaTeX: \begin{aligned}\text{maximize}_{\pi} & \operatorname{Ret}^{\pi}\\\text{subject to} & \operatorname{Cost}_i^{\pi}\le d_i \text{for every declared } i.\end{aligned}

Scroll sideways for the whole equation

Ret is the expected reward of a policy and Cost its expected measured cost. The multiplier lambda_safe is the exchange rate: how many reward units one cost unit is worth. d_1 is the declared cost limit. The penalty rule picks the larger blended score; the constraint rule keeps only policies with cost at most the limit, then picks the larger reward. The chapter's table has Careful at reward 8 and cost 0.5, Aggressive at 12 and 1.5, limit 1. Workbench VI.1 has Careful at 7 and 0.5, Aggressive at 11 and 2, limit 1. The laboratory's transfer case has act at reward 20 and cost 0.01, wait at -1 and 0, limit 0. The best mix is the largest chance of the costlier policy that keeps expected cost within the limit, if randomising once before the episode is allowed (the workbench's reading).

Predict first. In the chapter's table with limit 1, raise the multiplier through 2, 4 and 6. What does the penalty rule do, and does the constraint rule ever switch?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A penalty hides an exchange rate. Left: a reward against cost plane with the two policies and a dashed cost limit at 1; constraint rule picks Careful. Right: bars of blended score at multiplier 2; penalty rule picks Aggressive.
Policies and limit: Chapter table (limit 1), Penalty multiplier: 2
Constructed example: the chapter's own two-policy table, Mathematical Workbench problem VI.1 (Careful 7 and 0.5, Aggressive 11 and 2) and the laboratory's transfer case (act and wait, limit 0, multiplier 100). The transfer case is also evaluated with the laboratory's risk and authority function; the other two have costs above 1 and are computed directly.

Calculated values

Careful score
7.00
Aggressive score
9.00
Penalty rule picks
Aggressive
Policies within the limit
Careful
Constraint rule picks
Careful
Multiplier at which the penalty rule flips
4
Best randomised mix
workbench reading, outside the policy family: probability 1/2 of Aggressive, reward 10.00

Careful: 8 - 2 x 0.5 = 7.00; Aggressive: 12 - 2 x 1.5 = 9.00. The penalty rule picks Aggressive; the flip point is (12 - 8) / (1.5 - 0.5) = 4 / 1 = 4, and a multiplier above it favors Careful. Aggressive costs 1.5, above the limit 1, so it is removed whatever it earns; the constraint rule picks Careful. If randomising is allowed, expected cost 0.5 + p x 1 must stay at most 1, so p = 1/2 and the reward is 8 + 1/2 x 4 = 10.00. The two rules read the same returns and answer different questions.

Worked steps

  1. Score of Careful: 8 - 2 x 0.5 = 7.00.
  2. Score of Aggressive: 12 - 2 x 1.5 = 9.00.
  3. Penalty rule: Aggressive has the larger score.
  4. Flip multiplier = (12 - 8) / (1.5 - 0.5) = 4 / 1 = 4.
  5. Limit d = 1: policies with cost at most d are Careful.
  6. Constraint rule: Careful.

Use the idea

When a team reports that a penalty setting keeps cost down, ask whether cost is a preference with a defensible exchange rate or a limit. If it is a limit, state it as a limit and check each policy against it.

Where the conclusion applies

One step, two policies, one cost, and expected values read as the whole story. Each flip point belongs to its two policies and reward scale, not to any other problem. Mixing assumes the choice of policy is made once before the episode so rewards and costs mix linearly; it does not make a forbidden action permitted. Even the constrained choice says nothing about a single bad episode or a hazard the cost leaves out.

Common wrong turn: A large enough penalty works as a limit
The chapter warns against presenting a fixed penalty as a satisfaction guarantee for a limit it was never designed to enforce. In the transfer case at multiplier 100 the penalty rule still picks act: 20 - 100 x 0.01 = 19.00 beats -1.00, although act's cost 0.01 breaks the limit 0.
Check your understanding: Work this one by hand (3 is not an offered multiplier). In workbench VI.1 with the limit at 1 and the multiplier at 3, which policy does the penalty rule pick, and which does the constraint rule pick?
Careful: 7 - 3 x 0.5 = 5.5. Aggressive: 11 - 3 x 2 = 5. The penalty rule picks Careful (the flip point is 4 / 1.5 = 2.67, and 3 is above it). Aggressive costs 2, above the limit 1, so the constraint rule picks Careful too: here the two rules agree.

Chapter 22 source: section "Fixed penalties make a hidden exchange rate". Demonstration C22-D01.

Demonstration 2 of 4

A bound permits excess, so leave a margin

How much can the modeled cost exceed the limit after one update, how large a step does a margin allow, and when does the margin cover it?

The allowance is sqrt(2 delta) times gamma times epsilon, divided by (1 - gamma) squared. A larger trust region lets the update treat its local picture as informative farther away, which raises the allowance in proportion to the square root of the step, and a discount nearer 1 blows it up through the squared denominator. The margin is a separate choice: it only helps when it is at least as large as the allowance, and the right-hand curve shows the largest step it covers.

Equation (22.3), written in LaTeX: \operatorname{Cost}_i^{\pi_{k+1}}\le d_i+\frac{\sqrt{2\delta} \gamma \epsilon_i}{(1-\gamma)^2}.

Equation (22.4), written in LaTeX: \operatorname{Cost}_i^{\pi}\le d_i-\varepsilon_{\mathrm{safe}}, \varepsilon_{\mathrm{safe}}>0.

Scroll sideways for the whole equation

d is the cost limit (10 here). delta is the step-size parameter of one update, the trust-region size measured by a KL-divergence condition. gamma is the discount factor. epsilon is the largest absolute expected constraint advantage over states (0.1 here). The fraction after d is the allowance, the most the bound lets the cost return rise above d. The margin, written epsilon_safe in the chapter, is room reserved below d (0.5 or 1.0 here), so the optimizer aims at the ceiling d minus margin.

Predict first. At discount 0.85 and margin 0.5, will a step size of 0.005 fit inside the margin? What about 0.02?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A bound permits excess, so leave a margin. Left: two bars from the aimed cost level to the worst modeled cost; aiming at the limit ends at 10.38 and aiming at the ceiling at 9.88. Right: the allowance rising with step size for discount 0.85, with the margin 0.5 as a horizontal line and the marker at step 0.005 with allowance 0.378.
Step size delta: 0.005, Discount factor gamma: 0.85, Safety margin: 0.5
Constructed example: the chapter's bound evaluated with values defined for this reader (limit 10, advantage 0.1, margins 0.5 and 1.0).

Calculated values

Allowance above the aimed level
0.378
Worst case when aiming at the limit
10.378
Ceiling d minus margin
9.5
Worst case when aiming at the ceiling
9.878
Margin covers the allowance
yes
Largest step the margin covers
0.0088

Allowance = sqrt(2 x 0.005) x 0.85 x 0.1 / (1 - 0.85)^2 = 0.10 x 0.85 x 0.1 / 0.0225 = 0.378. Aiming at the limit: 10.0 + 0.378 = 10.378. Aiming at the ceiling, which means the bound is applied with the ceiling in place of d: 9.5 + 0.378 = 9.878. The allowance is below the margin 0.5, so the modeled worst case stays within d. Setting the allowance equal to the margin gives a largest covered step of ((0.5 x 0.0225) / (0.85 x 0.1))^2 / 2 = 0.0088. A bigger step or a longer horizon (discount nearer 1) enlarges the allowance; the bound permits excess, it does not rule it out. The lower row is this reader's illustration of a margin, not a result of the chapter.

Worked steps

  1. Root: sqrt(2 x 0.005) = 0.100.
  2. Allowance = 0.100 x 0.85 x 0.1 / 0.0225 = 0.378.
  3. Aim at the limit: 10.0 + 0.378 = 10.378.
  4. Ceiling = 10.0 - 0.5 = 9.5; worst case 9.5 + 0.378 = 9.878.
  5. Margin 0.5 against allowance 0.378: covered.
  6. Largest covered step = 0.0088.

Use the idea

Before citing the proposition to justify a safety margin, compute the allowance for your own step size, discount and advantage, and set the margin from that number rather than from habit.

Where the conclusion applies

The bound concerns one ideal update inside the modeled problem; it is not a statement about a training run, a sampled update or a changed deployment. The values here are constructed. The lower row applies the bound with the ceiling in the role of the limit (d replaced by d minus the margin), which is this reader's illustration of the margin idea. The target (what the optimizer asks for), the bound (what the theorem permits) and observed behavior (what an experiment records) are three different things; this figure shows the first two only.

Common wrong turn: Constrained means the cost stays below the limit
Proposition 2 gives a bound on how far above the limit the new cost return may be, not a promise that it is below the limit. The right side of Equation (22.3) is d plus a nonnegative allowance, so aiming at the limit always leaves a positive modeled overshoot.
What this does not settle

CPO's one-update result does not prove that a whole training run, or an altered deployment, satisfies a constraint. A clear optimization claim can grant limited authority only within the threat model, measurement, margin and update assumptions that make it true.

Chapter 22 source: "What this does not settle".

Check your understanding: Work this one by hand, using the control for the margin. With step size 0.005, discount 0.9 and advantage 0.1, what is the allowance, and does a margin of 0.5 cover it? What about a margin of 1.0?
sqrt(2 x 0.005) = 0.1, so 0.1 x 0.9 x 0.1 / (1 - 0.9)^2 = 0.009 / 0.01 = 0.9. That is above 0.5, so a margin of 0.5 does not cover it, and below 1.0, so a margin of 1.0 does.

Chapter 22 source: section "A constraint is not a promise". Demonstration C22-D02.

Demonstration 3 of 4

Two other risk contracts: a tail and a budget

For a rare bad episode, how far apart are mean cost and tail severity, and how does a risk budget shrink as charged actions spend it?

For this lottery the bracket is a bent line in z, so its smallest value sits at z = 0 or z = 100. At z = 0 it equals the mean divided by the tail share; at z = 100 it equals 100. CVaR is the smaller of the two, so a tail wider than the bad episode dilutes it with zeros. The budget is a different contract: it is an accounting state that falls by each charge, and a charge larger than what remains cannot be taken, even when a smaller one proposed later still fits.

Equation (22.5), written in LaTeX: \operatorname{CVaR}_{\alpha}(C)= \inf_{z\in\mathbb R}[z+\frac{1}{1-\alpha} \mathbb E[(C-z)_+]].

Equation (22.6), written in LaTeX: b_{t+1}=b_t-\widehat r_t(a_t), b_0=B, \widehat r_t(a_t)\le b_t.

Scroll sideways for the whole equation

C is the cost of one episode: 100 with probability 0.01, otherwise 0 (the chapter's lottery). alpha sets the tail: the worst share 1 - alpha of episodes. z is a cutoff, and (C - z)+ is the amount by which C exceeds z, or 0. CVaR is the smallest value the bracket takes as z varies, which equals the average cost over the worst 1 - alpha share. In the budget, B is the stated cumulative risk budget, b_t what remains before decision t, and r-hat the declared conservative charge of the chosen action; the four charges 0.04, 0.05, 0.03 and 0.05 are values defined for the reader, and a refused action is replaced by a fallback that charges 0.

Predict first. For a bad episode with probability 0.01, is the tail average larger at alpha 0.9 or at alpha 0.99?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Two other risk contracts: a tail and a budget. Left: the objective of Equation 22.5 against the cutoff, with its minimum (CVaR 100.0), the mean cost 1.0 and the worst outcome 100 marked. Right: four decisions with budget B 0.10; 2 charges are accepted and 2 refused, leaving 0.01.
Tail level alpha: 0.99, Risk budget B: 0.1
Constructed example: the chapter's lottery (cost 100 with probability 0.01, else 0) with other tail levels, and a four-decision budget whose charges are defined for this reader.

Calculated values

Mean cost
1.0
Tail share 1 minus alpha
0.01
Tail average (CVaR)
100.0
Budget B
0.10
Charges accepted
2 of 4
Budget left at the end
0.01

Mean cost = 0.01 x 100 + 0.99 x 0 = 1.0. Tail share = 1 - 0.99 = 0.01. Objective at z = 0: 0 + 1.0 / 0.01 = 100.0; at z = 100: 100 + 0 / 0.01 = 100.0. Both corners give 100.0, so the minimum is flat and CVaR is 100.0: the tail is exactly the bad episode. The curve rises above the top of the left plot (the axis stops at 230; the curve reaches 1090 at z = -10). Budget B = 0.10, with a refused action replaced by a fallback that charges 0: t = 1: 0.04 <= 0.10, charged, b = 0.10 - 0.04 = 0.06; t = 2: 0.05 <= 0.06, charged, b = 0.06 - 0.05 = 0.01; t = 3: 0.03 > 0.01, refused, b stays 0.01; t = 4: 0.05 > 0.01, refused, b stays 0.01. Spent 0.09 of 0.10, left 0.01.

Worked steps

  1. Mean cost = 0.01 x 100 = 1.0; tail share = 0.01.
  2. Cutoff z = 0 gives 1.0 / 0.01 = 100.0; z = 100 gives 100.0.
  3. CVaR is the smaller corner: 100.0.
  4. Budget t = 1: 0.04 <= 0.10, charged, b = 0.10 - 0.04 = 0.06.
  5. Budget t = 2: 0.05 <= 0.06, charged, b = 0.06 - 0.05 = 0.01.
  6. Budget t = 3: 0.03 > 0.01, refused, b stays 0.01.
  7. Budget t = 4: 0.05 > 0.01, refused, b stays 0.01.
  8. Left at the end: 0.01.

Use the idea

If one rare, severe episode is the worry, a limit on mean cost can look small while the tail measure is large. State which one the limit is about, and why. For a bounded horizon, keep a running budget and refuse a charged action that exceeds what remains.

Where the conclusion applies

A single two-outcome lottery and four charges defined for the reader. CVaR asks for average severity in a declared upper tail, not a promise that no episode is bad. The budget rule prevents one action from exceeding the remaining authority; it does not prove the charges are calibrated, independent, complete or safe beyond the declared horizon.

Common wrong turn: A tail measure promises that no episode is bad
The chapter says CVaR asks for average severity in a declared upper tail, not a promise that no episode is bad, and that none of the measures covers a hazard the cost variable omits. Here CVaR is 100 at alpha 0.99, yet the worst episode still happens with probability 0.01.
What this does not settle

Expected auxiliary-cost constraints are not sufficient for every safety problem: a rare catastrophic outcome, a distribution shift, an unobserved state, a human authorization boundary or an adversarial interaction may require a different risk measure, monitoring system, physical safeguard or refusal rule.

Chapter 22 source: "What this does not settle".

Check your understanding: Work this one by hand (a budget of 0.08 is not offered). With budget 0.08 and the charges 0.04, 0.05, 0.03, 0.05 proposed in order, which are charged and what is left? And a bad episode costs 100 with probability 0.02: what is CVaR at alpha 0.97?
Budget: 0.04 <= 0.08 is charged, leaving 0.04; 0.05 > 0.04 is refused; 0.03 <= 0.04 is charged, leaving 0.01; 0.05 > 0.01 is refused. Left 0.01. CVaR: tail share 0.03; at z = 0 the bracket is 2 / 0.03 = 66.67, at z = 100 it is 100, so CVaR is 66.67 while the mean cost is 2.

Chapter 22 source: section "Other risk contracts". Demonstration C22-D03.

Demonstration 4 of 4

Authority first, risk limit second, reward last

Which actions may be considered at all, which of those has the best reward, and what happens when one route to the actuator skips the check?

The penalty rule scores every row by reward minus penalty times risk, so a high-reward forbidden row can win. The gate removes rows before any score is used: authority is a yes or no condition, so no limit or penalty setting grants it. A looser risk limit can bring back an over-limit row but never a forbidden one. A gate only constrains routes that pass through it: with a bypass, the check exists but the unauthorized row can still run.

Equation (22.1), written in LaTeX: \begin{aligned}\text{maximize}_{\pi} & \operatorname{Ret}^{\pi}\\\text{subject to} & \operatorname{Cost}_i^{\pi}\le d_i \text{for every declared } i.\end{aligned}

Equation (22.2), written in LaTeX: \operatorname{Score}_{\lambda_{\mathrm{safe}}}(\pi)=\operatorname{Ret}^{\pi}-\lambda_{\mathrm{safe}}\operatorname{Cost}_1^{\pi}.

Scroll sideways for the whole equation

Each action has a reward, an expected adverse-event probability (risk) and a yes or no current permission. The risk limit is the largest probability allowed. The penalty is the reward deducted per unit of risk. The gate keeps only actions that are authorized and within the risk limit; the gate pick is the highest reward among those. In the displayed equations, Ret is the reward, Cost is the risk, d_i is the risk limit and lambda_safe is the penalty, with each action standing in for a policy pi. Authority, the yes or no permission, is an extra condition that the equations do not show. A bypass is a route to the actuator that does not pass the check.

Predict first. Raise the risk limit from 0.1 to 0.25. Does fast-release become eligible? Which action does the gate now pick?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Authority first, risk limit second, reward last. Left: reward against risk for four actions with a dashed risk limit at 0.1; the gate picks reviewed-release. Right: blended score of each action at penalty 5; the penalty rule picks fast-release. Route: every route is checked.
Risk limit: 0.1, Penalty per unit of risk: 5, Routes to the actuator: Every route passes the check
Constructed example: the laboratory's four-action example, evaluated with the laboratory's risk and authority function. The four rewards, risks and permissions are not taken from the chapter, whose section supports only the distinction between an objective and a gate. The bypass is defined for this reader.

Calculated values

Penalty pick over all rows
fast-release
Penalty pick among authorized rows
risky-authorized
Gate pick (authority, then limit, then reward)
reviewed-release
Eligible actions
2 of 4
Removed by authority
fast-release
Removed by the risk limit
risky-authorized
Row that executes
reviewed-release

fast-release: 10 - 5 x 0.3 = 8.50; reviewed-release: 5 - 5 x 0.05 = 4.75; risky-authorized: 8 - 5 x 0.2 = 7.00; abstain: 0 - 5 x 0 = 0.00. The penalty rule picks fast-release from all rows, without asking about authority. Eligible rows are reviewed-release (reward 5), abstain (reward 0); the highest reward among them is reviewed-release. Fast-release is not authorized, so no limit or penalty makes it eligible. Every route passes the check, so the row that executes is the gate pick, reviewed-release.

Worked steps

  1. Score every row: reward - penalty x risk (the sums above).
  2. Penalty rule over all rows: fast-release.
  3. Authority removes: fast-release.
  4. Risk limit 0.1 then removes: risky-authorized.
  5. Highest reward among the 2 eligible rows: reviewed-release.
  6. Row that executes: reviewed-release.

Use the idea

For an agent that can call tools, list each action with its permission and risk, remove what is not permitted, then rank the rest. Keep an explicit abstain row, provided abstain is authorized and its risk is within the limit, so the set is never empty. Confirm that every route to the actuator passes the gate.

Where the conclusion applies

Rewards, risks and permissions are supplied values; nothing here estimates risk. An expected-risk limit is a statement about an average, so an action with risk 0.05 can still end badly. A large enough penalty can mimic a gate in one case and fail in the next. In the bypass states the unchecked route is assumed to run the controller's own penalty pick; that is this reader's illustration of the chapter's complete-mediation requirement.

Common wrong turn: A system with only an objective has a gate
The chapter separates a safety objective, which changes how outcomes are ranked (Equation 22.2), from a safety gate, which decides whether an action path may proceed at all. Confusion begins when a system that has only an objective is described as though it had a gate.
What this does not settle

Expected auxiliary-cost constraints are not sufficient for every safety problem: a rare catastrophic outcome, a distribution shift, an unobserved state, a human authorization boundary or an adversarial interaction may require a different risk measure, monitoring system, physical safeguard or refusal rule.

Chapter 22 source: "What this does not settle".

Check your understanding: Work this one by hand (0.2 is not an offered limit). With the risk limit at 0.2 and the penalty at 5, which action does the gate pick, and is fast-release eligible?
Eligible rows are authorized with risk at most 0.2: reviewed-release (0.05), risky-authorized (0.2, which meets the limit exactly) and abstain (0). Rewards 5, 8 and 0, so risky-authorized wins. Fast-release is not authorized, so it stays out.

Chapter 22 source: section "Authority is an operating envelope". Demonstration C22-D04.