The Mathematics of AI Agents, laboratory reader ยท Chapter 25

Improving an Agent Without Trusting the Improvement

An agent can propose a useful change without supplying independent evidence that the change should be released.

These four demonstrations follow one proposed change to an agent's procedure: a canary with its rollback rule fixed first and the weights frozen throughout, why the development winner cannot certify itself, how much a frozen guard can tell you and how quickly repeated use erodes it, and why release needs four separate conditions. All values are constructed teaching examples, not measurements of any real system.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

A canary is monitored evidence, and its rules come first

If an accepted version enters a limited release, when does it roll back, and what changes if the rule is chosen after the numbers are seen?

A canary is a limited release to a declared population, action set, duration and resource budget. Its rollback trigger is set before it starts, here a cost boundary of c minus a margin, together with a named fallback. When the trigger fires the declared fallback applies and the weights were never involved, because Equation (25.1) holds theta fixed. If the boundary is moved after the costs are seen, the same trace can fire late or never.

Equation (25.1), written in LaTeX: \phi'\in\mathcal M(\phi_t,e_t), \theta_{t+1}=\theta_t.

Scroll sideways for the whole equation

theta is the set of model parameters, frozen, so theta at step t + 1 equals theta at step t. phi_t is the parent version of the procedure: the prompt template, tool-routing rule, memory policy, verifier and retry cap assembled around the model. phi' is the accepted new version, so rollback restores a procedure version, not weights. The cost per period is a constructed resource measure. c is the cost limit (1.00) and the margin is 0.10, so the declared boundary is c minus the margin. An alert is the first period with cost strictly above the boundary.

Predict first. Take the creeping-cost trace under the rule declared first. In which period does the alert fire? Then switch to the rule chosen after looking.

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A canary is monitored evidence, and its rules come first. Cost per period against a dashed boundary at 0.90; an alert marker sits at period 5. Beside it, the procedure in service is phi prime for 5 periods and then the restored parent version; the weights row never changes.
Cost trace during the canary: Cost creeping up, When the boundary was set: Declared before the canary (0.90), Fallback named in advance: Restore the parent version
Constructed example: canary cost traces, an eight-period duration and a cost limit with margin defined for this reader around the chapter's rollback and monitoring rules.

Calculated values

Boundary used
0.90 (declared before the canary)
Alert under the declared 0.90
period 5
Alert under the boundary used
period 5
Periods under phi'
5 of 8
Fallback after an alert
restore phi_t
Model parameters changed
0 (theta stays frozen)

Boundary = c - margin = 1.00 - 0.10 = 0.90, fixed before period 1. Costs run 0.80, 0.83, 0.86, 0.89, 0.92; the first above 0.90 is period 5 (0.92 > 0.90), so the alert fires there and the declared fallback, restore the parent version phi_t, applies from period 6. That is 5 of 8 periods under phi' and theta unchanged throughout.

Worked steps

  1. Boundary fixed before period 1: c - margin = 1.00 - 0.10 = 0.90.
  2. Observed cost per period: 0.80, 0.83, 0.86, 0.89, 0.92, 0.95, 0.98, 1.01.
  3. First period above 0.90: period 5 with cost 0.92.
  4. The alert fires in period 5; the declared fallback is restore the parent version phi_t.
  5. theta at step t + 1 equals theta at step t in every period, so rollback only changes the procedure version.

Use the idea

Before a canary starts, write the monitoring cadence, population, metric, alert threshold, maximum duration, decision after an alert, and the fallback route. A canary is monitored evidence after acceptance, not a cheap second chance to pass the guard.

Where the conclusion applies

One constructed cost metric, eight periods, one trigger type. The chapter also names a safety violation, a verifier failure, a decline in a monitored outcome and a missing authority condition as triggers; each needs its own pre-declared rule. A pre-declared boundary protects against moving the rule after the outcomes, but it does not by itself stop leakage through redesign, metric changes or holdout peeking, and a valid sequential procedure is a separate matter.

Common wrong turn: Self-improvement means the model rewrites its weights
The chapter notes that the Reflexion paper calls its method verbal reinforcement learning but does not report a model silently rewriting its weights after every failure. A reflection is text added to memory. The versioned object is the procedure, so theta at step t + 1 equals theta at step t and rollback restores the parent procedure phi_t.
What this does not settle

Training-time model updates are a separate question. This chapter held theta fixed; changing it involves its own training and evaluation controls, not covered here.

Chapter 25 source: "What this does not settle".

Check your understanding: A canary declares its boundary as c - margin with c = 1.20 and margin 0.15. Costs by period are 1.00, 1.05, 1.08 and 1.12. In which period does the alert fire, and what happens if the boundary had been moved to c after looking?
The boundary is 1.20 - 0.15 = 1.05. A cost of 1.05 is not above 1.05, so the first cost above it is 1.08 in period 3. With the boundary moved to 1.20 none of the four costs is above it, so no alert fires.

Chapter 25 source: section "Acceptance is a constrained release decision". Demonstration C25-D01.

Demonstration 2 of 4

Development chooses, the guard checks one frozen pick

Once development has picked a winner, what does a separate guard set decide, and what breaks if the guard picks or is reused?

Development picks the candidate with the best development score. Only that frozen candidate is then compared with the parent on guard cases, case by case, and the average of the differences Z is its guard gain. The other candidate may look better on the guard, but letting the guard choose would make the guard part of selection, so its best-of-several score stops being a test. If the guard was already reused, the gate rejects regardless of the gain.

Equation (25.2), written in LaTeX: \begin{aligned}&\widehat\Delta_D(\phi_j)=\widehat V_D(\phi_j)-\widehat V_D(\phi_t), \phi'=\arg\max_{\phi_j\in\Phi_t}\widehat\Delta_D(\phi_j),\\&Z_i(\phi')\in[-1,1], \widehat\Delta_G(\phi')=\frac{1}{m}\sum_{i=1}^{m}Z_i(\phi').\end{aligned}

Scroll sideways for the whole equation

D is the development set and G the guard set. V-hat_D(phi) is a procedure's estimated success on D, and Delta-hat_D(phi_j) is candidate j's development uplift over the parent phi_t. The arg max picks the winner phi'. For guard case i, Z_i is the candidate's result minus the parent's, between -1 and 1 (here +1, 0 or -1), and Delta-hat_G is the average of Z_i over the m guard cases. The threshold is the smallest guard gain accepted. The candidates are named flashy and steady (and proposal in the transfer case); the names are labels only.

Predict first. In the laboratory's default case with a clean guard, threshold 0.20 and the development winner tested, which candidate is frozen and what happens?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Development chooses, the guard checks one frozen pick. Left, development success rates for flashy, steady with flashy picked. Right, a grid of 4 paired guard cases for flashy: parent outcome, candidate outcome and their difference, averaging 0.00 against a threshold of 0.20; release is rejected (gain below threshold).
Teaching case: Default (winner matches the parent, threshold 0.20), Guard history: Clean (first use), Who picks the candidate to test: Development winner (the rule)
Constructed example: the laboratory's default, changed and transfer teaching cases, computed with its own gate function.

Calculated values

Development rates
flashy 1.00, steady 0.75
Selected on development
flashy
Candidate tested on the guard
flashy
Guard cases m
4
Guard gain of the tested candidate
0.00
Threshold
0.20
Guard status
clean (used once)
Release decision
rejected (gain below threshold)

Development rates: flashy 4/4 = 1.00, steady 3/4 = 0.75; the highest is flashy. Its paired guard differences (candidate minus parent) are 0 + 0 + 0 + 0, so guard gain = (0 + 0 + 0 + 0) / 4 = 0.00. 0.00 is below 0.20, so flashy fails this finite gate.

Worked steps

  1. Development rates: flashy 4/4 = 1.00, steady 3/4 = 0.75.
  2. Development picks flashy and it is frozen.
  3. Paired guard differences Z (candidate minus parent), case by case: 0 + 0 + 0 + 0.
  4. Guard gain = (0 + 0 + 0 + 0) / 4 = 0.00.
  5. Threshold 0.20 is fixed before guard access; 0.00 is below it.
  6. Guard status: clean, used once. Release: rejected (gain below threshold).

Use the idea

Write down which candidate is frozen, and the threshold, before anyone opens guard outcomes. Report guard numbers for the frozen candidate only, and treat any guard that has informed a redesign as development data.

Where the conclusion applies

Two or four paired cases with binary outcomes, as in the laboratory's teaching cases, so the gain can only be a multiple of 0.25 or 0.5 and the gate is a teaching device. A threshold is a rule, not a confidence statement. The reuse flag records what you declare, not the true history, and the threshold is only meaningful if fixed before guard access. A tie in guard gain leaves the development pick in place.

Common wrong turn: The top of the development table certifies the winner
The chapter says the problem is selection, not bad faith: if several changes have no true benefit, random variation can still place one at the top of the table. The winning development result records both the candidate and the selection process that chose it, so it cannot also serve as untouched evidence for that candidate.
What this does not settle

The chapter does not promise a clean separation between procedure and environment. A new tool, memory store, evaluator or task distribution can change the assembled system in ways that demand fresh evaluation.

Chapter 25 source: "What this does not settle".

Check your understanding: A frozen candidate beats its parent on 2 of 8 guard cases, loses on 1 and ties on 5. Its threshold is 0.20. Does it pass a clean guard?
Gain = (2 x 1 + 1 x (-1) + 5 x 0) / 8 = 1/8 = 0.125, which is below 0.20, so it does not pass.

Chapter 25 source: section "Development can choose a candidate; it cannot certify one". Demonstration C25-D02.

Demonstration 3 of 4

How little a guard promises, and how it wears out

How likely is a candidate with no true benefit to clear a guard threshold, and which guard decisions does that number still cover after a guard result reaches the designer?

The bound exp(-m x 0.25^2 / 2) shrinks quickly as guard cases are added. For several candidates, each fixed before it meets the guard, the chances add up, so ten decisions cost ten times the single bound. The exact tail for one concrete no-benefit guard lies below the bound, which is a worst-case guarantee, not a prediction. Once a guard result changes the next candidate, that candidate was not fixed before guard access and the adding-up no longer applies.

Equation (25.3), written in LaTeX: \Pr(\widehat\Delta_G(\phi')\geq 0.25) \leq \exp(-\frac{200(0.25)^2}{2}) \approx 0.00193 \text{for every fixed }\phi'\text{ with }\Delta(\phi')\leq0, 10(0.00193)\approx0.0193.

Scroll sideways for the whole equation

m is the number of guard cases. Each paired difference Z_i lies between -1 and 1 and has true mean Delta, which is at most 0 for a candidate with no benefit. Delta-hat_G is the guard average. The threshold tau is 0.25, as in the chapter. exp(x) is e raised to the power x. The decision count q is how many separately fixed candidates use the guard, each added once. In the redesign states the second guard result is shown to the designer, who revises the third candidate.

Predict first. At 200 guard cases and 10 decisions the total is 0.0193. Switch to the state where guard result 2 reaches the designer. What happens to the total, and what happens to decisions 3 to 10?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How little a guard promises, and how it wears out. Left, a log-scale plot of the bound and one exact tail against guard cases, marked at m = 200 with bound 0.00193. Right, 10 boxes for guard decisions: the first 10 are covered by the bound.
Number of guard cases m: 200, Candidates tested on the guard: 10, Did a guard result reach the designer?: No, every candidate fixed first
Constructed example: the chapter's own 200-case, 0.25 threshold, ten-decision values and its three-query ledger, with other sizes and the redesign state defined for this reader.

Calculated values

Guard cases m
200
Bound, one decision
0.00193
Decisions fixed before guard access
10 of 10
Total bound over those decisions
0.0193
Decisions needing a fresh guard
0
Exact tail, fair plus or minus 1 case
0.000250
Bound divided by exact tail
7.7

Exponent = m x tau^2 / 2 = 200 x 0.0625 / 2 = 6.250, so the bound is exp(-6.250) = 0.00193. Over the 10 decisions fixed before any guard access the total is at most 10 x 0.00193 = 0.0193. Every decision was fixed before it met the guard, which is the condition the bound needs. One concrete no-benefit guard (each paired outcome plus 1 or minus 1 with equal chance) clears 0.25 with exact probability 0.000250, below the bound, as it must be. Every number holds only if the candidate was fixed before guard access, outcomes lie in [-1, 1] and are independent.

Worked steps

  1. Exponent = m x tau^2 / 2 = 200 x 0.0625 / 2 = 6.250.
  2. Bound for one fixed candidate: exp(-6.250) = 0.00193.
  3. Decisions fixed before any guard access: 10 of 10.
  4. Total over those decisions: 10 x 0.00193 = 0.0193.
  5. No guard result reached a designer, so no decision needs a new guard.
  6. Exact tail for one fair plus or minus 1 guard: 0.000250, below the bound 0.00193.

Use the idea

Set a guard query budget before use. Count how many candidates will be tested, size the guard so the total stays small, and retire the guard for release evidence once its results start steering redesign. Track who has seen which outcomes, not whether the bank file changed.

Where the conclusion applies

The chapter's constructed conditions: every counted candidate is fixed before guard access, outcomes lie in [-1, 1], and outcomes are independent across cases. Shared tool state or a common evaluator can break independence. The exact tail uses one convenient distribution, plus or minus 1 with equal chance; it is not a real agent. The guard is evidence for one version, one task population and one evaluation protocol, and a shift in the task distribution changes what a clean result describes.

Common wrong turn: A frozen guard stays fresh while its file is unchanged
The chapter says the word frozen describes access, not a file format. A test bank on a server still becomes development evidence if operators learn which changes passed it and redesign from that information.
Check your understanding: With m = 300 guard cases and threshold 0.25, what is the bound for one decision?
exp(-300 x 0.0625 / 2) = exp(-9.375) = 0.0000848, about 8.5 in 100,000.

Chapter 25 source: section "A guard becomes development evidence when it is consulted repeatedly". Demonstration C25-D03.

Demonstration 4 of 4

Acceptance needs four passes, not a high score

Can a large guard uplift make up for a cost that is over the limit, a missing permission or an untested rollback?

Equation (25.4) turns each condition into a pass or fail, then requires all four. Cost is held to c minus a margin, so a cost of 0.95 fails even though it is under c. The conditions multiply instead of adding, so a zero anywhere gives zero overall, however large the uplift. Authority and rollback are yes or no, not scores that can be traded against performance.

Equation (25.4), written in LaTeX: \operatorname{accept}(\phi')=1 \iff \begin{cases}\widehat\Delta_G(\phi')\geq\tau,\\\widehat C_G(\phi')\leq c-\varepsilon_{\mathrm{safe}},\\\operatorname{Authorized}_{\mathcal G}(\phi')=1,\\\operatorname{Rollback}(\phi',\phi_t)=1.\end{cases}

Scroll sideways for the whole equation

phi' is the candidate version. Delta-hat_G is its guard uplift and tau the required uplift (0.25). C-hat_G is its guard cost estimate, c the current cost limit (1.00) and epsilon_safe the reserved margin (0.10). Authorized is 1 if the version is permitted in the proposed canary scope and 0 otherwise (a canary is a small monitored release). Rollback is 1 if a tested path back to the parent version exists and 0 otherwise.

Predict first. With uplift 0.45 and everything else passing, switch to cost 0.95. That is under the limit c = 1.00. Is the version accepted?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Acceptance needs four passes, not a high score. Left, bars for guard uplift 0.45 against a dashed line at 0.25 and guard cost 0.80 against a dashed limit at 0.90. Right, four pass or fail rows and an overall ACCEPT row.
Guard uplift estimate: 0.45, Cost, authority and rollback: All pass (cost 0.80)
Constructed example: the chapter's threshold 0.25 with cost limit, margin and candidate values defined for this reader; the authority case follows the workbench's revised third candidate.

Calculated values

1. Uplift condition
pass (0.45 vs 0.25)
2. Cost condition
pass (0.80 vs 0.90)
3. Authorized in scope
pass
4. Rollback path tested
pass
Accept
yes (1)

Limit = c - margin = 1.00 - 0.10 = 0.90. Conditions: uplift 0.45 >= 0.25 gives 1; cost 0.80 <= 0.90 gives 1; authorized gives 1; rollback gives 1. accept = 1 x 1 x 1 x 1 = 1. All four conditions hold, so the version may enter the limited canary of Demonstration 1. Acceptance permits that monitored release; it does not show the change is safe at scale.

Worked steps

  1. Cost limit with margin: c - margin = 1.00 - 0.10 = 0.90.
  2. Uplift: 0.45 is at least tau = 0.25, so condition 1 = 1.
  3. Cost: 0.80 is at most 0.90, so condition 2 = 1.
  4. Authority for this scope = 1; tested rollback path = 1.
  5. accept = 1 x 1 x 1 x 1 = 1; the conditions multiply, so one zero decides.

Use the idea

Write a release record with four lines: uplift against its threshold, cost against its limit minus margin, the named authority holder and scope, and the tested rollback route. Name the owner for each, for example the workflow owner for the guard result and the incident owner for rollback alerts. A missing line blocks release.

Where the conclusion applies

Constructed values for one candidate, with uplift and cost already estimated on a clean guard (see Demonstration 3 for what that requires). The laboratory's gate in Demonstration 2 checks only the uplift threshold and the reuse flag, so cost, authority and rollback are computed directly here. A passing record permits a limited canary; it does not show that the change is safe at scale.

Common wrong turn: A high score can pay for a missing permission
The chapter says the four conditions are not four terms of a score to be traded against one another. A procedure that improves a score but lacks permission to call a tool or lacks a recovery route is not accepted.
What this does not settle

An accepted procedure may later generate a new candidate. That does not convert the system into a source of guaranteed recursive gains. A positive result for phi' does not prove that phi'' will improve, and it does not license the system to expand its own change family or authority envelope.

Chapter 25 source: "What this does not settle".

Check your understanding: A candidate has uplift 0.40, cost 0.92, authority granted and rollback tested, with c = 1.00 and margin 0.10. Is it accepted?
The limit is 1.00 - 0.10 = 0.90 and 0.92 > 0.90, so the cost condition is 0 and accept = 1 x 0 x 1 x 1 = 0.

Chapter 25 source: section "Acceptance is a constrained release decision". Demonstration C25-D04.