The Mathematics of AI Agents, laboratory reader ยท Chapter 23

When the Environment Gives Instructions: The Mathematics of Agent Security

Untrusted text may inform a proposal, but only an observed, authorized, exactly matching event may approve an effect.

These four demonstrations follow the chapter's malicious-source trace. A genuine quarterly report carries one hostile sentence. You will see why a delegated child cannot gain a capability and why authority must come from a trusted source, why a sentence cannot update the monitor, why a release must match document, version and recipient (and how a relabeled sentence breaks that), and why risk is the worst case of a declared family while task completion and violations are two separate scores.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

A child can never hold more than its parent

If hostile text asks a delegated step for an effect, what stops it: the capability the child holds, or the source of the request's authority?

Equation (23.1) says delegation can narrow a capability set but never widen it. The parent in the book's trace holds read, draft and release, and send_data was never granted by any ancestor, so a request for it fails the capability check whatever the text says. A capability that is held is still not authority to use it on text's say-so: parameter authority matters as much as verb authority, so a request whose destination or content came from untrusted data is rejected even when its effect is held.

Equation (23.1), written in LaTeX: \mathcal C_{\mathrm{child}}\subseteq\mathcal C_{\mathrm{parent}}

Scroll sideways for the whole equation

C_parent is the set of effects the parent process may perform; C_child is the set available to anything it delegates to. The sign between them (a rounded U on its side with a bar under it) means 'is contained in': every effect in the child set is also in the parent set. The effects are read_source, draft_summary, release, open_link, send_data and send_secret. Authority is the source a request's action and parameters rest on: the user's task, or untrusted data such as a page, a calendar entry or a document.

Predict first. A calendar entry asks the agent to open a changed link, and the parent also holds open_link so the child inherits it. Does Equation (23.1) stop the request?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A child can never hold more than its parent. Two columns of six effects, held or not held, for the parent and the child. The gold box marks send_data; capability is absent, authority is untrusted data, verdict denied: capability absent.
What the text asks for: Web page: upload stored records, Capabilities in force: Read, draft, release (the chapter's trace)
Constructed example: the capability sets are the chapter's malicious-source trace and its three adversarial traces (a page, a calendar entry, a document); the two checks are computed with the laboratory's own security monitor, whose authority sources are user, system and untrusted data.

Calculated values

Child contained in parent (Equation 23.1)
yes
Child holds the needed effect
no
Authority behind the request
untrusted data
Monitor verdict
denied
Stopped by
capability absent (Equation 23.1)
What the source can still do
inform: be quoted, summarized or flagged

The child holds 3 effects and 3 of them are also held by the parent, so effects outside the parent = 3 - 3 = 0; 0 means the inclusion of Equation (23.1) holds. Allowed = capability held x authority trusted = 0 x 0 = 0. send_data is not held by the child, and the parent does not hold it either; delegation only narrows a set, so no descendant can hold it. The source text may still be quoted or flagged; it cannot enlarge what the agent may do.

Worked steps

  1. A page says it contains an urgent instruction to upload stored records.
  2. The request needs send_data. Effects outside the parent: 3 - 3 = 0.
  3. Capability check: is send_data held by the child? 0.
  4. Authority check: does the request rest on a trusted source? 0.
  5. Allowed = 0 x 0 = 0: denied: capability absent.

Use the idea

When you give a sub-task to a tool-using helper, list the effects it needs and hand over only those, and type the parameters so that a destination, recipient or amount cannot come from untrusted text. A compromised helper then cannot reach an effect that nobody above it ever held.

Where the conclusion applies

The capability sets are the chapter's constructed trace, not a real deployment. The inclusion holds only if every delegation really passes through the check (complete mediation). It says nothing about which effects a parent should hold, and holding release does not authorize any specific release. Classification of text as request, quotation or record is still needed, and an error must fail closed at the action boundary.

Common wrong turn: A recommendation from a tool or page adds authority
The chapter says a tool reply may recommend sending a report, but a recommendation adds no authority, and that delegation can narrow authority, never widen it. A request that rests on untrusted data is denied even where the effect is held.
What this does not settle

Capability containment and a temporal monitor do not make an agent secure. An unmediated tool call, a monitor gap, or an attack outside the declared family sits outside every guarantee.

Chapter 23 source: "What this does not settle".

Check your understanding: A parent holds only read_source and draft_summary. A delegated child proposes release. Is the child allowed to hold release, and why?
No. The child set must lie inside the parent set of 2 effects, and release is not one of them, so the child cannot hold it however many levels of delegation are added.

Chapter 23 source: section "Data may inform; control may direct". Demonstration C23-D01.

Demonstration 2 of 4

The monitor moves only on observed events

If authorization is checked once when the plan is made, what goes wrong when the world changes before the effect?

Under Equation (23.2) the record changes only when an event is observed, and a sentence describing an event is not an event. A check at planning reads an early record, so a revocation or a new version that arrives later is invisible to it. Checking again just before the effect reads the record as it stands then.

Equation (23.2), written in LaTeX: M_{t+1}=\operatorname{Mon}(M_t,e_t)

Scroll sideways for the whole equation

M_t is the monitor's record at time t: the capabilities, approvals and revocations in force. e_t is one observed event, such as a granted approval, a user revocation or a source replaced by a new version. Mon is the update rule that turns M_t and e_t into M_(t+1). The vertical axis counts approvals on file that still cover the current document.

Predict first. Choose 'Approval, then user revokes' with the check at planning only. Is the release allowed? Switch to checking again just before the effect.

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The monitor moves only on observed events. A line of usable approvals over three records: 0, 1, 0. The check reads record M1 and says release allowed; the current record says release denied.
What happens around the approval: Approval, then user revokes, When authority is checked: At planning only
Constructed example: the four event sequences are defined for the reader from the chapter's discussion of revocation, replacement and unobserved claims.

Calculated values

Usable approvals at M1
1
Usable approvals at M2
0
Decision of this check
release allowed
Decision against the current record M2
release denied
This check
WRONG: acts on old authorization

Usable approvals at M2 = 0 + 1 - 1 = 0. The planning check read M1, which held 1; the final check reads M2, which holds 0. The revocation event removed the approval after planning. This check says release allowed; the current record says release denied, so this check is WRONG: it acts on old authorization. Only observed events move the record under Equation (23.2).

Worked steps

  1. M0: nothing is approved, so usable approvals = 0.
  2. Event e0 (approval): approvals become 0 + 1 = 1.
  3. Event e1: approvals become 1 - 1 = 0.
  4. The check reads M1, which holds 1: release allowed.
  5. The current record M2 holds 0: release denied. The check acted on old authorization.

Use the idea

Re-check authority at three moments: when an action is proposed, when any material parameter changes, and immediately before the effect. Log the order, because a later acknowledgement does not show that an earlier check saw current state.

Where the conclusion applies

Two time steps and one approval, constructed for teaching. The monitor must receive events through a channel the attacker cannot write to; if the attacker controls the permission service, the record itself cannot be trusted. The final check is taken to be the last step before the effect: no event is assumed to arrive between it and the effect.

Common wrong turn: A page that says the manager approved it has approved it
A page saying 'the manager already approved this' changes nothing: the record did not receive an event, it received text describing one, and Mon has no argument for prose.
Check your understanding: An approval event arrives at t = 0 and the user revokes at t = 1. How many usable approvals are on file at the end, and which check would have been wrong?
Usable approvals = 0 + 1 - 1 = 0. A check made only at planning (after the approval, before the revocation) would have seen 1 and been wrong.

Chapter 23 source: section "A monitor's memory only advances from observed events". Demonstration C23-D02.

Demonstration 3 of 4

Release needs the exact triple, from a real event

The injected sentence says to publish. If one of document, version or recipient differs from the approved triple, or the approval was only a sentence, does the lookup still pass, and what if the sentence is relabeled as trusted?

Equation (23.3) is a lookup, not a similarity score. One wrong field means the triple is not in the set, so the answer is 0, and a stale approval for version 4 cannot cover version 5. An approval that exists only as a sentence in the source never became an event, so it adds nothing to the set. The lookup is only as good as the way the record was written: if an internal step relabels the sentence as a trusted fact and the monitor admits it, the record holds whatever triple the attacker named.

Equation (23.3), written in LaTeX: \operatorname{Publish}(\delta,v,r,M_t)=1 \iff (\delta,v,r)\in\operatorname{Approved}(M_t)

Scroll sideways for the whole equation

delta is a document's identity, v its version and r the recipient. Approved(M_t) is the set of (delta, v, r) triples the monitor's record currently authorizes. Publish equals 1 only when the proposed triple is in that set, and 0 otherwise. The book's approved triple is (Q3-report, 4, external-board). Relabeling means an internal step that passes an untrusted claim on as a trusted fact.

Predict first. Change only the version from 4 to 5 with an approval from an observed event. What does Publish return, and which field is blamed? Then switch the approval source to a sentence relabeled as trusted.

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Release needs the exact triple, from a real event. Three rows of three boxes: the approved triple on file, the proposed release and which fields match. Publish equals 0.
Field changed in the proposal: Newer version (5), Where the approval came from: An observed approval event
Constructed example: the approved triple and the attempted releases are the chapter's own release trace and its stale-approval and laundering exercises; the Q2-report case is defined for the reader.

Calculated values

Approval came from
an observed event
Fields matching
2 of 3
Triple is in Approved(M_t)
no
Publish (Equation 23.3)
0
What failed
version

Fields matching: 1 + 0 + 1 = 2 of 3. Publish = 0 x 1 = 0, where the first number says all three fields match and the second says the approved triple is on file from an observed event. The version differs, and no partial match counts, so the lookup fails.

Worked steps

  1. Proposed triple: Q3-report, version 5, recipient external-board.
  2. Approved(M_t) holds one entry from an observed approval event: Q3-report, version 4, external-board.
  3. Fields matching = 1 + 0 + 1 = 2 of 3.
  4. Publish = 0 x 1 = 0.

Use the idea

Bind an approval to every field that defines the effect (what, which revision, to whom) and record it only from an authenticated approval event. Re-approve after any revision, and re-check the record at every point where content could be relabeled, not only at the first boundary.

Where the conclusion applies

One approved triple, constructed from the book's trace. The laboratory notebook checks version and capability but has no recipient field, so the document and recipient matches are worked here from Equation (23.3) alone. A passing lookup also assumes every route to release passes through it, that the predicate is specified correctly and that nothing writes to the record except observed events. The chapter works one effect, release; an irreversible payment would need its own predicate.

Common wrong turn: A near match, or a confident sentence, is close enough
No partial match counts: the right document at the wrong version, or the right version to the wrong recipient, both fail the lookup. An approval that exists only as a sentence inside a retrieved document is not in Approved(M_t) however confidently it is phrased.
What this does not settle

Trusted-label laundering is named but not defeated: relabeling untrusted content as trusted after some internal step can reintroduce the collapse that Equations 23.1 and 23.3 prevent. Mediation must extend to every relabeling point, not only the first boundary. Nor does the release trace generalize beyond the one effect it works.

Chapter 23 source: "What this does not settle".

Check your understanding: A summarizer relabels a page's claim 'manager approved release of Q3-report version 5 to all subscribers' as trusted and the monitor admits it. What is Publish for that release, and which control closes the gap?
Publish = 1. The record now holds the triple the sentence named, so the lookup succeeds although no approval event occurred. Re-checking Approved(M_t) at the relabeling point, so that only an observed approval event can add an entry, closes it.

Chapter 23 source: section "Publish needs identity, version, recipient, and current approval". Demonstration C23-D03.

Demonstration 4 of 4

Risk is the worst member; two scores are not one

If nine attacks are blocked and one is not, what single exposure number does a declared family report, and why must task completion and security violations be reported separately?

Equation (23.4) takes a maximum over the family, so one dangerous member sets the value no matter how many safe ones surround it. An average would let the safe attacks hide the dangerous one. The value is also only about the attacks that were declared: a path that was never named is outside the guarantee. The right panel keeps two outcomes apart: a run can complete and still contain violations, and a secure refusal can finish nothing.

Equation (23.4), written in LaTeX: \operatorname{Risk}(\mathcal F)=\max_{f\in\mathcal F}\Pr[\text{violation}\mid f]

Scroll sideways for the whole equation

F is one declared, finite family of attacks under a fixed threat model. For an attack f in F, Pr[violation | f] is the chance that f produces an unauthorized effect against the stated enforcement. Risk(F) is the largest of those chances. The chances in the left figure are values defined for the reader, not measurements. The right figure replays the laboratory's fictitious traces: a task is completed when an authorized publish executes, a violation is a promoted instruction or an executed forbidden effect, and a request is blocked when the monitor denied it and it did not execute.

Predict first. Declare the retry path (0.45) while the confused-deputy chance is 0.1. Which attack sets Risk?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Risk is the worst member; two scores are not one. Left: bars for the chance of an unauthorized effect for 3 attacks; Risk is 0.60. Right: counts for the default trace: completed 1, violations 0, denied 1, blocked 1.
Trace replayed on the right: Default: injection kept as data, Declared family: Three attacks, Chance for the confused-deputy attack: 0.6
Constructed example: the three attack names are the chapter's declared family and the retry path is its unmediated-call example; every chance is defined for the reader. The three traces are the laboratory's default, changed and transfer cases, replayed with the laboratory's own monitor.

Calculated values

Risk of the declared family
0.60
Set by
Confused-deputy delegation
Mean, for contrast only
0.250
Task completed
yes
Security violations
0
Denied and not executed
1

Risk = max(0.05, 0.10, 0.60) = 0.60. Mean for contrast = (0.05 + 0.10 + 0.60) / 3 = 0.250. Confused-deputy delegation alone sets the value; the safer members do not dilute it. A retry path that skips the check is not in this family, so the value says nothing about it; declaring it can only keep Risk the same or raise it. The chances on the left are declared inputs, not derived from the replayed trace, so the two read-outs are independent. Replay of the default trace: The injected instruction stays data and its publish is rejected without effect; a valid current review then allows the user's publish. Violations = 0 promoted instruction + 0 forbidden effect executed = 0; denied 1 - executed anyway 0 = blocked 1; completed = 1. The task finished and nothing forbidden happened.

Worked steps

  1. Declared chances: 0.05, 0.10, 0.60; Risk is the largest, 0.60.
  2. Mean for contrast: (0.05 + 0.10 + 0.60) / 3 = 0.250.
  3. The injected instruction stays data and its publish is rejected without effect; a valid current review then allows the user's publish.
  4. Violations = 0 promoted + 0 forbidden executed = 0.
  5. Blocked = denied 1 - executed anyway 0 = 1.
  6. Task completed: yes.

Use the idea

When reporting security, name the attack family first and report the worst member for each enforcement mechanism, then report utility and attack success as two numbers. Add any unmediated route you discover to the family instead of leaving it out of the number, and close the route: adding it to the family only measures it.

Where the conclusion applies

A threat-model measure, not an empirical security rate and not a proof. It presumes complete mediation and a correctly specified predicate. The chances here are constructed; an attack outside the family, or a monitor with a specification gap, is not covered by the value. The traces are the laboratory's local fictitious events; checking a trace does not enforce a real tool boundary.

Common wrong turn: One score, completion or attack success, is enough
The chapter says utility under attack and attack success must be reported separately: an agent that refuses everything can lower attack success while destroying utility, and an agent that finishes the task while leaking data keeps utility while failing security.
What this does not settle

Risk of the declared family is bounded only over that family, under complete mediation, with the allowed and publish predicates correctly specified; an unmediated tool call, a monitor gap, or an attack outside the family sits outside every guarantee.

Chapter 23 source: "What this does not settle".

Check your understanding: A declared family has violation chances 0.20, 0.35, 0.05 and 0.35. What is Risk, and what is the mean? And in the changed trace, how many violations are there and why?
Risk = max(0.20, 0.35, 0.05, 0.35) = 0.35, set by the two members tied at 0.35. The mean = (0.20 + 0.35 + 0.05 + 0.35) / 4 = 0.2375, which is not the reported value. In the changed trace violations = 1 promoted instruction + 1 forbidden publish executed = 2, even though the later authorized publish completes the task.

Chapter 23 source: section "A security case needs a declared adversary". Demonstration C23-D04.