The Mathematics of AI Agents, laboratory reader ยท Chapter 18

When an Action Has Coordinates

A click names a position, not an object. What it does depends on the screen that is current when it arrives.

These four demonstrations follow the chapter's release console, where the rows can swap places after the controller has looked. They show how the interface enters the transition law for a coordinate, a semantic and a transactional command, which commands survive when the screen is uncertain, how age, version and permission combine with a steady rate of change, and when another look is worth its cost.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

One command, two layouts, three interfaces

If the controller is unsure which layout is current, what does each kind of command do, and what does a permission check change?

In layout 1 the saved click realizes release v2 for certain; in layout 2 the coordinate realizes release v3 for certain, or a denial if the permission check is on. Equation (18.1) multiplies each realized operation by its world effect and adds, and the belief averages over the two layouts. Uncertainty sits in which layout is current, not in the click. The semantic request is bound to the object, so the swap does not move it, and the transactional request adds the saved version, so a changed layout produces a refusal instead of an effect on a new record.

Equation (18.1), written in LaTeX: P_{\mathrm{ui}}(x'\mid x,a)=\sum_{\tilde a\in\operatorname{Act}_{\mathrm{ui}}}\operatorname{Exec}_{\mathrm{ui}}(\tilde a\mid x,a)P(x'\mid x,\tilde a)

Scroll sideways for the whole equation

x is the current state (here, which layout is on screen), a is the proposed command, and a-tilde is the operation the interface actually realizes: release v2, release v3 or a denial. Exec(a-tilde | x, a) is the chance the interface realizes a-tilde, and P(x' | x, a-tilde) is the world's law for the next state x'. Act with subscript ui is the set of operations the interface can realize, and Exec with subscript ui is the interface's own chance of realizing one. The belief is the controller's probability that layout 1 is current. A coordinate command is a position; a semantic command names a control bound to v2; a transactional command names v2 and the saved screen version, which the service checks at commit.

Predict first. With belief 0.6 in layout 1, no permission check and the coordinate command, what is the chance of a prohibited release?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: One command, two layouts, three interfaces. Left: two screens, layout 1 with v2 on top and layout 2 with the rows swapped, and a star where the coordinate command lands in each. Right: predicted probabilities of releasing v2 (0.80), releasing v3 (0.20) and no release (0.00).
Kind of command: Coordinate: click the saved point, Belief that layout 1 (v2 on top) is current: 0.8, Current permission check at the service: Off
Constructed example: the chapter's two-layout console with its 0.8 and 0.2 beliefs and workbench exercise 1's 0.6 and 0.4; the three interfaces are the chapter's, with issuance and target correctness computed by the laboratory's interface calculation.

Calculated values

Release v2 (authorized completion)
0.80
Release v3 (prohibited release)
0.20
Denied or refused, no release
0.00
Total
1.00

Layout 1 (v2 on top) has belief 0.80; layout 2 (rows swapped) has 0.20. The coordinate command takes its target from whatever occupies the saved point when it arrives. P(release v2) = 0.80 x 1 + 0.20 x 0 = 0.80. P(release v3) = 0.80 x 0 + 0.20 x 1 = 0.20. P(no release) = 0.80 x 0 + 0.20 x 0 = 0.00. The three add to 1.00, as a normalized transition law must. With no check, the wrong-layout branch becomes a prohibited release with probability equal to the belief in that layout. Later detection could identify it but cannot make the recipient unsee the document.

Worked steps

  1. Layout 1: the coordinate command realizes release v2; layout 2: release v3.
  2. P(release v2) = 0.80 x 1 + 0.20 x 0 = 0.80.
  3. P(release v3) = 0.80 x 0 + 0.20 x 1 = 0.20.
  4. P(no release) = 0.80 x 0 + 0.20 x 0 = 0.00.
  5. Check: 0.80 + 0.20 + 0.00 = 1.00.

Use the idea

Before trusting a command, write down what it realizes under every layout you consider possible, and which of those realizations a current check would block. Compare interfaces with the task, the document, the authority and the deadline held fixed.

Where the conclusion applies

Two layouts, each with a deterministic realized operation that leads to one next state, and complete mediation when the check is on. The beliefs are probabilities in a constructed model, not measured accuracies of any visual model. The semantic request is taken to resolve v2 correctly; a mislabeled or ambiguous control would break that, and a transactional request can fail because its precondition expired even when a person would accept the newer version. If an unmediated route to the service exists, the check does not apply.

Common wrong turn: A better visual model fixes the wrong target
The chapter notes that the improvement from a permission check comes from enforcement: the controller has not become better at locating the button. A better visual model might reduce misreading, but it cannot establish that a previously observed interface has remained unchanged. Turn the check on: the belief is the same, the prohibited release disappears.
Check your understanding: With belief 0.7 in layout 1, what is the chance of a prohibited release for the coordinate command with no check, and with the check on?
No check: 0.7 x 0 + 0.3 x 1 = 0.30. Check on: 0.7 x 0 + 0.3 x 0 = 0. Authorized completion stays 0.7 x 1 + 0.3 x 0 = 0.70 in both.

Chapter 18 source: section "The interface belongs in the transition law". Demonstration C18-D01.

Demonstration 2 of 4

What survives when the screen is uncertain

Which commands stay authorized in every layout the controller still considers possible, which state removes the rest, and which observation restores them?

A command is kept only if every state with positive weight authorizes it. Layout 2 does not authorize the saved-point click, so any positive weight on layout 2 removes it, while zero weight leaves layout 2 out of the intersection entirely. The weight itself never enters the result, only whether it is positive. The last column names the state that removed each command, and a perfect read narrows the support: seeing layout 1 restores the click, seeing layout 2 confirms its removal.

Equation (18.2), written in LaTeX: \mathcal A_{\mathrm{rob}}(\mathbf b)=\bigcap_{x: \mathbf b(x)>0}\mathcal A_{\mathrm{auth}}(x)

Scroll sideways for the whole equation

b(x) is the belief weight of state x, here layout 1 or layout 2. The set written A with subscript auth, at x, holds the commands whose protected effect is authorized in state x. The set written A with subscript rob, at b, is the intersection of those sets over every state with positive belief. A perfect read is an observation that shows which layout is current.

Predict first. Layout 2 has belief only 0.01 and nothing further is observed. Does the saved-point click survive the intersection?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: What survives when the screen is uncertain. A grid of five commands against layout 1, layout 2, the intersection and the state that removed each command. 4 of 5 commands survive; the saved-point click is removed.
Belief that layout 2 (rows swapped) is current: 0.01, Further observation: No further observation
Constructed example: the chapter's two-layout case, with authorization sets defined for this reader and belief weights chosen for illustration.

Calculated values

Belief in layout 1
0.99
Belief in layout 2
0.01
Layouts counted
layout 1 and layout 2
Commands kept
4 of 5
Saved-point click
removed

Beliefs add to 0.99 + 0.01 = 1.00. Both beliefs are above zero, so both layouts count, whatever their size. The click reaches an unapproved record in layout 2, so it is not authorized throughout the support. Commands removed = 5 - 4 = 1, so the saved-point click is removed and 4 of 5 commands remain. The request for v2 survives in either layout only because the service checks identity and approval at commit; belief alone never authorizes anything.

Worked steps

  1. Beliefs: layout 1 = 0.99, layout 2 = 0.01; they add to 1.00.
  2. States with positive belief: layout 1 and layout 2.
  3. Each command must be authorized in every counted layout; the saved-point click is not authorized in layout 2.
  4. Commands removed = 5 - 4 = 1; the saved-point click is removed.

Use the idea

When a decision is removed by a state you think unlikely, the record shows which state removed it and which observation (a fresh screenshot, a metadata read) could narrow the support and restore it. Reading can often be permitted over a broader set of states than releasing, which is why it stays available.

Where the conclusion applies

The true layout must lie inside the counted support, and the authorization sets must be specified correctly; if either fails, the intersection guarantees nothing. The sets here are constructed, and the construction is conservative: it is not a substitute for the service's own current permission check. A read only helps if it is current and correct.

Common wrong turn: A likely target is authority for it
The chapter says a high-probability intended target is not current authority for a different target. At belief 0.99 in layout 1 the click is probably right, yet the 0.01 on layout 2 removes it, because the intersection asks what is authorized throughout the support, not what is most likely.
Check your understanding: If layout 2 has belief 0.001 and nothing further is observed, does the saved-point click survive? What if the belief is exactly 0?
At 0.001 the weight is positive, so layout 2 counts and removes the click (5 - 4 = 1 removed, 4 of 5 remain). At 0 layout 2 is not counted, the click survives and 5 of 5 remain, but only if layout 1 is truly current.

Chapter 18 source: section "A belief over screens". Demonstration C18-D02.

Demonstration 3 of 4

How long a picture stays useful

Given its age, its saved version and the current permission, which interface may act, and how fast does a coordinate go stale if changes arrive at a steady rate?

If changes arrive at a constant rate, the chance of none in a delay falls as an exponential of rate times delay. Doubling the delay squares the freshness: at 0.02 per second, 15 seconds gives 0.7408 and 30 seconds gives 0.5488, which is 0.7408 squared. The table shows what the age rule, the version and the permission do on top of that: in the default case the age is within its limit but the layout changed, so only the version-bound request refuses; in the changed case the age passes its limit and all three refuse; in the transfer case permission is false and every interface refuses although the observation is fresh and the versions match.

Equation (18.3), written in LaTeX: \operatorname{Fresh}(\Delta)=\exp(-\lambda_{\mathrm{ui}}\Delta)

Scroll sideways for the whole equation

Delta is the delay between observing and acting, in seconds. Lambda with subscript ui is the rate of invalidating interface changes, in changes per second. Their product is a pure number. Fresh(Delta) is the chance that no invalidating change occurred during the delay. Separately, the age rule compares the observation's age with a limit; the version-bound interface also compares the saved version with the current one; every interface needs current permission.

Predict first. At 0.02 changes per second, is the chance of still being fresh after 30 seconds above or below one half?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How long a picture stays useful. Left: a table of coordinate, semantic and version-bound interfaces against fresh, version, allowed, issued, target and done; interfaces that issue: coordinate, semantic. Right: freshness falls from 1 to 0.55 at 30 seconds at 0.02 changes per second; the table does not depend on the delay.
Case: Default: age 1 of limit 2, layout changed, Delay before acting (seconds): 30
Constructed example: the notebook's default case (age 1 against a limit of 2, layout-1 against layout-2, rate 0.02), changed case (age 3) and transfer case (permission false, rate 0.05), computed with the laboratory's interface function, and the chapter's delays of 5 and 30 seconds with workbench exercise 2's 15; the delay of 10 seconds is defined for this reader.

Calculated values

Rate
0.02 per second
Delay
30 seconds
Rate x delay
0.60
Chance still fresh
0.5488
Chance of an invalidating change
0.4512
Delay at which freshness is one half
34.7 seconds
Observation age against the limit
1 s against 2 s: fresh
Interfaces that issue
coordinate, semantic
Confirmed completions
semantic

Rate x delay = 0.02 x 30 = 0.60, a pure number with no unit left. Fresh = exp(-0.60) = 0.5488, so the chance of at least one invalidating change is 1 - 0.5488 = 0.4512. Freshness reaches one half at 0.6931 / 0.02 = 34.7 seconds. The table uses the case's observation age (1 s) and the curve uses the delay you choose (30 s): they are separate inputs, so changing the delay moves the curve and not the table. The age rule passes (1 <= 2) yet the saved layout-1 differs from the current layout-2, so only the version-bound request refuses. The coordinate request is issued and reaches delete, not release, so it does not complete; the semantic request reaches release and completes.

Worked steps

  1. Rate x delay = 0.02 x 30 = 0.60.
  2. Fresh = exp(-0.60) = 0.5488; an invalidating change has chance 1 - 0.5488 = 0.4512.
  3. Age rule: observation age 1 s against a limit of 2 s is fresh.
  4. Version: saved layout-1, current layout-2; only the version-bound request compares them.
  5. Current permission is true; issued = fresh and allowed (and, for version-bound, same version).
  6. Issued: coordinate, semantic. Confirmed completion: semantic.

Use the idea

Choose a maximum observation age by deciding what chance of an invalidating change you will accept, then refresh the observation or bind the command to a state version before that age is reached, and check current permission at issuance, not from memory.

Where the conclusion applies

The table and the curve are independent: the table uses the case's observation age, the curve uses the chosen delay. In the default and changed cases the coordinate target is called delete, the notebook's name for the wrong object (v3 in Demonstration 1). Constant-rate, memoryless changes, and a clear definition of which changes invalidate the binding. Real interfaces can update periodically, in bursts, or in response to the agent's own actions, and a rate taken from quiet periods can mislead. The rates and limits here are teaching inputs. Semantic resolution can itself be wrong, and a click receipt is not an effect receipt.

Common wrong turn: A freshly taken screenshot is current
The chapter notes that a fresh observation has latency: if observation, interpretation and execution take several seconds, a controller can keep acquiring a new image and still act on an old state. Age by itself does not decide validity; a state version, a lock, a commit precondition or a service action that checks the object may be needed.
Check your understanding: At 0.05 invalidating changes per second and a 20 second delay, what is Fresh?
Rate x delay = 0.05 x 20 = 1.0, so Fresh = exp(-1) = 0.3679. The chance of an invalidating change is 1 - 0.3679 = 0.6321.

Chapter 18 source: section "How long a picture stays useful". Demonstration C18-D03.

Demonstration 4 of 4

The price of another look

When does one more observation before the click earn its cost?

Clicking now costs the chance of a wrong target times its loss. Looking first costs its price and removes that loss. The look pays when chance x loss exceeds the price, which gives a break-even chance of the price divided by the loss. Enforcement matters because it changes the loss, not the chance; binding the command to v2 with a version check removes the chance for document identity altogether, leaving the look nothing to buy on that question.

Equation (18.1), written in LaTeX: P_{\mathrm{ui}}(x'\mid x,a)=\sum_{\tilde a\in\operatorname{Act}_{\mathrm{ui}}}\operatorname{Exec}_{\mathrm{ui}}(\tilde a\mid x,a)P(x'\mid x,\tilde a)

Scroll sideways for the whole equation

The chance the saved point is wrong is the probability of the layout in which the click reaches the wrong record. The loss is what that outcome costs, in declared utility units. A second observation costs 1 or 2 units and is assumed to settle which layout is current. Equation (18.1) describes the possible outcomes of the click (its symbols are defined in Demonstration 1: a-tilde is a realized operation, Exec with subscript ui its chance, Act with subscript ui the set of realizable operations, x' the next state). This demonstration's own comparison is: expected loss of clicking now = chance x loss, and looking pays when chance x loss is more than its cost.

Predict first. With a wrong-target loss of 3 (a denial) instead of 40, a look costing 2 and a 0.1 chance of being wrong, does the extra look still pay?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: The price of another look. Left: bars for the expected loss of clicking now (4.00) and the cost of looking (2.00). Right: the expected loss rising with the chance of a wrong target against the flat cost of looking, crossing at 0.050; the chosen chance is 0.10.
Loss if the click reaches the wrong record: 40: unguarded harmful release, Chance the saved point is wrong: 0.1, Cost of one more look: 2
Constructed example: the chapter's 0.1 chance of a wrong target, loss of 40 or 3 and observation cost of 2, workbench exercise 3's look costing 1 with a 0.2 chance, and the chapter's third interface with chance 0.

Calculated values

Expected loss of clicking now
4.00
Cost of one more look
2.00
Net advantage of looking
2.00
Break-even chance of a wrong target
0.050
Better option
Observe first

Expected loss of clicking now = 0.10 x 40 = 4.00. Looking first costs 2.0 and removes that loss, so its net advantage is 4.00 - 2.0 = 2.00. Break-even chance = 2.0 / 40 = 0.050. Observing first removes an expected loss of 4.00 for a cost of 2.0, so it pays.

Worked steps

  1. Expected loss of clicking now: 0.10 x 40 = 4.00.
  2. A perfect look costs 2.0 and removes that loss.
  3. Net advantage of looking: 4.00 - 2.0 = 2.00.
  4. Break-even chance: 2.0 / 40 = 0.050; the chosen chance is 0.10.
  5. Better option: observe first.

Use the idea

Do not buy information because uncertainty exists. Price the wrong-target branch under the interface you actually have, and look again only when that expected loss is larger than the cost of looking.

Where the conclusion applies

The look is perfect, execution follows at once, and the loss lumps delay, retries and side effects into one number. A denied attempt can also use up a rate limit, lock an account or reveal information; if so the loss of 3 is too small. A deadline can raise the cost of looking. The number 2 is not a general price for visual safety.

Common wrong turn: Buy information whenever you are uncertain
The chapter says a controller should not buy information merely because uncertainty exists, but when the information can change an authorized decision enough to justify its cost. With complete permission mediation a wrong click only costs a denial, and the same perfect look no longer pays.
What this does not settle

Physical agents extend the problem further: contact, inertia, uncertain actuators and irreversible material effects can require additional state and control models, and this chapter establishes no transfer guarantee from desktop completion to robotics. For the release controller, the useful result is narrower: the model may choose the right intention, and the interface must still connect that intention to the right object, current permission and confirmed consequence.

Chapter 18 source: "What this does not settle".

Check your understanding: If one more look cost 3 units, the wrong-target loss were 40 and the chance of a wrong target 0.1, would looking first pay?
Clicking now costs 0.1 x 40 = 4.0. Looking costs 3, so the net advantage is 4.0 - 3 = 1.0 and it pays. The break-even chance is 3 / 40 = 0.075, below 0.1.

Chapter 18 source: section "The price of another look". Demonstration C18-D04.