The Mathematics of AI Agents, laboratory reader · Chapter 10

Plans Within Plans

A skill that lasts several steps needs a start rule, an inner policy and a stopping rule, and its value must count how long it really ran.

These four demonstrations follow the chapter's document-release workflow. First the three parts of an option, then how its value changes with duration, then how time and cost add up differently (and what a stale completion adds), and last what an interruption does to completion. Every number is a constructed teaching value.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

An option has three parts

Where may an option start, how long does it run once it has started, and what does a blocked start earn?

Equation (10.1) lists three separate things. The start rule only decides which rows are blocked. The inner policy fixes what each step does. The stopping rule decides where each run ends and how many steps it never takes. Changing the stopping rule changes durations without changing where the option may start, and the reverse.

Equation (10.1), written in LaTeX: o=(\mathcal{I}_o,\pi_o,\beta_o)

Scroll sideways for the whole equation

o is the option. ℐ (script I) is its initiation set: the states where it may start. The policy π (pi) says which primitive action to take at each step while it runs; here it checks one record per step. β (beta) is the termination rule that says when control returns. The state x is the number of records still to check; the duration is counted in primitive steps. In written sums, x between numbers means multiply.

Predict first. Require at least 4 records and add an early stop when 2 records remain. How long does the option run from x = 6, and can it start from x = 3?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: An option has three parts. Eight rows, one per start state from 1 to 8 records left. Allowed rows show one filled cell per step taken; the longest run from an allowed start is 8 steps. No early stop, so every allowed row runs one step per record.
Start rule: at least this many records left: 1, Early stop on a missing approval: No early stop
Constructed example: a record-checking option defined for this reader to illustrate the three parts in Equation (10.1).

Calculated values

Initiation set
x = 1 to 8
Start states allowed
8 of 8
Steps run from x = 6
6 steps
Steps run from x = 8
8 steps
Longest run from an allowed start
8 steps

From x = 6 it checks one record per step, so it runs 6 x 1 = 6 steps. From x = 8 it runs 8 x 1 = 8 steps. The start rule x of at least 1 allows every state shown. Without an early stop, the duration always equals the number of records. The termination rule, not only the start state, fixes the duration.

Worked steps

  1. Initiation set: states x of at least 1, so x = 1 to 8.
  2. Inner policy: check one record per primitive step.
  3. Termination: stop only when no record is left.
  4. From x = 6 it checks one record per step, so it runs 6 x 1 = 6 steps.
  5. From x = 8 it runs 8 x 1 = 8 steps.

Use the idea

Before calling a workflow a reusable skill, write down all three: where it may begin, what it does at each step, and what event hands control back. A workflow with a name but no stated stopping rule has not yet been defined as an option.

Where the conclusion applies

A row of 8 start states, one record checked per step, and a missing approval that is met at a fixed number of records remaining. A start state at or below the stop level runs to the end, so durations need not rise steadily with x. Real options act on richer state, and their stopping may depend on what they see. The start rules and stop levels here are declared examples, not rules from the chapter.

Common wrong turn: A named workflow is already an option
The chapter says "Retrieve, then decide" is not an option merely because it has a name in an interface. One must state where it may begin, what primitive choices it makes in response to state, and what event returns control. Change the start rule or the stop here and watch the rows: each of the three parts changes a different thing.
Check your understanding: With no early stop and a start rule of at least 4, how many start states are allowed, and how long does the option run from x = 8?
The allowed states are 4, 5, 6, 7 and 8, so 8 - 4 + 1 = 5 states. From x = 8 it checks one record per step, so it runs 8 x 1 = 8 steps.

Chapter 10 source: section "A skill is not a label". Demonstration C10-D01.

Demonstration 2 of 4

Discount by the steps the option really took

How much does the value rise when the continuation after an option is discounted as if it took one step?

Equation (10.2) discounts each inner reward by how many steps passed and then discounts the value that remains by the whole duration. The shortcut discounts it by a single step. The longer the option runs, the more the shortcut overstates the continuation. With these inputs the overstatement is V x (gamma - gamma^tau): a smaller gamma gives a larger gap at durations 2 to 4, and the gap grows with duration for either discount. The right panel shows why: the weight at the step where the option ends falls with every extra step.

Equation (10.2), written in LaTeX: Q^{\mu}(x,o)=\operatorname E_\mu[\sum_{k=0}^{\tau-1}\gamma^k r_{t+k}+\gamma^\tau V^\mu(x_{t+\tau}) \middle| x_t=x,o_t=o]

Scroll sideways for the whole equation

Q(x, o) is the value of starting option o in state x under the parent policy μ (mu). τ (tau) is the number of primitive steps the option lasts. r is the reward earned at a step inside the option, γ (gamma) the discount per primitive step, and V is the value of the state where the option ends. In the chapter example the rewards are 2 then 3, and V is 10; in the workbook transfer case they are 0 then 4, and V is 8. Rewards and the end state are deterministic here, so the expectation is a single value. In written sums, x between numbers means multiply.

Predict first. With discount 0.9 and a duration of 2, the chapter says 12.8 against 13.7. Make the option last 4 steps. Does the overstatement of the one-step shortcut grow or shrink?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Discount by the steps the option really took. Left: five bars. Rewards inside the option 4.70, continuation 8.10, value 12.80, and hatched bars for the one-step shortcut, continuation 9.00 and value 13.70. Right: discount weights for steps 0 to 4, with the weight at step 2 marked as where the option ends.
Case: Chapter example: rewards 2 and 3, value 10, Duration of the option (primitive steps): 2, Discount per primitive step: 0.9
Constructed example: the chapter's worked numbers (rewards 2 and 3, continuation 10, discount 0.9, also its Exercise 1) and the workbook transfer case (rewards 0 and 4, continuation 8, discount 0.5), with the duration and discount varied, computed with the laboratory's option function.

Calculated values

Rewards inside the option
4.700
Continuation discounted by gamma^2
8.100
Value by Equation (10.2)
12.800
Value if discounted as one step
13.700
Overstatement
0.900

Value = 1 x 2 + 0.9 x 3 + 0.81 x 10 = 12.800. Treating the option as one step instead gives 4.700 + 0.9 x 10 = 13.700, which is 0.900 too high. The exponent counts the steps the option really used, so a longer option has its continuation worth less. Ignoring duration makes delay look too cheap.

Worked steps

  1. Rewards inside: 1 x 2 + 0.9 x 3 = 4.700.
  2. The option lasts 2 steps, so the continuation is discounted by gamma^2 = 0.81.
  3. Continuation: 0.81 x 10 = 8.100.
  4. Value = 4.700 + 8.100 = 12.800.
  5. One-step shortcut: 4.700 + 0.9 x 10 = 13.700.
  6. Overstatement = 13.700 - 12.800 = 0.900.

Use the idea

When comparing a one-step action with a multi-step skill, put both on the same clock. A skill that takes five steps to return should not get the continuation value of one that returns immediately.

Where the conclusion applies

The duration is known and the option always reaches its end. Rewards are paid at primitive steps and the first is undiscounted. The discount is per step, not per minute: wall-clock time needs its own model. Steps after the second pay 0 here, a constructed choice to make the duration vary.

Common wrong turn: Discount the continuation as if the option took one step
The chapter calls that calculation inconsistent: keeping both inner rewards but discounting the continuation by one step returns 13.7 instead of 12.8, and it makes delay look too cheap. The exponent must be the number of primitive steps the option really used, and the clock must be named: wall-clock time would need its own model.
What this does not settle

Equation (10.2)'s value presumes the option's initiation, policy, and termination are already correctly declared, not discovered by the option itself. The numbers are stipulated for teaching, not measured from a deployed system.

Chapter 10 source: "What this does not settle".

Check your understanding: With discount 0.9, rewards 2 and 3, and continuation 10, what is the value if the option takes 3 steps and the third step pays 0?
2 + 0.9 x 3 + 0.81 x 0 + 0.729 x 10 = 2 + 2.7 + 0 + 7.29 = 11.99.

Chapter 10 source: section "Value when duration varies". Demonstration C10-D02.

Demonstration 3 of 4

Release time is not release cost

If one activity runs beside the others, how can the release be early while the cost stays the same, and what does a stale verification add?

Time follows the longest dependency chain: a maximum. Cost follows everything that must be paid for: a sum. The two totals therefore respond to different things. Shortening an activity off the longest chain buys no time, and running activities in parallel saves time without saving a single credit. A stale completion adds required work to the chain, so it raises both totals.

Equation (10.3), written in LaTeX: T_{\mathrm{release}}=\max\{2+3+1, 4\}=6\text{ minutes}, C_{\mathrm{tools}}=4+7+2+3=16\text{ credits}

Scroll sideways for the whole equation

T is the time until release, in minutes, and C the total tool cost, in credits. Draft (2 minutes, 4 credits), Verify (3, 7) and Publish (1, 2) must run one after another. Audit (4 minutes by default, 3 credits) is independent. Release needs both Publish and Audit. The longest chain of required activities sets T; every required activity adds to C. If the source changes from v17 to v18 while Verify runs, Verify's completion is about v17, so a Revalidate activity against v18 is added; its 3 minutes and 7 credits are a value set for this reader (it repeats Verify).

Predict first. Make the audit take 8 minutes, in parallel, with the source unchanged. Which chain sets the release time, and does the cost change?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Release time is not release cost. A bar chart of minutes for Draft, Verify, Publish, Audit. The release line is at minute 6 and the total cost is 16 credits. The audit bar starts at minute 0.
Audit duration (minutes): 4, Audit schedule: In parallel from minute 0, Source during Verify: Unchanged
Constructed example: the chapter's release workflow (Equation 10.3, 6 minutes and 16 credits) with the audit length and the schedule varied; the stale completion follows the chapter's v17 and v18 account, and its revalidation cost is defined for this reader.

Calculated values

Release time
6 minutes
Total tool cost
16 credits
Longest chain
Draft, Verify, Publish
Chain Draft, Verify, Publish
6 minutes

Time T = max{2 + 3 + 1, 4} = max{6, 4} = 6 minutes. Cost C = 4 + 7 + 2 + 3 = 16 credits in every schedule, because every required activity is paid for. The audit has 2 minutes of slack, so speeding it up would not release earlier.

Worked steps

  1. Chain minutes: 2 + 3 + 1 = 6.
  2. Audit minutes: 4, starting at minute 0.
  3. Time T = max{6, 4} = 6 minutes.
  4. Cost C = 4 + 7 + 2 + 3 = 16 credits.

Use the idea

When a controller says a workflow is cheap or fast, ask which total is meant. A plan can be fast by running work side by side; running it in sequence is slower but, in this example, costs the same.

Where the conclusion applies

Fixed durations and costs with no waiting for tools, no failures and no shared resource between the parallel activities. The sequential schedule and the revalidation (3 minutes, 7 credits) are variations added here, not in the chapter. The chapter's own case is a 4 minute audit in parallel with an unchanged source. Real durations vary and credits may depend on the order of work.

Common wrong turn: One step count, or one total, for the whole workflow
The chapter warns that agent systems often report a single step count although four different counts are mixed: token emissions, tool invocations, test runs and delegated tasks, and none determines another. Here two clocks already differ: minutes follow a maximum and credits follow a sum. A controller must state which resource its value measures before it combines them.
What this does not settle

Shared-state conflicts between options are named, not solved: a lock, snapshot, or retry rule is a separate transition-law decision this chapter does not make.

Chapter 10 source: "What this does not settle".

Check your understanding: In the parallel schedule with an unchanged source, how long must the audit take before it, rather than the chain Draft, Verify, Publish, sets the release time?
The chain takes 2 + 3 + 1 = 6 minutes, so the audit must take more than 6 minutes. At exactly 6 the two chains tie. The cost stays 4 + 7 + 2 + 3 = 16 credits.

Chapter 10 source: section "A release workflow has two totals". Demonstration C10-D03.

Demonstration 4 of 4

An interrupted option is scored here without continuation

If an option is stopped early, how much of its value survives, can the preferred option change, and what does an option that cannot start earn?

Equation (10.2) adds γ^τ V for the state where the option ends. In this demonstration the laboratory stipulates V = 0 for an interrupted option, so only the rewards of the steps that ran are counted; Equation (10.2) itself would add γ^τ V at the state actually reached, which is not computed here. The interruption does not change the payoff of releasing; it removes access to the later reward. An option whose start rule fails (the disabled option) is the extreme case: it runs no steps at all.

Equation (10.2), written in LaTeX: Q^{\mu}(x,o)=\operatorname E_\mu[\sum_{k=0}^{\tau-1}\gamma^k r_{t+k}+\gamma^\tau V^\mu(x_{t+\tau}) \middle| x_t=x,o_t=o]

Equation (10.1), written in LaTeX: o=(\mathcal{I}_o,\pi_o,\beta_o)

Scroll sideways for the whole equation

review-release is a three-step option with rewards -1, -1 and 8 and a continuation value of 2. archive is a one-step option that pays 1. In the workbook case, inspect has rewards 0 and 4 and a continuation value of 8, and disabled is a one-step option that would pay 9 but whose start rule fails. γ (gamma) is the discount per primitive step. The deadline is 5 steps (2 in the workbook case) and does not bind. An interruption stops every option after the chosen number of steps. The laboratory's rule, used in this demonstration, gives an option that did not reach its end no continuation, which is the same as setting the value of the state it reached to 0.

Predict first. At discount 0.9, interrupt review-release after 2 steps. What is its value, and which option is now higher?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: An interrupted option is scored here without continuation. Two bars. Review-release has inside rewards 4.58 plus a hatched continuation 1.46, total 6.04. Archive is a single bar of 1.
Case: No interruption, Discount per primitive step: 0.9
Constructed example: the laboratory's review-release and archive options, using the laboratory's example values (6.038 complete, -1.9 interrupted after 2 steps, at discount 0.9), a second discount, and the workbook transfer case (inspect worth 4 at discount 0.5, a disabled option worth 0), computed with the laboratory's option function.

Calculated values

review-release value
6.038
archive value
1.000
review-release completed
yes
Higher executed value
review-release

Review-release = 1 x (-1) + 0.9 x (-1) + 0.81 x 8 + 0.729 x 2 = 6.038. The option reached its end, so the continuation value 2 is added after discounting by 0.9 x 0.9 x 0.9 = 0.729. Review-release is higher, 6.038 against 1.000 for archive.

Worked steps

  1. Inside rewards: 1 x (-1) + 0.9 x (-1) + 0.81 x 8 = 4.580.
  2. Continuation: 0.729 x 2 = 1.458.
  3. Review-release = 4.580 + 1.458 = 6.038.
  4. Archive = 1 x 1 = 1.000.
  5. Review-release is higher, 6.038 against 1.000 for archive.

Use the idea

A parent controller should treat an interrupted skill as a different outcome from a finished one and plan from the state it actually reached, with a stated cancellation or recovery step.

Where the conclusion applies

Rewards, durations and the interruption point are given, not random. The chapter warns that an interruption can leave effects behind (spend, writes) and that cancellation needs an acknowledgement. This calculation shows only the value of the executed steps (the reached state is valued at 0 by stipulation) and says nothing about how to clean up.

Common wrong turn: A timeout or a cancel request means the option stopped
The chapter says timeout alone neither acknowledges cancellation nor fences off outstanding effects, and that "cancel requested" is not proof of non-application. Without an acknowledgement the parent loses the right to assume the option ended; the stopped option here is only scored as stopped because the interruption point is given.
What this does not settle

This chapter does not prove that a compact completion interface loses no distinction a later decision needs; that check belongs to the interface designer, not to the option contract itself.

Chapter 10 source: "What this does not settle".

Check your understanding: At discount 0.9, what is the value of review-release if it is stopped after 1 step?
Only the first reward is earned and nothing is discounted yet: 1 x (-1) = -1. This demonstration adds no continuation, so archive (1) is higher.

Chapter 10 source: section "Interruption changes what completion means". Demonstration C10-D04.