The Mathematics of AI Agents, laboratory reader ยท Chapter 15

What an Agent Should Remember: Retrieval, Compression, and Experience

A memory earns its place by improving a later decision, and a similar old sentence is not current authority.

These four demonstrations follow the chapter's release controller. It has a small budget for what it may bring back into view: a current revocation, an old approval, a summary and an irrelevant note. They show which records a budget buys once expired and superseded records are screened out, what a summary can lose and when the bound stops applying, why wording similarity and incomplete deletion cannot stand in for authority, and when a record that may no longer be current is worth its tokens. Every number is a constructed teaching value.

Every example in these readers is a constructed teaching example. The probabilities, utilities and cases are declared inputs chosen to make the mathematics visible. They are not measurements of any deployed agent, product or team.

Demonstration 1 of 4

Records chosen under a token budget

Which records does a limited budget buy once expired and superseded records are screened out, and how does the price of a token change the answer?

Each eligible record gets a net value: what it adds to the decision minus what its tokens cost. Equation (15.1) looks for the set of records, no larger than the budget, with the greatest total net value. A record with negative net value is left out even when room remains, and a useful record can be left out because the budget is spent. Records that are expired or tied to an old version are excluded before any value is counted. Similarity to the request is not part of the calculation.

Equation (15.1), written in LaTeX: R^\star\in\arg\max_{R: \sum_{m\in R}c(m)\le B}\{\operatorname E[\text{decision value}\mid R]-\lambda\sum_{m\in R}c(m)\}

Scroll sideways for the whole equation

R is the set of records brought into view, m one record, c(m) its size in tokens, and B the token budget. The expected decision value is how much R improves the present decision, in declared units. ฮป (lambda) is the price of one token in those same units. Net value of a record is its decision value minus ฮป times its tokens. The arg max picks the set with the largest total. Before any value is counted, a record must be fresh (its age is at most its time to live) and, if it carries authority, bound to the current version.

Predict first. In the default case with ฮป = 0 (budget 6), which records are retrieved?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Records chosen under a token budget. Chapter example. Horizontal bars of net value for 4 records at budget 2 and lambda 1; chosen: Current revocation; excluded before scoring: Old approval.
Case: Chapter example: budget 2, Price of a token, lambda (value units per token): 1
Constructed example: the chapter's revocation (2 tokens, value 8) and irrelevant note (value 0), plus the laboratory's default, changed and transfer cases for the memory-budget function; the thread summary and the old approval's value are defined for this reader, and every selection is computed with the laboratory's function.

Calculated values

Chosen by Equation (15.1)
Current revocation
Total net value
6.0
Tokens used
2 of 2
Excluded before scoring
Old approval
Best sets
1

Chapter example. Net value = decision value - lambda x tokens. Current revocation 8 - 1 x 2 = 6.0; Thread summary 3 - 1 x 2 = 1.0; Old approval excluded: stale authority (v1, current v2); Irrelevant note 0 - 1 x 2 = -2.0. Equation (15.1) keeps Current revocation, for a total of 6.0 using 2 of 2 tokens. Thread summary has net value 1.0 but would push the total past the budget.

Worked steps

  1. Check each record first: it must be fresh (age at most ttl) and any authority record must match version v2.
  2. Excluded: Old approval (stale authority (v1, current v2)).
  3. Net values: Current revocation 8 - 1 x 2 = 6.0; Thread summary 3 - 1 x 2 = 1.0; Irrelevant note 0 - 1 x 2 = -2.0.
  4. Best set within 2 tokens: Current revocation, total 6.0, 2 tokens.

Use the idea

When a context window is limited, give each candidate record a stated value for the decision at hand and a cost, screen out what is expired or superseded, then keep the set with the best total instead of the closest wording.

Where the conclusion applies

Values are declared by the designer and add up across records, which overcounts two records that repeat the same fact. The chapter example values the old approval 0 because it is superseded; the thread summary and its value 3 are defined for this reader, and the laboratory cases use their own declared values. A net value of exactly 0 is a tie and is reported as one.

Common wrong turn: Retrieve the most similar or most valuable-looking text
The chapter says the memory problem is not to retrieve the most similar text, and that embedding similarity is a selection score, not a probability of truth or utility. A high declared value on a superseded record does not survive the freshness and version checks.
What this does not settle

The context-budget and stale-approval numbers are stipulated for teaching rather than drawn from a deployed system.

Chapter 15 source: "What this does not settle".

Check your understanding: In the transfer case (budget 3, now 20, current version doc-C), which record is retrieved, and what is the total value at ฮป = 0?
expired (value 10, written at 10, lifetime 2) is excluded because its age 20 - 10 = 10 exceeds 2. current (3 tokens, value 4) is fresh and bound to doc-C, so it is retrieved for a total of 4.

Chapter 15 source: section "Retrieval is an action under budget". Demonstration C15-D01.

Demonstration 2 of 4

How much can a summary lose?

If a summary scores each action within epsilon of the full history, how much value can following it lose, and when does that claim stop applying?

The summary may undercount the best action by epsilon and overcount a rival by epsilon, so a rival can match or overtake only when the gap is at most 2 epsilon. That is why the regret cannot exceed 2 epsilon, provided every error really is at most epsilon: if the largest error is bigger than the claimed epsilon, the premise fails and the claim is silent. The argument also needs both versions to list the same permitted actions. Switch to the summary that dropped a revocation: the full history no longer permits release (drawn here as a value of -50) while the summary still scores it 10, so the guarantee has nothing to say.

Equation, written in LaTeX: 2\varepsilon

Scroll sideways for the whole equation

ฮต (epsilon) is the largest error the summary is claimed to make on any action that both the summary and the full history list. Regret is the value of the best action under the full history minus the full-history value of the action the summary picks. The equation block shows the limit: the chapter's claim is that this regret is at most 2 epsilon when the conditions below hold. Values are in declared units.

Predict first. In the changed case (full values 4 and 6, summary values 6 and 5) with ฮต = 2, does the summary keep the best action, and is the 2ฮต claim satisfied?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: How much can a summary lose?. Errors placed against the best action. Left, bars of full-history and summary values for 3 actions. Right, the regret of the summary's action, 0.0, against the limit 2 epsilon, 2.0, with epsilon 1; the claim holds: regret within the limit.
What the summary did: Errors placed against the best action, Claimed summary error epsilon: 1
Constructed example: action values 10, 6 and 3 and an unauthorized release worth -50 are defined for this reader; the 2 epsilon bound is the chapter's, and the laboratory's default and changed compression cases (full values 4 and 6, summary values 3 and 5, then 6 and 5) are checked with its compression comparison.

Calculated values

Largest error of the summary
1.0
Claimed epsilon
1
Summary follows
Release
Regret of that action
0.0
Limit 2ฮต
2.0
2ฮต claim
holds: regret within the limit

Summary scores: release 10 - 1 = 9.0, ask owner 6 + 1 = 7.0, abstain 3. It still follows release, so regret = 10 - 10 = 0.0, far inside the limit 2 x 1 = 2.0. Every error here is at most epsilon on actions both versions list, which is what the bound requires.

Worked steps

  1. Full-history values: Release 10, Ask owner 6, Abstain 3.
  2. Summary values: Release 9.0, Ask owner 7.0, Abstain 3.0.
  3. Largest error on shared actions: 1.0; claimed epsilon 1; premise holds.
  4. The summary follows Release.
  5. Regret = best full value 10 minus the full value of that action = 0.0.
  6. Limit 2 x 1 = 2.0; claim: holds: regret within the limit.

Use the idea

Before replacing a long history with a short summary, check whether the summary errs only in its scores, or whether it has changed which actions are permitted. Only the first case is covered by the bound.

Where the conclusion applies

The bound needs a shared way of valuing what comes after the decision and covers only actions both versions list, including the full-history best. The laboratory's cases are one supplied comparison each, so preserving one decision vector does not prove that every future task is preserved. Worst-case errors are placed by hand to show the limit; they are not a claim about any real summary.

Common wrong turn: The 2ฮต bound covers a summary that dropped a revocation
The chapter says that if the full history has approval but the summary omits a new revocation, the action sets differ and no 2ฮต claim applies: query the authoritative external state or choose abstention.
What this does not settle

The compression section's 2ฮต bound presumes a shared continuation convention and covers only common feasible actions, not actions a summary omits entirely; no particular retrieval system is shown to satisfy it here.

Chapter 15 source: "What this does not settle".

Check your understanding: If epsilon is 2.5 and the worst-case errors are applied to scores 10 and 6, what are the summary's two scores and the largest possible regret?
The scores are 10 - 2.5 = 7.5 and 6 + 2.5 = 8.5, so the summary follows the second action, with regret 10 - 6 = 4, within the limit 2 x 2.5 = 5.

Chapter 15 source: section "Compression that preserves action". Demonstration C15-D02.

Demonstration 3 of 4

Stale approval is not support

When an old approval is the most similar record, and deletion may reach only some views, which retrieval policy still blocks an unauthorized release?

The old approval reads closest to the request, so a similarity ranking retrieves it and releases without permission, even though every retrieved sentence was accurate when written. Recency retrieves the newest record. The decision-aware rule gives the revocation net value 8 - 2 = 6 and the others 0 - 2, so it keeps the revocation; if the revocation cannot be reached, nothing has positive net value and the controller abstains. Deletion that reaches only the visible store leaves derived copies competing; deletion from every view removes the approval but does not choose among the rest.

Equation (15.1), written in LaTeX: R^\star\in\arg\max_{R: \sum_{m\in R}c(m)\le B}\{\operatorname E[\text{decision value}\mid R]-\lambda\sum_{m\in R}c(m)\}

Scroll sideways for the whole equation

Similarity is a score for how closely a record's wording matches the request; it is not a probability of truth. Recency means the newest record. Decision aware means choosing the record with the best net value in Equation (15.1), here with room for one record of 2 tokens and ฮป = 1 per token. A view is any place a record's content can survive: the visible store, a selection index, a derived summary, a cache, the model context or a downstream copy.

Predict first. Delete the old approval from the visible store only (a summary index keeps a copy of the same wording). Which policy still releases without permission?

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: Stale approval is not support. A grid of three retrieval policies against four columns: the old approval (Old approval (in the store)), the note, the revocation and no record. Recency only: Release blocked (correct). Similarity only: Unauthorized release. Decision aware: Release blocked (correct).
Similarity of the current revocation: 0.72, Can the current authority be reached?: Yes, How far deletion of the old approval reached: Not deleted
Constructed example: the chapter's similarities 0.98, 0.72 and 0.80 and its values 8 and 0; the other revocation similarity, the unreachable case and the deletion states are defined for this reader, computed with the laboratory's memory-budget function.

Calculated values

Recency only
Release blocked (correct)
Similarity only
Unauthorized release
Decision aware
Release blocked (correct)

Net value with lambda = 1 and 2 tokens per record: old approval 0 - 1 x 2 = (-2.0); note 0 - 1 x 2 = (-2.0); revocation 8 - 1 x 2 = 6.0. Decision aware retrieves the revocation, the only record with positive net value, and blocks the release. Similarity only follows the highest score (old approval, 0.98 > 0.80). Recency only retrieves the newest record it can reach, revocation. Two reports for Similarity only: the recall report says the retrieved old approval is accurate when it was written; the decision report says: unauthorized release.

Worked steps

  1. Candidates: old approval (similarity 0.98), note (similarity 0.80), revocation (similarity 0.72).
  2. Net values: old approval 0 - 1 x 2 = (-2.0); note 0 - 1 x 2 = (-2.0); revocation 8 - 1 x 2 = 6.0.
  3. Decision aware keeps the record with positive net value: release blocked (correct).
  4. Similarity only keeps the closest wording, old approval (0.98): unauthorized release.
  5. Recency only keeps the newest reachable record, revocation: release blocked (correct).
  6. Two reports for Similarity only: the recall report says the retrieved old approval is accurate when it was written; the decision report says: unauthorized release.

Use the idea

Test a memory component by what the controller does next, not only by whether the retrieved text reads well, and query derived views as well as the source store after a deletion.

Where the conclusion applies

Room for one record, the book's three similarity scores for the approval and the note, and a revocation valued 8 against 0 for the others. The index copy of the approval is assumed to keep the approval's wording and so its similarity; the 0.99 revocation similarity and the note's recency are defined for this reader. Recency fails if the newest record is not the authority, and the decision-aware rule is only as good as the declared values and the authority lookup. 'Correct' means consistent with the chapter's rule that a revoked release must not go ahead.

Common wrong turn: Recalling every sentence means the decision is right, and a deleted record is gone
The chapter says a high recall score with unauthorized release is a memory-system failure even if every retrieved sentence was historically accurate, and that absence from one query result is not evidence that later agents cannot use a deleted record.
What this does not settle

The deletion-aware index and conflict-resolution rule appear only as named requirements; building either is left to the implementer.

Chapter 15 source: "What this does not settle".

Check your understanding: A new record has value 5 and 2 tokens, with ฮป = 3. What is its net value, and is it retrieved?
5 - 3 x 2 = (-1). The net value is negative, so Equation (15.1) leaves it out.

Chapter 15 source: section "Stale approval is not support". Demonstration C15-D03.

Demonstration 4 of 4

A record that may no longer be current

How likely must a record be to still be current before it is worth its tokens?

A record's status, expiry and version tell a controller how likely the record is to still hold. That probability scales its value: at p = 0.6 the expected value is 4.8, and one token at ฮป = 1 leaves a net of 3.8. The same record at a higher token price needs a higher p, and in this model, where a record is worth 8 only while it is current, a deleted or expired record has p = 0 and cannot pay for itself.

Equation, written in LaTeX: 0.6\times 8+0.4\times 0=4.8

Scroll sideways for the whole equation

p is the probability that the record is still current. A current record is worth 8 decision-value units and an invalid one is worth 0. The record costs 1 token, and ฮป (lambda) is the price of a token in the same units. The expected value is p x 8 + (1 - p) x 0, and the net value subtracts ฮป times the tokens.

Predict first. At ฮป = 4, where does the net value cross zero? That value of p is the break-even probability.

Your prediction

Choose an example

Scroll sideways for the whole figure

Figure: A record that may no longer be current. Left, three bars: expected value 4.8, token cost 1.0 and net value 3.8 at p 0.6 and lambda 1. Right, net value against p for lambda 1, 2 and 4, with a dashed break-even line at p 0.125.
Probability the record is still current, p: 0.6, Price of a token, lambda: 1
Constructed example: the chapter's Exercise 1 values (value 8, one token, probability 0.6, ฮป = 1); the other probabilities and token prices are defined for this reader.

Calculated values

Expected value
4.80
Token cost
1.00
Net value
3.80
Break-even p
0.125
Decision
Retrieve

Expected value = 0.60 x 8 + 0.40 x 0 = 4.80. Net value = 4.80 - 1 x 1 = 3.80. Break-even p = 1 x 1 / 8 = 0.125. The net value is positive, so the record is worth retrieving (p is above the break-even 0.125). In this model a record is worth 8 only while it is current, so a deleted, expired or superseded one has p = 0: its expected value is 0 x 8 + 1 x 0 = 0 and its net value is (-1.00): the life-cycle fields exist to move p, and the rule then does the rest.

Worked steps

  1. A current record is worth 8 and an invalid one 0; it costs 1 token at lambda 1.
  2. Expected value = p x 8 + (1 - p) x 0 = 0.60 x 8 + 0.40 x 0 = 4.80.
  3. Token cost = lambda x tokens = 1 x 1 = 1.00.
  4. Net value = 4.80 - 1.00 = 3.80.
  5. Break-even p = cost / value = 1.00 / 8 = 0.125; decision: Retrieve.

Use the idea

Record who wrote an item, when it was seen, and when it expires, so the controller can estimate how likely it is to be current and charge retrieval accordingly.

Where the conclusion applies

One record, valued 8 if current and 0 if invalid, with no partial credit. The probability p is assumed to come from the record's fields; an honest estimate of it is the hard part and is not shown here. A record that can authorize an action should need a much higher p than this simple price implies.

Common wrong turn: Durable authority is as safe as durable knowledge
The chapter says broad durable knowledge may be useful but broad durable authority is dangerous: the more a record can authorize, the shorter its expiry and the stronger its required provenance should be.
Check your understanding: A record is worth 10 if current and 0 if invalid, costs 2 tokens, with ฮป = 1 and p = 0.3. What is its net value?
Expected value 0.3 x 10 + 0.7 x 0 = 3.0; net 3.0 - 1 x 2 = 1.0. It is positive, so retrieve it (break-even p = 2 / 10 = 0.2).

Chapter 15 source: section "Exercises". Demonstration C15-D04.