Chapter links in this companion copy open the corresponding laboratory reading page. The symbols and explanations are retained from the book; its original notation source is preserved in reference-sources/notation-guide.md.
Core symbols carry their meanings across chapters. Explicitly declared local aliases and diagram labels are scoped where they appear; a local count or vertex label does not redefine a global quantity.
That rule costs something. It means a few symbols are less conventional than the ones a specialist in one subfield would expect, because a letter that is standard in information theory may already be doing a job in decision theory. Where this book departs from a field's usual notation, it is to avoid a collision, and the departure is noted at the point of use.
This guide is for returning to, not for memorizing. Every symbol is also defined where it first appears.
dec, Mem, and Reach are roman, never italic single letters.E (\(\operatorname{E}\) or \(\mathrm{E}\)). Both forms denote one operator. A plain E used as a subscript label, such as a delegated destination in Chapter 27, names a thing and is not an operator.t is time, except where a chapter declares otherwise (Chapter 19's matrix M_{st} uses s and t as run indices). Subscript m indexes a member of a declared model family.C, z, u, \(\Phi\), and \(\lambda\). Each meaning is scoped to the chapter listed beside it and restated where it applies.Two objects recur from Part I onward and are worth reading before Chapter 1.
The score of an assembly. \(U_\mathcal V(S)=\mathbb E[Y_S\mid\mathcal V]\) is the population score of
the system built from enabled component set S, under declared evaluation contract \(\mathcal V\);
Y_S is the score of one run of that system.
\(\widehat U_\mathcal V(S)\) is its finite-sample estimate. The contract fixes task distribution,
scoring rule and units, sampling settings, resource budget, and stopping rule. When the set holds
only the frozen model, U describes the model-alone system. As components are
added to the set, the same notation carries through unchanged, which is the point of writing it this
way.
The interaction contrast. \(\Gamma(A;B)\) is the signed difference between the joint gain and the sum of isolated gains under one matched evaluation contract. Four scores identify this nonadditivity, not a mechanism or the necessity of either component. Chapter 2 defines it. \(\rho_\Gamma\) divides the contrast by the positive total joint gain. Nonnegative isolated gains and a nonnegative contrast additionally make the ratio a fraction between zero and one. Negative contrasts remain meaningful signed quantities.
| Symbol | Meaning | Introduced |
|---|---|---|
| \(K_\theta(z\mid c)\) | Model conditional law over outputs given context | 1 |
| \(U_\mathcal V(S)\) | Population task score of assembly S |
1 |
| \(\widehat U_\mathcal V(S)\) | Finite-sample estimate of that task score | 1 |
| \(\operatorname{dec}\) | Decoding map from a distribution to a candidate output | 1 |
| \(\operatorname{eval}_\tau\) | Evaluator with threshold, parser, or tolerance \(\tau\) | 1 |
| \(\theta\) | Model parameters, frozen unless a chapter says otherwise | 1 |
| \(\mathbb E\) | Expectation, as in Chapters 1 and 22 (Chapters 6, 7, 10, 11, 13, 15, 16, and 19 write the same operator as roman E; a plain subscript label such as Chapter 27's p_E is not this operator) |
1 |
| \(\mathcal V\) | Declared evaluation contract: task distribution, scoring rule and units, sampling settings, budget, and stopping rule | 1 |
q, n |
Per-token success probability and target length in the exact-match example, so the score is q^n |
1 |
| \(\Gamma(A;B)\) | Interaction contrast between component sets | 2 |
| \(\rho_\Gamma\) | Interaction share of total joint gain | 2 |
q_m |
Per-attempt success probability for family member m |
3 |
| \(\operatorname{pass}_\tau(q_m)\) | Thresholded pass indicator on q_m |
3 |
L(C)=aC^b+c |
Fitted pretraining-loss curve; C is training compute and a, b, c are fitted constants |
3 |
| \(\psi\) | Fitted forecast parameters | 3 |
| \(f_\psi\), \(\epsilon_m\) | Fitted forecast function and its error for member m |
3 |
| \(\Gamma_m\) | Interaction contrast measured for family member m |
3 |
z_m |
Declared predictor features for member m |
3 |
G |
Action graph: vertices are controller states, edges are allowed transitions | 4 |
V_0 |
Initial vertex set | 4 |
| \(\operatorname{step}\) | One-transition successor map on vertex sets | 4 |
| \(\operatorname{Reach}_G(V_0)\) | Reachable set | 4 |
| \(\operatorname{Act}(v)\) | Candidate actions at vertex v |
4 |
| \(\mathcal{G}\) | Permission grant rule: maps a set of candidate actions to the permitted subset | 4 |
| \(\operatorname{Allowed}_{\mathcal G}(v)\) | Permitted actions after filtering | 4 |
| \(\Pi_{G,v}(p)\) | Infinite-cluster probability | 4 |
p_c(G,v) |
Critical probability | 4 |
d |
Branching factor or degree | 4 |
x_t=(c_t,m_t,w_t,b_t) |
Composite agent state realization | 5 |
w_t |
World or environment state | 5 |
| \(P(x_{t+1}\mid x_t)\) | Composite transition kernel | 5 |
| \(\operatorname{Mem}\) | Memory map from the whole history to the retained memory m_t; the update itself is part of \(\operatorname{Upd}\) |
5 |
| \(\pi_\theta(z,a\mid x)\) | Proposal policy: model law times decoder | 5 |
| \(\kappa\) | Coarse-graining map | 5 |
H(X), I(X;Y) |
Shannon entropy, mutual information | 5 |
| \(\operatorname{Grant}(g\mid x,u)\) | Normalized permission-outcome law, including denial | 5 |
| \(\operatorname{Env}(w',o\mid x,u,g)\) | Joint next-world and observation law | 5 |
| \(\operatorname{Upd}(c',m',b'\mid x,u,g,w',o)\) | Joint context, memory, and budget update law | 5 |
| \(P_{\mathrm{env}}(x'\mid x,u)\) | Controlled transition given a full proposal | 5 |
| \(P_\theta(x'\mid x)\) | Closed-loop transition after averaging over proposal policy | 5 |
u=(z,a) |
Full proposal: model output and decoded action | 5 |
| Symbol | Meaning | Introduced |
|---|---|---|
u(y) |
Utility of a single outcome | 6 |
y |
An outcome of an action | 6 |
| \(\mathcal{A}\) | The action set available at a state | 6 |
| \(\operatorname{EU}(a)\) | Expected utility of an action | 6 |
| \(\operatorname{CE}\) | Certainty equivalent | 6 |
| \(\operatorname{Ret}_t\) | Return, the discounted sum of future rewards | 7 |
r_t |
Reward received on the transition out of x_t |
7 |
| \(\gamma\) | Discount factor | 7 |
| \(\pi(a\mid x)\) | Stochastic policy | 7 |
| \(V^{\pi}\), \(V^{\star}\) | State value under a policy, and under an optimal policy | 7 |
| \(Q^{\pi}\) | Action value under a policy | 7 |
| \(\Phi(x)\) | Local shaping potential | 7 |
| \(\mathcal{X}\) | The finite state set | 8 |
| \(\mathbf{b}_t\) | Belief, a distribution over \(\mathcal{X}\) at time t |
8 |
| \(\mathcal{B}\) | The belief simplex, the set of distributions over \(\mathcal{X}\) | 8 |
| \(\operatorname{Obs}(o\mid x,a)\) | Observation kernel: the chance of seeing o on arriving in x after a |
8 |
| \(\bar r(\mathbf{b},a)\) | Expected reward of action a under belief \(\mathbf{b}\) |
8 |
| \(J(x',o\mid x,a)\) | Filtered joint next-state and observation law | 8 |
| \(\Gamma_H\), \(\alpha\) | Finite-horizon policy-vector set and one vector | 8 |
| \(\operatorname{VOI}(O)\) | Value of observing O before a declared decision; Chapter 27 writes \(\operatorname{VOI}(o)\) for one realized observation |
8 |
g(v), h(v) |
True cost from the start to vertex v, and true cost from v to the nearest goal |
9 |
f(v) |
Cost of the best path constrained to pass through v |
9 |
| \(\hat g(v)\), \(\hat h(v)\), \(\hat f(v)\) | The computable estimates of each | 9 |
| \(\operatorname{cost}(u,v)\) | Cost of the edge from u to v |
9 |
s, \(\mathcal{T}\) |
Start vertex and goal set | 9 |
| \(\operatorname{Exp}(A,G)\) | Number of vertices algorithm A expands on graph G |
9 |
h(u,v) |
True minimum path cost from u to v, distinct from the heuristic estimate \(\hat h(v)\) |
9 |
| \(\mathcal{I}_o\) | Initiation set of option o: the states where the option may start |
10 |
| \(\pi_o\) | Option o's internal policy, mapping execution information to a primitive action |
10 |
| \(\beta_o\) | Option o's termination condition |
10 |
| \(\mu\) | Policy over options: the parent policy that selects which option to run | 10 |
a_t |
Action selected at step t; in Chapter 11, the next action selected after t completed pulls |
11 |
| \(\operatorname{Reg}(T)\) | Regret accumulated over T rounds |
11 |
| \(\hat\mu_a\) | Empirical mean payoff of action a |
11 |
N_t(a) |
Number of pulls of action a among the t completed pulls, before the next decision |
11 |
| \(\operatorname{Gain}(O)\) | Decision-relevant information in an observation | 11 |
| \(\hat V_t(x)\) | The agent's current estimate of the value of state x |
12 |
| \(\delta_t\) | Temporal-difference error at step t |
12 |
| \(\alpha\) | Step size for an update | 12 |
| \(\lambda\) | Weighting between the one-step estimate and the realized outcome | 12 |
| \(\operatorname{Ret}^{\lambda}_t\) | The weighted return that interpolates between them | 12 |
| \(\operatorname{Pot}(x)\) | Bounded shaping potential on declared state | 13 |
r'_t, G'_0 |
Shaped reward and return; the original reward r and the return keep their meaning (Chapter 13 gives the return a local alias for Chapter 7's \(\operatorname{Ret}_t\)) |
13 |
| \(\phi\) | Trainable policy parameters, distinct from the frozen model parameters \(\theta\) | 13 |
| \(\hat P(x'\mid x,a)\) | Learned transition-model estimate | 14 |
| \(\widehat{\operatorname{Obs}}\) | Learned observation-kernel estimate | 14 |
| \(\epsilon\) | Uniform one-step total-variation model error | 14 |
| \(\tau_{\text{samp}}\) | Sampling temperature of a generative model | 14 |
R |
Upper bound on the common expected immediate state-action reward | 14 |
| \(H_{\max}\) | Declared finite planning cap | 14 |
| \(R^\star\) | Retained record set chosen under the token budget | 15 |
c(m), B, \(\lambda\) (Chapter 15) |
Raw token cost of record m, the token budget, and decision-value units per token |
15 |
| \(\varepsilon\), \(2\varepsilon\) (Chapter 15) | Summary value tolerance and the resulting regret bound | 15 |
| \(\operatorname{Cov}(k)\) | Coverage at k: the chance that at least one of k samples is correct |
16 |
| \(\operatorname{Sel}(k)\) | Selection at k: the chance that the returned sample is correct |
16 |
| \(\widehat c_{\mathrm{success}}\) | Total attempt cost divided by authorized confirmed completions | 16 |
| \(\pi(k)\) | Measured chance that the selector's top-ranked sample is correct given k candidates that contain at least one correct candidate |
16 |
d |
Probability that the selector scores one correct sample above one incorrect sample | 16 |
| \(\operatorname{eff}(T,x)\) | State resulting from applying tool T in state x |
17 |
| \(\operatorname{resp}(T,x)\) | Report returned by tool T in state x |
17 |
| \(\operatorname{undo}(T)\) | Recovery tool for an effect of T |
17 |
| \(\operatorname{Rep}(x)\) | Conservative set of eligible repeat attempts | 17 |
| \(\operatorname{Ready}(T,x)\) | Current authorization and preconditions for a repeat attempt | 17 |
| \(\operatorname{Absent}(T,x)\) | Authoritative terminal non-application of the original attempt | 17 |
| \(\operatorname{Restored}(T,x)\) | Verified recovery of relevant consequences and preconditions | 17 |
| \(\beta\), \(c_{\text{dup}}\), \(c_{\text{miss}}\) | Belief that the effect already landed, cost of a duplicate effect, and cost of the effect never happening | 17 |
| \(\operatorname{Exec}_{\mathrm{ui}}(\tilde a\mid x,a)\) | Normalized realized-operation law | 18 |
| \(P_{\mathrm{ui}}\) | Execution-composed world kernel | 18 |
| \(\operatorname{Act}_{\mathrm{ui}}\) | Finite realized-operation set including declared denial/no-effect operations | 18 |
| \(\mathcal A_{\mathrm{rob}}(\mathbf b)\) | Intersection of correctly specified authorization sets over belief support | 18 |
| \(\lambda_{\mathrm{ui}}, \operatorname{Fresh}(\Delta)\) | Material-change rate and no-invalidating-change probability | 18 |
| \(\mathcal A_{\mathrm{auth}}(x)\) | Commands whose protected effect is authorized in state x |
18 |
| \(u_i(\pi_1,\pi_2)\) | Payoff to party i under a pair of policies |
19 |
| \(\pi_i^{(s)}\) | Party i policy from training run s |
19 |
| \(\operatorname{diag}(M)\), \(\operatorname{off}(M)\) | Matched-run and cross-run mean returns | 19 |
| \(\operatorname{JPC}(M)\) | Proportional cross-play loss | 19 |
| \(\operatorname{BR}_i(\pi_{-i})\) | Best response by party i to the other policy |
19 |
M |
Matched matrix whose entry M_{st} pairs run s against run t |
19 |
| \(\operatorname{Know}_i(F)\) | Party i knows proposition F |
20 |
| \(\operatorname{MK}_{\mathcal P}^n(F)\) | n levels of mutual knowledge among party set \(\mathcal P\) |
20 |
| \(\operatorname{CK}_{\mathcal P}(F)\) | Common knowledge of F among party set \(\mathcal P\) |
20 |
q_e |
Nonnegative flow on traffic-network edge e |
21 |
| \(\ell_e(q_e)\) | Latency incurred on edge e at flow q_e |
21 |
| \(\operatorname{TL}(q)\) | Total latency of a feasible network flow q |
21 |
| \(q^{\mathrm{NE}}\) | Nash-equilibrium flow | 21 |
| \(q^{\star}\) | Total-latency-minimizing flow | 21 |
| \(\ell^{\mathrm{mc}}_e(q_e)\) | Marginal-cost latency imposed by edge e |
21 |
| \(q_{\mathrm{upper}}\), \(q_{\mathrm{lower}}\), \(q_{\mathrm{middle}}\) | Flows on the three directed Braess routes S-U-T, S-L-T, S-U-L-T | 21 |
| \(\gamma_{\mathrm{cap}}\) | Extra-traffic fraction in the bicriteria routing comparison | 21 |
C |
Declared episode-cost random variable for tail-risk accounting | 22 |
| \(\operatorname{CVaR}_{\alpha}(C)\) | Conditional value at risk of C at declared tail level \(\alpha\) |
22 |
z |
Candidate tail cutoff in CVaR's infimum representation | 22 |
B, b_t |
Declared bounded-horizon risk budget and remaining budget before decision t |
22 |
| \(\widehat r_t(a_t)\) | Conservative charged risk estimate for action a_t at time t |
22 |
| \(\operatorname{Ret}^{\pi}\) | Expected discounted reward under policy \(\pi\) | 22 |
| \(\mathcal I\) | Set of states in which a named invariant holds | 22 |
| \(\operatorname{Cost}_i^{\pi}\) | Expected discounted auxiliary-cost return | 22 |
d_i |
Declared limit on auxiliary cost i |
22 |
| \(\mathcal T\) | Declared threat model | 22 |
| \(\varepsilon_{\mathrm{safe}}\) | Deliberate safety margin | 22 |
| \(A_{C_i}^{\pi_k}(s,a)\) | Constraint advantage for auxiliary cost i |
22 |
| \(\epsilon_i\) | Largest candidate-policy expected constraint advantage | 22 |
| \(\delta\) | Ideal CPO step-size bound | 22 |
| \(\lambda_{\mathrm{safe}}\) | Fixed multiplier between reward and one auxiliary cost | 22 |
| \(\mathcal C_{\mathrm{parent}}\), \(\mathcal C_{\mathrm{child}}\) | Declared capability sets of a delegating parent and a delegated child | 23 |
M_t, e_t |
Monitor's enforcement-layer record at time t, and one observed event |
23 |
| \(\operatorname{Mon}\) | Monitor update map from prior record and one observed event to the next record | 23 |
\(\delta\), v, r |
Document identity, version, and recipient in the publish predicate | 23 |
| \(\operatorname{Approved}(M_t)\) | Set of \((\delta,v,r)\) triples the monitor's current record authorizes for release | 23 |
| \(\operatorname{Publish}(\delta,v,r,M_t)\) | Boolean predicate: is this exact document-version-recipient triple currently approved | 23 |
| \(\mathcal F\) | One declared, finite attack family under a fixed threat model | 23 |
| \(\operatorname{Risk}(\mathcal F)\) | Worst-case violation probability across \(\mathcal F\) | 23 |
k |
Number of samples drawn | 24 |
p |
Probability that one sample is correct | 24 |
| \(N_{\mathrm{eval}}\) | Number of evaluated task instances | 24 |
| \(\Delta_{\mathrm{match}}(\theta)\) | Score difference for frozen model \(\theta\) between a named benchmark and its matched evaluation | 24 |
| \(\zeta\) | Declared threshold on a continuous metric when defining a thresholded scale onset | 24 |
| \(\operatorname{Onset}_{\zeta}\) | First declared family scale at which the thresholded metric reaches \(\zeta\) | 24 |
| \(\mathcal K\) | Fixed same-bank task set used to compare coverage with deployed selection | 24 |
| \(\widehat{\operatorname{Cov}}_{\mathcal K}(k)\), \(\widehat{\operatorname{Sel}}_{\mathcal K}(k)\) | Empirical same-bank coverage and deployed-selector success after k candidates per task |
24 |
| \(\mathcal C\), \(\xi\) | Finite tested configuration set and one configuration index | 24 |
| \(\Pr\) | Probability over Equation 24.4's declared sampling or resampling design | 24 |
| \(\operatorname{Lower}^{\mathrm{sim}}_{\alpha_{\mathrm{tail}}}(\xi)\) | Lower score bound for configuration \(\xi\) from a procedure with joint coverage across all tested configurations | 24 |
| \(\operatorname{SE}_{\mathrm{match}}\) | Estimated standard error of a difference of independent bank pass proportions | 24 |
| \(\alpha_{\mathrm{tail}}\) | Tail probability level for the simultaneous lower-bound guarantee | 24 |
| \(\operatorname{LCF}_{\alpha_{\mathrm{tail}}}\) | Simultaneous lower confidence frontier across tested configurations | 24 |
| \(\phi_t\), \(\phi'\) | Parent and candidate versions of the agent procedure around frozen weights | 25 |
\(\widehat\Delta_G(\phi')\), G |
Guard-set estimate and its frozen guard cases for the named candidate | 25 |
| \(\widehat C_G(\phi')\), \(\mathcal G\) | Guard cost estimate and authority envelope for the proposed release | 25 |
\(\tau\), c, \(\varepsilon_{\rm safe}\) |
Local uplift threshold, cost limit, and reserved safety margin | 25 |
D, \(\widehat V_D\) |
Development set and a procedure's estimated value on it | 25 |
| \(\mathcal M\), \(\Phi_t\) | Declared bounded family of procedure changes and the finite candidate collection compared on D |
25 |
m, \(Z_i(\phi')\) |
Number of guard cases and the paired candidate-minus-parent outcome on guard case i, bounded in [-1,1] |
25 |
| \(\Delta(\phi')\) | Population mean of \(Z_i(\phi')\): the candidate's true uplift, an unknown parameter | 25 |
| \(\operatorname{Authorized}_{\mathcal G}\), \(\operatorname{Rollback}\) | Boolean authority predicate for the proposed release and the check for a tested path back to the parent | 25 |
n, c |
Single-task aliases of n_i, c_i, the source-pool size and the number of passing samples for task i |
26 |
| \(\operatorname{Del}(d,\Delta)\) | Delegation action to destination d, with stated maximum delay \(\Delta\) before a response, expiration, or fallback |
27 |
p_E(x) |
Probability that delegated destination E returns a correct resolution at input or state x |
27 |
| \(\mathcal R_{\mathrm{auth}}(x)\) | Authorized decision routes in state x |
27 |
| \(\varrho\) | One member of \(\mathcal R_{\mathrm{auth}}(x)\), an authorized decision route | 27 |
| \(V(\varrho,x)\) | Declared value of authorized route \(\varrho\) in state x |
27 |
Chapter 25 uses \(\phi_t\), \(\phi'\), G, \(\widehat\Delta_G\), \(\widehat C_G\), \(\tau\),
c, \(\varepsilon_{\rm safe}\), and \(\mathcal G\) only for its constructed procedure-release rule.
\(\theta\) keeps its book-wide meaning: frozen model parameters. These local release symbols do not
rename the book's standing state, action, score, or authority notation.
\(\operatorname{Pot}(x)\) in Chapter 13 is a bounded shaping potential; its terminal boundary affects policy comparisons. Chapter 16's cost per authorized confirmed success is undefined when no run succeeds. In Chapter 18, \(\operatorname{Exec}_{\mathrm{ui}}\) connects a proposed command to a realized operation. \(\mathcal A_{\mathrm{rob}}\) retains actions feasible throughout the declared belief support. \(\operatorname{Fresh}(\Delta)\) describes no material interface change under a stated constant-rate model.
Local indices, declared aliases, and diagram labels retain their stated scope; return to the defining chapter for domains and assumptions.