Demonstration 1 of 4
The belief update, one stage at a time
How does a belief change when an action is taken and a report arrives, and which part of the change is prediction and which is correction?
The prediction pushes each state's weight along the action, which is the inner sum of Equation (8.1). The report then multiplies each predicted weight by its likelihood and the denominator rescales what is left. When the action leaves the state fixed, prediction changes nothing and the prior is already the predicted belief. When it moves the state, the prior must not be inserted straight into the weighting step. The west end of the corridor never reaches 0 because nothing the robot has heard rules it out.
Scroll sideways for the whole equation
b is the belief: one probability per state. In the corridor there are four places numbered 1 to 4 from west to east, place 3 is the goal, and an EAST move goes east with chance 0.9 and west with chance 0.1 (a move into a wall leaves the robot in place). In Equation (8.1), a is the action, P(x' | x, a) the chance that action a moves state x to x', Obs(o | x', a) the chance of hearing report o in state x' (here the report depends only on the state), and x'' a stand-in state in the rescaling sum. The notebook cases use two states: the default keeps the state fixed and hears observation 0 from a sensor with rows (0.9, 0.1) and (0.1, 0.9); the transfer case moves the state with rows (0.7, 0.3) and (0.2, 0.8) and hears observation 1 from rows (0.8, 0.2) and (0.3, 0.7). The three steps are the prior, the prediction (inner sum) and the correction (weight by the report, rescale).
Predict first. In the notebook transfer case the prior is (0.8, 0.2). After the prediction step, before the report is used, is the belief still (0.8, 0.2)?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's corridor update (the book's own values for one and two moves at slip chance 0.1) and the laboratory notebook's default and transfer cases, computed with the laboratory's belief function.
Calculated values
- Prior belief
- 0.333, 0.333, 0.000, 0.333
- Predicted belief
- 0.067, 0.300, 0.333, 0.300
- Belief after the report
- 0.100, 0.450, 0.000, 0.450
- Weight the report keeps
- 0.667
- First coordinate (place 1) after the report
- 0.100
The non-goal report keeps weight 0.067 x 1.000 + 0.300 x 1.000 + 0.300 x 1.000 = 0.667. Place 1 after the report = 0.067 x 1.000 / 0.667 = 0.100; place 2 = 0.300 x 1.000 / 0.667 = 0.450. The belief after the report is (0.100, 0.450, 0.000, 0.450). A state the report rules out is exactly 0; a state it cannot rule out keeps weight, which is why the west end never reaches 0.
Worked steps
- Predicted belief (0.067, 0.300, 0.333, 0.300); the report's likelihood in each state is (1.000, 1.000, 0.000, 1.000).
- Weight kept = 0.067 x 1.000 + 0.300 x 1.000 + 0.300 x 1.000 = 0.667.
- Place 1 = 0.067 x 1.000 / 0.667 = 0.100.
- Place 2 = 0.300 x 1.000 / 0.667 = 0.450.
- The belief after the report is (0.100, 0.450, 0.000, 0.450).
Use the idea
Whenever a tool result leaves several explanations open, write down the possibilities, the chance the action moved each one, and the chance each would have produced the report. The three steps then give the new weights without rereading the whole history.
Where the conclusion applies
The corridor, its deterministic goal report and the move probabilities are the chapter's; the notebook cases are constructed. The model is only as good as the move and report probabilities supplied to it; if the real world differs, the update is exact for the wrong world.
Common wrong turn: More moves drive the belief to the truth
What this does not settle
The belief formulation assumes the transition and observation kernels are known, and an agent operating on real tools has neither.
Chapter 8 source: "What this does not settle".
Check your understanding: Start from (1/3, 1/3, 0, 1/3) with a move that goes west with chance 0.2 instead of 0.1. What is the west-end weight after the first non-goal report?
Chapter 8 source: section "The update, worked". Demonstration C08-D01.
Demonstration 2 of 4
The same report under different instruments
If an instrument says pass, how far does the belief move, and what decides that: the report or the instrument model?
Equation (8.1) weights the first state by its report chance and the second by its own, then rescales. When the report chances differ a lot the posterior curve bows far from the diagonal; when they are equal it lies on the diagonal and the report carries nothing. Noise corrupts a reading that still discriminates; aliasing gives distinct states the same reading, and no operation on that reading recovers the difference. Holding the prior fixed, only the assumed instrument changes, and with it the conclusion.
Scroll sideways for the whole equation
The runner has two modeled states: the tests genuinely pass, or they do not (failed or never ran). A passing state always reports pass; the false-positive probability is the chance that a not-passing state also reports pass. The sensor kernels have two states: the informative sensor reports observation 0 with chance 0.9 in the first state and 0.1 in the second; the aliased sensor has identical rows, so observation 0 has chance 0.5 in both. The prior is the probability of the first state before the report. The dotted diagonal is a report that carries no information (posterior equals prior). The dashed line is the belief at which the decision flips.
Predict first. Keep the prior at 0.50 and move from the runner with false-positive probability 0.03 to the one with 0.30. Does the posterior cross the 0.80 threshold, and in which direction?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's test-runner teaching assumptions (true-positive 1, false-positive 0.03 and 0.30, prior 0.5, giving 0.9709 and 0.7692) and the laboratory notebook's default and changed observation rows, with priors 0.2 and 0.8 defined for this reader.
Calculated values
- Chance of this report
- 0.5150
- Posterior, first state
- 0.9709
- Compared with the threshold
- above the threshold
- Change from the prior
- 0.4709
Posterior = 0.50 x 1.00 / (0.50 x 1.00 + 0.50 x 0.03) = 0.5000 / 0.5150 = 0.9709, above the threshold (the 0.80 release threshold). Only the assumed instrument differs between the kernels; the report is the same.
Worked steps
- The report has chance 1.00 in tests pass and 0.03 in not passing.
- Weight of tests pass = 0.50 x 1.00 = 0.5000.
- Weight of not passing = 0.50 x 0.03 = 0.0150.
- Chance of the report = 0.5000 + 0.0150 = 0.5150.
- Posterior = 0.5000 / 0.5150 = 0.9709.
- Compared with the 0.80 release threshold: above the threshold.
Use the idea
Before trusting a pass, a green check or a found-nothing result, ask how often the bad state produces the same output. That one comparison decides whether the report can carry a decision.
Where the conclusion applies
Two states, one report, and instrument probabilities that are teaching assumptions, not measurements of any runner or sensor. Correct arithmetic cannot repair a wrong likelihood. Repeated reports would need a model of how they depend on each other; rerunning the same broken or cached instrument does not automatically supply independent evidence.
Common wrong turn: More readings fix any ambiguity
What this does not settle
A correct application of Bayes' rule cannot repair an incorrect likelihood supplied to it.
Chapter 8 source: "Where the observation kernel comes from".
Check your understanding: With prior 0.50 and a runner that always reports pass when the tests pass, what false-positive probability puts the posterior exactly at 0.80?
Chapter 8 source: section "Where the observation kernel comes from". Demonstration C08-D02.
Demonstration 3 of 4
A belief is valued by its best plan
If every plan is a straight line in the belief, what does the belief's value look like, and how can two equal memories give different values?
Equation (8.2) makes each plan's value linear in the belief, so each plan is a line and the value of the belief is the highest line, a piecewise linear convex curve. The two memories show why the belief matters more than storage size: keeping authorization gives a belief of 1 or 0 and a mean utility of 2, while keeping the formatting letter leaves belief 0.5 and, because release also needs the authorization record, a mean utility of 0. A look is just another action with a negative immediate reward, so the detour is priced by the same maximum.
Scroll sideways for the whole equation
b is the belief, here the probability of the first state (the second has 1 - b). A plan has a reward in each state; its expected reward at the belief, Equation (8.2), is the weighted average, which is the dot product of the plan's vector alpha with b. V(b) is the best such value, the highest line, as in Equations (8.3) and (8.4). The settings are: the chapter's two plans alpha1 = (4, 1) and alpha2 = (0, 3); the one-bit memories of the document-release task (release earns 4 when authorized and -12 when not, holding earns 0); and the three-move detour (commit earns 100 with the chance that the committed route is right; the chapter stipulates 97 for the detour, which equals 100 - 3 if routing then succeeds for certain).
Predict first. At belief 0.25 on the first state, which of the chapter's plans (4, 1) and (0, 3) has the larger expected reward?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's alpha-vector pair, its four-history document-release table with the book's stipulated utilities (+4, -12 and 0) and its three-move detour (100, cost 3, value 97), with the break-even belief 0.97 derived from them; maxima checked with the laboratory's expected-reward function.
Calculated values
- Value of plan 1
- 2.50
- Value of plan 2
- 1.50
- Belief value V(b) (the larger)
- 2.50
- Chosen
- plan 1
- Crossing belief
- 0.333
Belief 0.50. plan 1: 4 x 0.50 + 1 x 0.50 = 2.50; plan 2: 0 x 0.50 + 3 x 0.50 = 1.50. The larger is 2.50, so plan 1 is chosen. The two plans cross at belief 0.333. The belief value is the highest line at the belief, as in Equation (8.4): a plan is a vector, the value its dot product with the belief.
Worked steps
- Belief on the first state = 0.50, so the second state has 0.50.
- plan 1: 4 x 0.50 + 1 x 0.50 = 2.50.
- plan 2: 0 x 0.50 + 3 x 0.50 = 1.50.
- The belief value is the larger: 2.50 (plan 1).
- The lines cross at belief 0.333.
Use the idea
When shrinking a summary or a memory, list the histories it merges and check whether the best allowed action agrees across each merged group. Equal size says nothing; what the retained bit distinguishes is what matters.
Where the conclusion applies
A stipulated toy task: four equally likely histories, authorization fixed between observation and action, and release permitted only with an authoritative authorization record. The detour value 97 is stipulated, not derived. This shows decision sufficiency for one task, not that short summaries are better in general, and the finite-horizon vector form does not make planning cheap.
Common wrong turn: The agent is paid for believing
What this does not settle
Equation (8.3) is a solution concept rather than an algorithm, and the space it ranges over cannot be enumerated.
Chapter 8 source: "What this does not settle".
Check your understanding: For plans (4, 1) and (0, 3), at what belief on the first state do they tie?
Chapter 8 source: section "Where an agent's belief actually lives". Demonstration C08-D03.
Demonstration 4 of 4
A look pays only near the threshold
For which beliefs can a report change the decision, and is the change worth its price?
After each report the controller picks its best action, and Equation (8.5) averages those best rewards by how likely each report is. If the same action is best after every report, the average equals the best reward without looking and the value is exactly 0. Only beliefs where a report can push the decision across its threshold give a positive value, and the chapter's tool peaks at 0.8 where the two actions tie. Free information never hurts; priced information must clear its price.
Scroll sideways for the whole equation
VOI(O) is the average of the best expected reward after each report minus the best expected reward without looking. Pr(o | b) is the chance of report o. In the chapter's verification tool, releasing a supported claim earns 10, an unsupported one loses 40, declining earns 0, the tool reports pass on 90 percent of supported claims and fail on 85 percent of unsupported ones, and releasing beats declining when the belief is above 0.8. The notebook cases use rewards (10, -10) and (-10, 10) for default and changed, and (5, -4) and (0, 2) for transfer. The curve shows the value against the belief at the moment of decision. The price is zero, the case's own price, or the break-even price equal to the value.
Predict first. In the notebook changed case the two observation rows are identical. Is one report worth anything?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's verification-tool instance (rewards 10, -40 and 0; pass 0.90 and fail 0.85; value 3.0 at prior 0.6), and the laboratory notebook's default, changed and transfer cases, computed with the laboratory's belief function; the price 1.0 is defined for this reader.
Calculated values
- Best action without looking
- decline (0.00)
- Value of looking (gross)
- 3.00
- Price of looking
- 0.00
- Net value of looking
- 3.00
- Decision
- look
Belief at the decision (0.60, 0.40). Acting now: release (-10.000), decline 0.000, best 0.00. Each term after a report is weighted by the chance of the report: report 0: release 3.00, decline 0.00, best release 3.00; report 1: release (-13.00), decline 0.00, best decline 0.00. Value of looking = 3.00 + 0.00 - 0.00 = 3.00. The look pays: 3.00 - 0.00 = 3.00 is above zero. Against this case's reward and report rows, a look can have positive value only for beliefs between 0.4000 and 0.9714.
Worked steps
- Belief when the decision is made: (0.60, 0.40).
- Acting now: release (-10.000), decline 0.000; best 0.00.
- After report 0 (weighted by its chance): release 3.00, decline 0.00; best 3.00.
- After report 1 (weighted by its chance): release (-13.00), decline 0.00; best 0.00.
- Value of looking = 3.00 + 0.00 - 0.00 = 3.00.
- Price 0.00: net = 3.00 - 0.00 = 3.00.
Use the idea
Before paying for a check, ask whether any answer it could give would change what you do. If not, its value is 0 however reassuring it sounds. If so, compare the value with the price of the check.
Where the conclusion applies
One report, one decision afterward, a fixed price for looking, and the stipulated rewards and rates. The band describes beliefs, not how often cases land in it. Real tools can fail in ways the rates do not describe, and checking takes time that a single price may not capture. The price 1.0 for the chapter's tool is defined for this reader.
Common wrong turn: A sharper belief is worth its price
What this does not settle
The worked release numbers are constructed for teaching.
Chapter 8 source: "What this does not settle".
Check your understanding: With the chapter's tool, prior 0.5 and no price, what is the value of the look?
Chapter 8 source: section "The condition that makes a look worthless". Demonstration C08-D04.