{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "7e0a382325cb",
   "metadata": {},
   "source": [
    "# Chapter 7: The Equation That Looks Ahead\n",
    "\n",
    "Preparation often looks worse than acting when you examine only the next reward. A document controller can collect an immediate modest payoff or spend one step preparing a larger release. The preparation cost is real, but so is the later opportunity. Which action wins depends on how much future remains.\n",
    "\n",
    "This notebook turns that comparison into a finite state model. Every reward, transition probability, terminal value, and available action is declared. The goal is to understand the equation that looks ahead, then inspect how a shortened horizon changes the policy. Optimal means optimal within this supplied model; it does not mean the controller has discovered every possible real-world alternative.\n",
    "\n",
    "**Outcome:** Compute a finite-horizon optimal policy and its value by backward induction.\n",
    "\n",
    "- Inspect the declared input contract\n",
    "- Predict the hand-checkable case\n",
    "- Run the shared computation\n",
    "- Change the critical assumption\n",
    "- Apply the method to the transfer data\n",
    "\n",
    "**Guided route:** Run the worked calculation, inspect its figure, change the stated assumption, and try the transfer case. Read the explanations beside each result before opening the answers.\n",
    "\n",
    "**Deeper route:** First read the mathematics and canonical equation reference. Audit the input contract, predict the changed result, then inspect the shared chapter implementation and solve the questions independently. Both routes use the same calculations and preserve the equations."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "a7781f3ee7dc",
   "metadata": {},
   "source": [
    "## Technical Requirements\n",
    "\n",
    "Python 3.11 or later, the complete laboratory folder, and the notebook dependencies listed in `requirements-notebooks.txt` (the launcher's **Install notebook tools** choice installs them; see START-HERE). Standard-library chapter commands also support Python 3.10. No API key, model account or network call is used by this experiment.\n",
    "\n",
    "Prior knowledge:\n",
    "\n",
    "- Python lists and dictionaries\n",
    "- The mapped chapter and its notation"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "2af031277723",
   "metadata": {},
   "source": [
    "## The question and its mathematics\n",
    "\n",
    "Let V_h(s) be the best expected return from state s with h decisions remaining. Terminal continuation defines V_0(s). For h>=1, Q_h(s,a)=r(s,a)+gamma sum_s' P(s'|s,a)V_(h-1)(s'), and V_h(s)=max_a Q_h(s,a). Recording the maximizing action gives a policy indexed by state and remaining horizon.\n",
    "\n",
    "Backward induction starts with V_0 and builds V_1,V_2,...,V_H. Each stage uses only the previous stage's values. This separation prevents a within-stage update from accidentally pretending that extra decisions remain. Rewards are expected immediate state-action rewards; their outcome dependence is already integrated into the declared number.\n",
    "\n",
    "Stopping must be represented as an explicit transition or action. The example's done state has a zero-reward self-loop, so extra horizon after termination adds nothing. Discount gamma scales later reward relative to present reward. Gamma=1 is valid for this finite computation, even though some infinite-horizon formulas require a strict discount. A terminal value can encode a declared continuation estimate, but it is not learned here."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "dbc90041bd06",
   "metadata": {},
   "source": [
    "## A calculation you can run\n",
    "\n",
    "Encode states as a JSON object, with an action list for every state. Each action supplies its immediate reward and a normalized transition vector aligned with next-state identifiers. Validation rejects unknown destination states, absent actions, invalid probabilities, and unsupported horizons.\n",
    "\n",
    "The code initializes terminal values and repeatedly evaluates all actions from each state. Its returned tables preserve every value stage and policy stage. The figure shows the start state's value against remaining decisions, making horizon effects visible. Compare the first stage with the full-horizon stage before reading the selected action. In the changed case only horizon changes from three decisions to one. All transition and reward assumptions remain fixed.\n",
    "\n",
    "The next cell finds the bundle and imports the same computation used by the chapter skill. It does not change your system Python."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 1,
   "id": "42ebefc0e5c1",
   "metadata": {},
   "outputs": [],
   "source": [
    "from pathlib import Path\n",
    "import sys, json\n",
    "LAB_ROOT = next((p for p in [Path.cwd(), *Path.cwd().parents] if (p / \"lab-manifest.json\").is_file()), None)\n",
    "if LAB_ROOT is None:\n",
    "    raise RuntimeError(\"Open this notebook from the complete extracted laboratory folder.\")\n",
    "sys.path.insert(0, str(LAB_ROOT / \"src\"))\n",
    "from math_ai_agents.core import analyze, report_text\n",
    "from math_ai_agents.plotting import figure_svg\n",
    "from IPython.display import SVG, display\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "7fffde8c0edb",
   "metadata": {},
   "source": [
    "Set the declared inputs below. These are constructed teaching values, not measurements from a production agent. Change a value only after predicting what it should change."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 2,
   "id": "11b79482cb45",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Chapter 7: finite-horizon-planning\n",
      "When should a controller accept an immediate payoff instead of preparing a better future?\n",
      "Evidence: constructed teaching example\n",
      "\n",
      "Calculated quantities:\n",
      "{\n",
      "  \"start_value\": 5.0,\n",
      "  \"policy_by_remaining_steps\": [\n",
      "    {\n",
      "      \"done\": \"stop\",\n",
      "      \"draft\": \"cash\",\n",
      "      \"ready\": \"release\"\n",
      "    },\n",
      "    {\n",
      "      \"done\": \"stop\",\n",
      "      \"draft\": \"prepare\",\n",
      "      \"ready\": \"release\"\n",
      "    },\n",
      "    {\n",
      "      \"done\": \"stop\",\n",
      "      \"draft\": \"prepare\",\n",
      "      \"ready\": \"release\"\n",
      "    }\n",
      "  ],\n",
      "  \"values_by_remaining_steps\": [\n",
      "    {\n",
      "      \"done\": 0.0,\n",
      "      \"draft\": 0.0,\n",
      "      \"ready\": 0.0\n",
      "    },\n",
      "    {\n",
      "      \"done\": 0.0,\n",
      "      \"draft\": 2.0,\n",
      "      \"ready\": 6.0\n",
      "    },\n",
      "    {\n",
      "      \"done\": 0.0,\n",
      "      \"draft\": 5.0,\n",
      "      \"ready\": 6.0\n",
      "    },\n",
      "    {\n",
      "      \"done\": 0.0,\n",
      "      \"draft\": 5.0,\n",
      "      \"ready\": 6.0\n",
      "    }\n",
      "  ]\n",
      "}\n",
      "\n",
      "Interpretation:\n",
      "Backward induction compares the immediate reward with the discounted value of the next state at each remaining horizon.\n",
      "\n",
      "Assumptions:\n",
      "- Finite fully observed state model.\n",
      "- Terminal values and discount are declared.\n",
      "- Stopping is represented by an explicit action and terminal state.\n",
      "\n",
      "Limitations:\n",
      "- The computed policy is optimal only inside the supplied finite model.\n",
      "\n",
      "Execution: completed locally; constructed inputs are not deployment measurements.\n"
     ]
    }
   ],
   "source": [
    "chapter = 7\n",
    "inputs = {'horizon': 3,\n",
    " 'discount': 1,\n",
    " 'start': 'draft',\n",
    " 'states': {'draft': {'actions': [{'name': 'cash',\n",
    "                                   'reward': 2,\n",
    "                                   'probabilities': [1],\n",
    "                                   'next_states': ['done']},\n",
    "                                  {'name': 'prepare',\n",
    "                                   'reward': -1,\n",
    "                                   'probabilities': [1],\n",
    "                                   'next_states': ['ready']}]},\n",
    "            'ready': {'actions': [{'name': 'release',\n",
    "                                   'reward': 6,\n",
    "                                   'probabilities': [1],\n",
    "                                   'next_states': ['done']}]},\n",
    "            'done': {'actions': [{'name': 'stop',\n",
    "                                  'reward': 0,\n",
    "                                  'probabilities': [1],\n",
    "                                  'next_states': ['done']}]}}}\n",
    "report = analyze(chapter, inputs)\n",
    "# This input was explicitly taken from the teaching fixture.\n",
    "report['evidence_kind'] = 'constructed teaching example'\n",
    "print(report_text(report))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b0face462b9d",
   "metadata": {},
   "source": [
    "At one remaining decision, cash gives 2 and prepare gives -1, so cash wins. At two remaining decisions, preparation gives -1+6=5, while cash still gives 2. With three decisions the value remains 5 because done adds no reward.\n",
    "\n",
    "The policy table therefore changes at draft when the remaining horizon crosses from one to two. The changed case returns cash at the start with value 2. This is a horizon-driven reversal, not inconsistent arithmetic. A controller that ignores its deadline can select preparation and then run out of opportunities before collecting the release reward.\n",
    "\n",
    "The plot below uses the calculated quantities. Read each panel's units before comparing its values."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 3,
   "id": "590f57e74048",
   "metadata": {},
   "outputs": [
    {
     "name": "stderr",
     "output_type": "stream",
     "text": [
      "Matplotlib is building the font cache; this may take a moment.\n"
     ]
    },
    {
     "data": {
      "image/svg+xml": [
       "<svg xmlns:xlink=\"http://www.w3.org/1999/xlink\" xmlns=\"http://www.w3.org/2000/svg\" width=\"576pt\" height=\"237.6pt\" viewBox=\"0 0 576 237.6\" version=\"1.1\"><title>Calculated chapter experiment</title><desc>Labeled plot of the explicitly supplied chapter inputs. See the adjacent explanation for assumptions.</desc>\n",
       " <metadata>\n",
       "  <rdf:RDF xmlns:dc=\"http://purl.org/dc/elements/1.1/\" xmlns:cc=\"http://creativecommons.org/ns#\" xmlns:rdf=\"http://www.w3.org/1999/02/22-rdf-syntax-ns#\">\n",
       "   <cc:Work>\n",
       "    <dc:type rdf:resource=\"http://purl.org/dc/dcmitype/StillImage\"/>\n",
       "    <dc:format>image/svg+xml</dc:format>\n",
       "    <dc:creator>\n",
       "     <cc:Agent>\n",
       "      <dc:title>Mathematics of AI Agents Laboratory</dc:title>\n",
       "     </cc:Agent>\n",
       "    </dc:creator>\n",
       "   </cc:Work>\n",
       "  </rdf:RDF>\n",
       " </metadata>\n",
       " <defs>\n",
       "  <style type=\"text/css\">*{stroke-linejoin: round; stroke-linecap: butt}</style>\n",
       " </defs>\n",
       " <g id=\"figure_1\">\n",
       "  <g id=\"patch_1\">\n",
       "   <path d=\"M 0 237.6  L 576 237.6  L 576 0  L 0 0  z \" style=\"fill: #ffffff\"/>\n",
       "  </g>\n",
       "  <g id=\"axes_1\">\n",
       "   <g id=\"patch_2\">\n",
       "    <path d=\"M 38.602344 195.477656  L 565.2 195.477656  L 565.2 60.080234  L 38.602344 60.080234  z \" style=\"fill: #ffffff\"/>\n",
       "   </g>\n",
       "   <g id=\"matplotlib.axis_1\">\n",
       "    <g id=\"xtick_1\">\n",
       "     <g id=\"line2d_1\">\n",
       "      <defs>\n",
       "       <path id=\"m356609981a\" d=\"M 0 0  L 0 3.5  \" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </defs>\n",
       "      <g>\n",
       "       <use xlink:href=\"#m356609981a\" x=\"62.538601\" y=\"195.477656\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_1\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"62.538601\" y=\"210.075312\" transform=\"rotate(-0 62.538601 210.075312)\">0</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"xtick_2\">\n",
       "     <g id=\"line2d_2\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m356609981a\" x=\"222.113648\" y=\"195.477656\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_2\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"222.113648\" y=\"210.075312\" transform=\"rotate(-0 222.113648 210.075312)\">1</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"xtick_3\">\n",
       "     <g id=\"line2d_3\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m356609981a\" x=\"381.688696\" y=\"195.477656\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_3\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"381.688696\" y=\"210.075312\" transform=\"rotate(-0 381.688696 210.075312)\">2</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"xtick_4\">\n",
       "     <g id=\"line2d_4\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m356609981a\" x=\"541.263743\" y=\"195.477656\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_4\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"541.263743\" y=\"210.075312\" transform=\"rotate(-0 541.263743 210.075312)\">3</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"text_5\">\n",
       "     <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"301.901172\" y=\"224.076094\" transform=\"rotate(-0 301.901172 224.076094)\">remaining decisions</text>\n",
       "    </g>\n",
       "   </g>\n",
       "   <g id=\"matplotlib.axis_2\">\n",
       "    <g id=\"ytick_1\">\n",
       "     <g id=\"line2d_5\">\n",
       "      <path d=\"M 38.602344 189.323228  L 565.2 189.323228  \" clip-path=\"url(#p5eb62942da)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_6\">\n",
       "      <defs>\n",
       "       <path id=\"m793aa3291f\" d=\"M 0 0  L -3.5 0  \" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </defs>\n",
       "      <g>\n",
       "       <use xlink:href=\"#m793aa3291f\" x=\"38.602344\" y=\"189.323228\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_6\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"31.602344\" y=\"193.122056\" transform=\"rotate(-0 31.602344 193.122056)\">0</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_2\">\n",
       "     <g id=\"line2d_7\">\n",
       "      <path d=\"M 38.602344 164.705515  L 565.2 164.705515  \" clip-path=\"url(#p5eb62942da)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_8\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m793aa3291f\" x=\"38.602344\" y=\"164.705515\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_7\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"31.602344\" y=\"168.504343\" transform=\"rotate(-0 31.602344 168.504343)\">1</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_3\">\n",
       "     <g id=\"line2d_9\">\n",
       "      <path d=\"M 38.602344 140.087802  L 565.2 140.087802  \" clip-path=\"url(#p5eb62942da)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_10\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m793aa3291f\" x=\"38.602344\" y=\"140.087802\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_8\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"31.602344\" y=\"143.88663\" transform=\"rotate(-0 31.602344 143.88663)\">2</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_4\">\n",
       "     <g id=\"line2d_11\">\n",
       "      <path d=\"M 38.602344 115.470089  L 565.2 115.470089  \" clip-path=\"url(#p5eb62942da)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_12\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m793aa3291f\" x=\"38.602344\" y=\"115.470089\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_9\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"31.602344\" y=\"119.268917\" transform=\"rotate(-0 31.602344 119.268917)\">3</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_5\">\n",
       "     <g id=\"line2d_13\">\n",
       "      <path d=\"M 38.602344 90.852376  L 565.2 90.852376  \" clip-path=\"url(#p5eb62942da)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_14\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m793aa3291f\" x=\"38.602344\" y=\"90.852376\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_10\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"31.602344\" y=\"94.651204\" transform=\"rotate(-0 31.602344 94.651204)\">4</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_6\">\n",
       "     <g id=\"line2d_15\">\n",
       "      <path d=\"M 38.602344 66.234663  L 565.2 66.234663  \" clip-path=\"url(#p5eb62942da)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_16\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m793aa3291f\" x=\"38.602344\" y=\"66.234663\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_11\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"31.602344\" y=\"70.033491\" transform=\"rotate(-0 31.602344 70.033491)\">5</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"text_12\">\n",
       "     <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"18.8375\" y=\"127.778945\" transform=\"rotate(-90 18.8375 127.778945)\">discounted reward</text>\n",
       "    </g>\n",
       "   </g>\n",
       "   <g id=\"line2d_17\">\n",
       "    <path d=\"M 62.538601 189.323228  L 222.113648 140.087802  L 381.688696 66.234663  L 541.263743 66.234663  \" clip-path=\"url(#p5eb62942da)\" style=\"fill: none; stroke: #31586b; stroke-width: 1.8; stroke-linecap: square\"/>\n",
       "    <defs>\n",
       "     <path id=\"mf8f1db4e83\" d=\"M 0 2  C 0.530406 2 1.03916 1.789267 1.414214 1.414214  C 1.789267 1.03916 2 0.530406 2 0  C 2 -0.530406 1.789267 -1.03916 1.414214 -1.414214  C 1.03916 -1.789267 0.530406 -2 0 -2  C -0.530406 -2 -1.03916 -1.789267 -1.414214 -1.414214  C -1.789267 -1.03916 -2 -0.530406 -2 0  C -2 0.530406 -1.789267 1.03916 -1.414214 1.414214  C -1.03916 1.789267 -0.530406 2 0 2  z \" style=\"stroke: #31586b\"/>\n",
       "    </defs>\n",
       "    <g clip-path=\"url(#p5eb62942da)\">\n",
       "     <use xlink:href=\"#mf8f1db4e83\" x=\"62.538601\" y=\"189.323228\" style=\"fill: #31586b; stroke: #31586b\"/>\n",
       "     <use xlink:href=\"#mf8f1db4e83\" x=\"222.113648\" y=\"140.087802\" style=\"fill: #31586b; stroke: #31586b\"/>\n",
       "     <use xlink:href=\"#mf8f1db4e83\" x=\"381.688696\" y=\"66.234663\" style=\"fill: #31586b; stroke: #31586b\"/>\n",
       "     <use xlink:href=\"#mf8f1db4e83\" x=\"541.263743\" y=\"66.234663\" style=\"fill: #31586b; stroke: #31586b\"/>\n",
       "    </g>\n",
       "   </g>\n",
       "   <g id=\"patch_3\">\n",
       "    <path d=\"M 38.602344 195.477656  L 38.602344 60.080234  \" style=\"fill: none; stroke: #000000; stroke-width: 0.8; stroke-linejoin: miter; stroke-linecap: square\"/>\n",
       "   </g>\n",
       "   <g id=\"patch_4\">\n",
       "    <path d=\"M 38.602344 195.477656  L 565.2 195.477656  \" style=\"fill: none; stroke: #000000; stroke-width: 0.8; stroke-linejoin: miter; stroke-linecap: square\"/>\n",
       "   </g>\n",
       "   <g id=\"text_13\">\n",
       "    <text style=\"font-size: 11px; font-family: 'DejaVu Sans'; text-anchor: start\" x=\"38.602344\" y=\"54.080234\" transform=\"rotate(-0 38.602344 54.080234)\">optimal start value</text>\n",
       "   </g>\n",
       "  </g>\n",
       "  <g id=\"text_14\">\n",
       "   <text style=\"font-size: 12px; font-family: 'DejaVu Sans'; text-anchor: start\" x=\"46.08\" y=\"13.870125\" transform=\"rotate(-0 46.08 13.870125)\">Chapter 7: finite horizon planning</text>\n",
       "  </g>\n",
       " </g>\n",
       " <defs>\n",
       "  <clipPath id=\"p5eb62942da\">\n",
       "   <rect x=\"38.602344\" y=\"60.080234\" width=\"526.597656\" height=\"135.397422\"/>\n",
       "  </clipPath>\n",
       " </defs>\n",
       "</svg>"
      ],
      "text/plain": [
       "<IPython.core.display.SVG object>"
      ]
     },
     "metadata": {},
     "output_type": "display_data"
    }
   ],
   "source": [
    "display(SVG(figure_svg(report)))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "76b4b757eb34",
   "metadata": {},
   "source": [
    "**Figure 7.L1:** Calculated chapter experiment. Each panel labels its input and output units; interpret it under the assumptions printed in the report."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b6fbdf5f7ea6",
   "metadata": {},
   "source": [
    "## Change the assumption\n",
    "\n",
    "A planning result can be internally exact and operationally wrong because its state model omits a precondition. If release needs current approval, ready must encode that approval or the action set must check it. Reward optimism cannot create a missing permission.\n",
    "\n",
    "Another failure comes from treating a finite model policy as robust to unknown transition changes. The Bellman recursion uses the supplied kernel at every stage. A tool outage, stale memory, or different counterparty changes that kernel and can reverse action rankings. Model auditing belongs beside planning, not after an inconvenient outcome.\n",
    "\n",
    "Finally, terminal values can hide unsupported future optimism. With a short horizon, a large terminal estimate may dominate the apparent plan. Report whether terminal continuation is zero, a measured estimate, or a judgment, and inspect sensitivity before acting."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 4,
   "id": "5a62b3250dc1",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Chapter 7: finite-horizon-planning\n",
      "When should a controller accept an immediate payoff instead of preparing a better future?\n",
      "Evidence: constructed changed-assumption example\n",
      "\n",
      "Calculated quantities:\n",
      "{\n",
      "  \"start_value\": 2.0,\n",
      "  \"policy_by_remaining_steps\": [\n",
      "    {\n",
      "      \"done\": \"stop\",\n",
      "      \"draft\": \"cash\",\n",
      "      \"ready\": \"release\"\n",
      "    }\n",
      "  ],\n",
      "  \"values_by_remaining_steps\": [\n",
      "    {\n",
      "      \"done\": 0.0,\n",
      "      \"draft\": 0.0,\n",
      "      \"ready\": 0.0\n",
      "    },\n",
      "    {\n",
      "      \"done\": 0.0,\n",
      "      \"draft\": 2.0,\n",
      "      \"ready\": 6.0\n",
      "    }\n",
      "  ]\n",
      "}\n",
      "\n",
      "Interpretation:\n",
      "Backward induction compares the immediate reward with the discounted value of the next state at each remaining horizon.\n",
      "\n",
      "Assumptions:\n",
      "- Finite fully observed state model.\n",
      "- Terminal values and discount are declared.\n",
      "- Stopping is represented by an explicit action and terminal state.\n",
      "\n",
      "Limitations:\n",
      "- The computed policy is optimal only inside the supplied finite model.\n",
      "\n",
      "Execution: completed locally; constructed inputs are not deployment measurements.\n"
     ]
    },
    {
     "data": {
      "image/svg+xml": [
       "<svg xmlns:xlink=\"http://www.w3.org/1999/xlink\" xmlns=\"http://www.w3.org/2000/svg\" width=\"576pt\" height=\"237.6pt\" viewBox=\"0 0 576 237.6\" version=\"1.1\"><title>Calculated chapter experiment</title><desc>Labeled plot of the explicitly supplied chapter inputs. See the adjacent explanation for assumptions.</desc>\n",
       " <metadata>\n",
       "  <rdf:RDF xmlns:dc=\"http://purl.org/dc/elements/1.1/\" xmlns:cc=\"http://creativecommons.org/ns#\" xmlns:rdf=\"http://www.w3.org/1999/02/22-rdf-syntax-ns#\">\n",
       "   <cc:Work>\n",
       "    <dc:type rdf:resource=\"http://purl.org/dc/dcmitype/StillImage\"/>\n",
       "    <dc:format>image/svg+xml</dc:format>\n",
       "    <dc:creator>\n",
       "     <cc:Agent>\n",
       "      <dc:title>Mathematics of AI Agents Laboratory</dc:title>\n",
       "     </cc:Agent>\n",
       "    </dc:creator>\n",
       "   </cc:Work>\n",
       "  </rdf:RDF>\n",
       " </metadata>\n",
       " <defs>\n",
       "  <style type=\"text/css\">*{stroke-linejoin: round; stroke-linecap: butt}</style>\n",
       " </defs>\n",
       " <g id=\"figure_1\">\n",
       "  <g id=\"patch_1\">\n",
       "   <path d=\"M 0 237.6  L 576 237.6  L 576 0  L 0 0  z \" style=\"fill: #ffffff\"/>\n",
       "  </g>\n",
       "  <g id=\"axes_1\">\n",
       "   <g id=\"patch_2\">\n",
       "    <path d=\"M 54.442344 195.477656  L 565.2 195.477656  L 565.2 60.080234  L 54.442344 60.080234  z \" style=\"fill: #ffffff\"/>\n",
       "   </g>\n",
       "   <g id=\"matplotlib.axis_1\">\n",
       "    <g id=\"xtick_1\">\n",
       "     <g id=\"line2d_1\">\n",
       "      <defs>\n",
       "       <path id=\"ma2b3bd7214\" d=\"M 0 0  L 0 3.5  \" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </defs>\n",
       "      <g>\n",
       "       <use xlink:href=\"#ma2b3bd7214\" x=\"77.658601\" y=\"195.477656\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_1\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"77.658601\" y=\"210.075312\" transform=\"rotate(-0 77.658601 210.075312)\">0</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"xtick_2\">\n",
       "     <g id=\"line2d_2\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#ma2b3bd7214\" x=\"541.983743\" y=\"195.477656\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_2\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"541.983743\" y=\"210.075312\" transform=\"rotate(-0 541.983743 210.075312)\">1</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"text_3\">\n",
       "     <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"309.821172\" y=\"224.076094\" transform=\"rotate(-0 309.821172 224.076094)\">remaining decisions</text>\n",
       "    </g>\n",
       "   </g>\n",
       "   <g id=\"matplotlib.axis_2\">\n",
       "    <g id=\"ytick_1\">\n",
       "     <g id=\"line2d_3\">\n",
       "      <path d=\"M 54.442344 189.323228  L 565.2 189.323228  \" clip-path=\"url(#p66c58ab04d)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_4\">\n",
       "      <defs>\n",
       "       <path id=\"m855850f29d\" d=\"M 0 0  L -3.5 0  \" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </defs>\n",
       "      <g>\n",
       "       <use xlink:href=\"#m855850f29d\" x=\"54.442344\" y=\"189.323228\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_4\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"47.442344\" y=\"193.122056\" transform=\"rotate(-0 47.442344 193.122056)\">0.0</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_2\">\n",
       "     <g id=\"line2d_5\">\n",
       "      <path d=\"M 54.442344 158.551087  L 565.2 158.551087  \" clip-path=\"url(#p66c58ab04d)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_6\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m855850f29d\" x=\"54.442344\" y=\"158.551087\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_5\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"47.442344\" y=\"162.349915\" transform=\"rotate(-0 47.442344 162.349915)\">0.5</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_3\">\n",
       "     <g id=\"line2d_7\">\n",
       "      <path d=\"M 54.442344 127.778945  L 565.2 127.778945  \" clip-path=\"url(#p66c58ab04d)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_8\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m855850f29d\" x=\"54.442344\" y=\"127.778945\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_6\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"47.442344\" y=\"131.577773\" transform=\"rotate(-0 47.442344 131.577773)\">1.0</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_4\">\n",
       "     <g id=\"line2d_9\">\n",
       "      <path d=\"M 54.442344 97.006804  L 565.2 97.006804  \" clip-path=\"url(#p66c58ab04d)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_10\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m855850f29d\" x=\"54.442344\" y=\"97.006804\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_7\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"47.442344\" y=\"100.805632\" transform=\"rotate(-0 47.442344 100.805632)\">1.5</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"ytick_5\">\n",
       "     <g id=\"line2d_11\">\n",
       "      <path d=\"M 54.442344 66.234663  L 565.2 66.234663  \" clip-path=\"url(#p66c58ab04d)\" style=\"fill: none; stroke: #d4d8da; stroke-width: 0.6; stroke-linecap: square\"/>\n",
       "     </g>\n",
       "     <g id=\"line2d_12\">\n",
       "      <g>\n",
       "       <use xlink:href=\"#m855850f29d\" x=\"54.442344\" y=\"66.234663\" style=\"stroke: #000000; stroke-width: 0.8\"/>\n",
       "      </g>\n",
       "     </g>\n",
       "     <g id=\"text_8\">\n",
       "      <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: end\" x=\"47.442344\" y=\"70.033491\" transform=\"rotate(-0 47.442344 70.033491)\">2.0</text>\n",
       "     </g>\n",
       "    </g>\n",
       "    <g id=\"text_9\">\n",
       "     <text style=\"font-size: 10px; font-family: 'DejaVu Sans'; text-anchor: middle\" x=\"25.136875\" y=\"127.778945\" transform=\"rotate(-90 25.136875 127.778945)\">discounted reward</text>\n",
       "    </g>\n",
       "   </g>\n",
       "   <g id=\"line2d_13\">\n",
       "    <path d=\"M 77.658601 189.323228  L 541.983743 66.234663  \" clip-path=\"url(#p66c58ab04d)\" style=\"fill: none; stroke: #31586b; stroke-width: 1.8; stroke-linecap: square\"/>\n",
       "    <defs>\n",
       "     <path id=\"mece881f7c5\" d=\"M 0 2  C 0.530406 2 1.03916 1.789267 1.414214 1.414214  C 1.789267 1.03916 2 0.530406 2 0  C 2 -0.530406 1.789267 -1.03916 1.414214 -1.414214  C 1.03916 -1.789267 0.530406 -2 0 -2  C -0.530406 -2 -1.03916 -1.789267 -1.414214 -1.414214  C -1.789267 -1.03916 -2 -0.530406 -2 0  C -2 0.530406 -1.789267 1.03916 -1.414214 1.414214  C -1.03916 1.789267 -0.530406 2 0 2  z \" style=\"stroke: #31586b\"/>\n",
       "    </defs>\n",
       "    <g clip-path=\"url(#p66c58ab04d)\">\n",
       "     <use xlink:href=\"#mece881f7c5\" x=\"77.658601\" y=\"189.323228\" style=\"fill: #31586b; stroke: #31586b\"/>\n",
       "     <use xlink:href=\"#mece881f7c5\" x=\"541.983743\" y=\"66.234663\" style=\"fill: #31586b; stroke: #31586b\"/>\n",
       "    </g>\n",
       "   </g>\n",
       "   <g id=\"patch_3\">\n",
       "    <path d=\"M 54.442344 195.477656  L 54.442344 60.080234  \" style=\"fill: none; stroke: #000000; stroke-width: 0.8; stroke-linejoin: miter; stroke-linecap: square\"/>\n",
       "   </g>\n",
       "   <g id=\"patch_4\">\n",
       "    <path d=\"M 54.442344 195.477656  L 565.2 195.477656  \" style=\"fill: none; stroke: #000000; stroke-width: 0.8; stroke-linejoin: miter; stroke-linecap: square\"/>\n",
       "   </g>\n",
       "   <g id=\"text_10\">\n",
       "    <text style=\"font-size: 11px; font-family: 'DejaVu Sans'; text-anchor: start\" x=\"54.442344\" y=\"54.080234\" transform=\"rotate(-0 54.442344 54.080234)\">optimal start value</text>\n",
       "   </g>\n",
       "  </g>\n",
       "  <g id=\"text_11\">\n",
       "   <text style=\"font-size: 12px; font-family: 'DejaVu Sans'; text-anchor: start\" x=\"46.08\" y=\"13.870125\" transform=\"rotate(-0 46.08 13.870125)\">Chapter 7: finite horizon planning</text>\n",
       "  </g>\n",
       " </g>\n",
       " <defs>\n",
       "  <clipPath id=\"p66c58ab04d\">\n",
       "   <rect x=\"54.442344\" y=\"60.080234\" width=\"510.757656\" height=\"135.397422\"/>\n",
       "  </clipPath>\n",
       " </defs>\n",
       "</svg>"
      ],
      "text/plain": [
       "<IPython.core.display.SVG object>"
      ]
     },
     "metadata": {},
     "output_type": "display_data"
    }
   ],
   "source": [
    "changed_inputs = {'horizon': 1,\n",
    " 'discount': 1,\n",
    " 'start': 'draft',\n",
    " 'states': {'draft': {'actions': [{'name': 'cash',\n",
    "                                   'reward': 2,\n",
    "                                   'probabilities': [1],\n",
    "                                   'next_states': ['done']},\n",
    "                                  {'name': 'prepare',\n",
    "                                   'reward': -1,\n",
    "                                   'probabilities': [1],\n",
    "                                   'next_states': ['ready']}]},\n",
    "            'ready': {'actions': [{'name': 'release',\n",
    "                                   'reward': 6,\n",
    "                                   'probabilities': [1],\n",
    "                                   'next_states': ['done']}]},\n",
    "            'done': {'actions': [{'name': 'stop',\n",
    "                                  'reward': 0,\n",
    "                                  'probabilities': [1],\n",
    "                                  'next_states': ['done']}]}}}\n",
    "changed = analyze(chapter, changed_inputs)\n",
    "changed['evidence_kind'] = 'constructed changed-assumption example'\n",
    "print(report_text(changed))\n",
    "display(SVG(figure_svg(changed)))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "e567da8ec11f",
   "metadata": {},
   "source": [
    "**Figure 7.L2:** The changed-assumption result. Compare the printed quantities and the stated assumptions with the first run. A different input need not imply a causal effect in a deployed agent."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "ff629959cf7b",
   "metadata": {},
   "source": [
    "## Try a new case\n",
    "\n",
    "The transfer problem discounts future reward by 0.5. Taking one now gives value 1. Waiting and collecting 4 next step gives 0+0.5(4)=2, so later wins with two decisions remaining. With one remaining decision, waiting gives zero and now wins.\n",
    "\n",
    "For local use, begin with a state model small enough to inspect. Include deadline, authority, or tool status when they affect future choices. Export the policy stage corresponding to the actual remaining horizon, together with its alternatives and transition assumptions. A policy for three remaining steps is not automatically valid after one step has already been spent."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 5,
   "id": "2196f5eab132",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Chapter 7: finite-horizon-planning\n",
      "When should a controller accept an immediate payoff instead of preparing a better future?\n",
      "Evidence: constructed transfer example\n",
      "\n",
      "Calculated quantities:\n",
      "{\n",
      "  \"start_value\": 2.0,\n",
      "  \"policy_by_remaining_steps\": [\n",
      "    {\n",
      "      \"end\": \"stop\",\n",
      "      \"paid\": \"collect\",\n",
      "      \"wait\": \"now\"\n",
      "    },\n",
      "    {\n",
      "      \"end\": \"stop\",\n",
      "      \"paid\": \"collect\",\n",
      "      \"wait\": \"later\"\n",
      "    }\n",
      "  ],\n",
      "  \"values_by_remaining_steps\": [\n",
      "    {\n",
      "      \"end\": 0.0,\n",
      "      \"paid\": 0.0,\n",
      "      \"wait\": 0.0\n",
      "    },\n",
      "    {\n",
      "      \"end\": 0.0,\n",
      "      \"paid\": 4.0,\n",
      "      \"wait\": 1.0\n",
      "    },\n",
      "    {\n",
      "      \"end\": 0.0,\n",
      "      \"paid\": 4.0,\n",
      "      \"wait\": 2.0\n",
      "    }\n",
      "  ]\n",
      "}\n",
      "\n",
      "Interpretation:\n",
      "Backward induction compares the immediate reward with the discounted value of the next state at each remaining horizon.\n",
      "\n",
      "Assumptions:\n",
      "- Finite fully observed state model.\n",
      "- Terminal values and discount are declared.\n",
      "- Stopping is represented by an explicit action and terminal state.\n",
      "\n",
      "Limitations:\n",
      "- The computed policy is optimal only inside the supplied finite model.\n",
      "\n",
      "Execution: completed locally; constructed inputs are not deployment measurements.\n"
     ]
    }
   ],
   "source": [
    "transfer_inputs = {'horizon': 2,\n",
    " 'discount': 0.5,\n",
    " 'start': 'wait',\n",
    " 'states': {'wait': {'actions': [{'name': 'now',\n",
    "                                  'reward': 1,\n",
    "                                  'probabilities': [1],\n",
    "                                  'next_states': ['end']},\n",
    "                                 {'name': 'later',\n",
    "                                  'reward': 0,\n",
    "                                  'probabilities': [1],\n",
    "                                  'next_states': ['paid']}]},\n",
    "            'paid': {'actions': [{'name': 'collect',\n",
    "                                  'reward': 4,\n",
    "                                  'probabilities': [1],\n",
    "                                  'next_states': ['end']}]},\n",
    "            'end': {'actions': [{'name': 'stop',\n",
    "                                 'reward': 0,\n",
    "                                 'probabilities': [1],\n",
    "                                 'next_states': ['end']}]}}}\n",
    "transfer = analyze(chapter, transfer_inputs)\n",
    "transfer['evidence_kind'] = 'constructed transfer example'\n",
    "print(report_text(transfer))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "898a9d14de2d",
   "metadata": {},
   "source": [
    "## Apply the method to your inputs\n",
    "\n",
    "The example file below has the exact input shape the method accepts. Copy it to a new file, replace its values, then point `reader_file` at your copy. Run the cell again. Supplied inputs retain their stated provenance; the program cannot establish that they are representative observations.\n",
    "\n",
    "- **horizon:** Integer decision count 1..200.\n",
    "- **discount:** Gamma in [0,1].\n",
    "- **start:** State key.\n",
    "- **states:** Object mapping states to optional terminal:number and actions:list. Action has name,reward:number,probabilities:normalized list,next_states:aligned registered state keys."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 6,
   "id": "6f9d3320c5ed",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "Chapter 7: finite-horizon-planning\n",
      "When should a controller accept an immediate payoff instead of preparing a better future?\n",
      "Evidence: supplied local inputs; provenance not independently verified\n",
      "\n",
      "Calculated quantities:\n",
      "{\n",
      "  \"start_value\": 2.0,\n",
      "  \"policy_by_remaining_steps\": [\n",
      "    {\n",
      "      \"end\": \"stop\",\n",
      "      \"paid\": \"collect\",\n",
      "      \"wait\": \"now\"\n",
      "    },\n",
      "    {\n",
      "      \"end\": \"stop\",\n",
      "      \"paid\": \"collect\",\n",
      "      \"wait\": \"later\"\n",
      "    }\n",
      "  ],\n",
      "  \"values_by_remaining_steps\": [\n",
      "    {\n",
      "      \"end\": 0.0,\n",
      "      \"paid\": 0.0,\n",
      "      \"wait\": 0.0\n",
      "    },\n",
      "    {\n",
      "      \"end\": 0.0,\n",
      "      \"paid\": 4.0,\n",
      "      \"wait\": 1.0\n",
      "    },\n",
      "    {\n",
      "      \"end\": 0.0,\n",
      "      \"paid\": 4.0,\n",
      "      \"wait\": 2.0\n",
      "    }\n",
      "  ]\n",
      "}\n",
      "\n",
      "Interpretation:\n",
      "Backward induction compares the immediate reward with the discounted value of the next state at each remaining horizon.\n",
      "\n",
      "Assumptions:\n",
      "- Finite fully observed state model.\n",
      "- Terminal values and discount are declared.\n",
      "- Stopping is represented by an explicit action and terminal state.\n",
      "\n",
      "Limitations:\n",
      "- The computed policy is optimal only inside the supplied finite model.\n",
      "\n",
      "Execution: completed locally; constructed inputs are not deployment measurements.\n"
     ]
    }
   ],
   "source": [
    "reader_file = LAB_ROOT / 'data/examples/ch07.json'\n",
    "reader_inputs = json.loads(reader_file.read_text())\n",
    "reader_report = analyze(chapter, reader_inputs)\n",
    "print(report_text(reader_report))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "184056e970b1",
   "metadata": {},
   "source": [
    "## Questions\n",
    "\n",
    "1. Compute draft value with two decisions.\n",
    "\n",
    "2. Why does horizon 3 not exceed horizon 2 here?\n",
    "\n",
    "3. What is transfer later value?\n",
    "\n",
    "Answers: [separate solutions](../solutions/ch07.md). Try the calculation before opening them."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "d85a2cdd02e7",
   "metadata": {},
   "source": [
    "## Summary\n",
    "\n",
    "Backward induction prices immediate reward together with discounted continuation. Its policy depends on state and remaining horizon, and explicit stopping prevents extra horizon from manufacturing reward. The default chooses preparation when time permits; the changed case chooses the immediate payoff; the transfer case shows discounting. Preserve the finite model boundary and terminal assumptions. This method solves a declared planning problem without estimating the accuracy of that model or authorizing its actions.\n",
    "\n",
    "Limits of this experiment:\n",
    "\n",
    "- This is a local calculation under declared inputs, not an empirical claim about a deployed agent.\n",
    "- Read the returned assumptions and limitations before applying the numerical result.\n",
    "\n",
    "The assistant skill is [`maa-07-finite-horizon-planning`](../skills/maa-07-finite-horizon-planning/SKILL.md). It uses this notebook's tested computation and input contract."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "9cc9f25ba898",
   "metadata": {},
   "source": [
    "## Equations from the chapter\n",
    "\n",
    "These are the unchanged display equations and their explanations from the canonical chapter. They are a reference for the experiment, not a claim that every equation is numerically implemented by this one method."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "a4753e4001d0",
   "metadata": {},
   "source": [
    "### Equation 7.1\n",
    "\n",
    "![Equation 7.1](../assets/math/676a39e05b5e3136feb4.svg)\n",
    "\n",
    "Equation (7.1) collapses a whole future of rewards into one number, weighting later rewards less than earlier ones.\n",
    "\n",
    "Add up every future reward, multiplying each by the discount factor raised to the number of steps away it is.\n",
    "\n",
    "LaTeX source, preserved for inspection:\n",
    "\n",
    "```latex\n",
    "\\operatorname{Ret}_t=\\sum_{k=0}^{\\infty}\\gamma^{k}r_{t+k}.\n",
    "\\tag{7.1}\n",
    "```"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "87d1b882ddfc",
   "metadata": {},
   "source": [
    "### Equation 7.2\n",
    "\n",
    "![Equation 7.2](../assets/math/29045c6784efcc8a7fd3.svg)\n",
    "\n",
    "Equation (7.2) attaches a single number to a state, saying what the agent can expect to collect from there onward under a given way of behaving.\n",
    "\n",
    "Average the return over everything that could happen from this state, under the policy the agent is following.\n",
    "\n",
    "LaTeX source, preserved for inspection:\n",
    "\n",
    "```latex\n",
    "V^{\\pi}(x)=\\operatorname{E}_{\\pi}\\!\\left[\\operatorname{Ret}_t \\mid x_t=x\\right].\n",
    "\\tag{7.2}\n",
    "```"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "db9da47f2db9",
   "metadata": {},
   "source": [
    "### Equation 7.3\n",
    "\n",
    "![Equation 7.3](../assets/math/f5b42c519e200f14b460.svg)\n",
    "\n",
    "Equation (7.3) expresses a state's value in terms of the values of the states it leads to, turning a question about an infinite future into a relationship between neighbours.\n",
    "\n",
    "Average over the actions the policy might take and the states each might lead to, taking the immediate reward plus the discounted value of where the agent arrives.\n",
    "\n",
    "LaTeX source, preserved for inspection:\n",
    "\n",
    "```latex\n",
    "\\begin{aligned}\n",
    "V^{\\pi}(x)\n",
    "&=\\sum_{a}\\pi(a\\mid x)\\sum_{x'}P(x'\\mid x,a)\\\\\n",
    "&\\quad\\cdot\\Bigl[r(x,a,x')+\\gamma V^{\\pi}(x')\\Bigr].\n",
    "\\end{aligned}\n",
    "\\tag{7.3}\n",
    "```"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "cb1e8208406a",
   "metadata": {},
   "source": [
    "### Equation 7.4\n",
    "\n",
    "![Equation 7.4](../assets/math/5bfe728026d5e6cb786c.svg)\n",
    "\n",
    "Equation (7.4) characterizes the best achievable value from every state, and the policy that achieves it falls out of the same expression.\n",
    "\n",
    "For each action, average the immediate reward plus discounted next-state value over where it might lead, then keep the largest.\n",
    "\n",
    "LaTeX source, preserved for inspection:\n",
    "\n",
    "```latex\n",
    "V^{\\star}(x)=\\max_{a}\\sum_{x'}P(x'\\mid x,a)\\Bigl[r(x,a,x')+\\gamma V^{\\star}(x')\\Bigr].\n",
    "\\tag{7.4}\n",
    "```"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "70cd30b7fd92",
   "metadata": {},
   "source": [
    "### Equation 7.5\n",
    "\n",
    "![Equation 7.5](../assets/math/bb891e918caf9b7356c6.svg)\n",
    "\n",
    "Equation (7.5) turns a Bellman residual into an upper bound on value error for declared discounted model.\n",
    "\n",
    "Divide largest one-step inconsistency by remaining contraction margin $1-\\gamma$.\n",
    "\n",
    "LaTeX source, preserved for inspection:\n",
    "\n",
    "```latex\n",
    "\\|V-V^\\star\\|_\\infty\\leq \\frac{\\|TV-V\\|_\\infty}{1-\\gamma}.\n",
    "\\tag{7.5}\n",
    "```"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Mathematics of AI Agents",
   "language": "python",
   "name": "maa-lab"
  },
  "lab_chapter": 7,
  "lab_execution": {
   "code_cells": 6,
   "created_utc": "2026-10-02T05:13:57.034636+00:00",
   "elapsed_seconds": 8.839985917089507,
   "method": "fresh process; new ipykernel InProcessKernelManager; cells submitted as Jupyter execute requests",
   "network_transport_tested": false,
   "python": "3.11.15",
   "source_sha256": "16c01e9296a1f514171238e07a47a7be3c3900c3544f917a2446aa9cf870392e"
  },
  "language_info": {
   "name": "python"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
