# Chapter 15 notebook: Information Theory: Measuring Information and Its Limits

**Goal:** Compute information quantities on a declared scale and identify the operational limit they express.

**Start with:** Rules 15.1.1, 15.1.2, 15.1.4. Use a calculator or the optional Python lab. Programming is optional for the workbook; the Jupyter version requires a Python 3 kernel. All lab numbers are constructed practice inputs, not observed data.

**Source:** [Finished Chapter 15](../skills/math-thumb-information-theory/references/chapter.md) from *Mathematical Rules of Thumb*. Numbered rules, classifications, and the decision path come from the book. The lab, exercises, answer key, and worksheet prompts are companion additions.

## 1. Frame your decision

Write a question in which an answer would change something you do. Gather: Probability distribution; log base and units; alphabet; conditioning or side information; channel/source assumptions; common evaluation data.

- My question and intended decision:
- Known inputs and units:
- Required accuracy or threshold:
- What I expect before calculating:
- What I still need to find out:

Use the lab as a worked starting point if you do not yet have your own problem. For notation or prerequisites, ask the chapter skill to explain only the concept blocking the next step.

## 2. Choose a route

- **Is the question about one outcome?** Use surprisal. If it asks for the average over a source, use entropy or cross-entropy instead.
- **Is the alphabet finite?** Apply the log-alphabet ceiling before accepting a larger uncertainty or mutual-information claim.
- **Is a Bernoulli event rare?** Use the small-\(p\) entropy approximation only after checking the regime and whether temporal dependence changes the entropy rate.
- **Are predictive models being compared?** Keep the tokenization, evaluation data, conditioning information, and log base fixed before comparing perplexity.
- **Is dependence the target?** Mutual information is broader than correlation, but finite-sample estimation and causal interpretation remain separate issues.
- **Has information passed through a transformation?** Draw the Markov diagram and identify any side information before invoking data processing.
- **Must an information discrepancy become an event error?** State the KL direction and units, then use Pinsker.
- **Is the question operational?** Match the exact source, channel, code, or probabilistic-data-structure assumptions before using the corresponding limit or tuning formula.

**Chapter-specific stop check:** Bits and nats are different units. Equal-looking perplexities from different tokenizations are not direct evidence of equal predictive quality.

## 3. Work the lab

For probabilities [0.5,0.25,0.25], outcome surprisals are 1,2,2 bits. Average entropy is 0.5 × 1+0.25 × 2+0.25 × 2=1.5 bits, below log2(3)≈1.585. The corresponding entropy-based effective alphabet size is 2^1.5≈2.828. Model perplexity instead exponentiates cross-entropy on the evaluation distribution, so keep the distinction visible.

Predict the sign and scale before running the code. Then change one input and explain why the result moves. The code checks the constructed example; it does not prove the rule for every possible input. Code assertions may describe the example's chosen regime, so inspect them before changing that regime.

```python
import math
probabilities = [0.5, 0.25, 0.25]
assert all(p > 0 for p in probabilities) and math.isclose(sum(probabilities), 1)
surprisals = [-math.log2(p) for p in probabilities]
entropy = sum(p*s for p, s in zip(probabilities, surprisals))
ceiling = math.log2(len(probabilities))
print("Surprisals (bits):", surprisals)
print(f"Entropy: {entropy:.6f} bits; ceiling: {ceiling:.6f} bits; effective size: {2**entropy:.6f}")
assert entropy <= ceiling+1e-14
```

**My prediction, observed result, and explanation:**

_Record your work here._

## 4. Practise without the answers

### Exercise 1

A model's cross-entropy on a fixed evaluation set is 2 bits per token. What is its perplexity?

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

### Exercise 2

Two binary variables have marginal entropies one bit each. Can their mutual information be 1.2 bits?

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

### Exercise 3

A transformation receives both X and extra information about Y. Can the simple Y→X→Z data-processing bound be assumed?

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

**Coaching prompt:** “Use the Chapter 15 skill to help me with Exercise 2. Ask for my attempt, give one useful hint if I need it, and help me check the assumptions before showing the answer.”

## 5. Answer key and reasoning

Read this after attempting the exercises, or use it immediately if you prefer a complete walkthrough. An answer is complete only when its assumptions and stopping point are clear.

### Answer 1

2^2=4. Comparisons require the same tokenization, data, conditioning, and units.

### Answer 2

No. Mutual information cannot exceed either marginal entropy, so it is at most one bit.

### Answer 3

No. The claimed Markov chain may fail because Z also receives side information. Specify the full information path before invoking the inequality.

## 6. Build the complete chapter toolkit

The new lab samples the chapter; the following checklist covers all 12 rules. Study one thematic group at a time. A large group can take several sessions.

- **A Numerical Language for Information:** work with rules 15.1.1, 15.1.2, 15.1.3, 15.1.4.
- **Dependence and Irrecoverable Loss:** work with rules 15.2.1, 15.2.2, 15.2.3.
- **Operational Limits and Design Rules:** work with rules 15.3.1, 15.3.2, 15.3.3, 15.3.4, 15.3.5.

For each selected rule, read its equation, explanation, and worked use in the source. Reproduce that example; change one input; then change one assumption so the rule is no longer justified. Record the result and what check catches the failure. Historical examples remain labeled and qualified as in the source.

Read each complete numbered profile in the [chapter reference](../skills/math-thumb-information-theory/references/chapter.md) before applying it. The cues below abbreviate the graph metadata; they are not complete conditions. Change the status only after doing the practice described below.

| Rule | Book role | First assumptions to inspect | Practice status |
|---|---|---|---|
| 15.1.1: Surprisal Is Minus Log Probability | Independent | positive modeled probability;  specified log base | new |
| 15.1.2: Entropy Cannot Exceed Log Alphabet Size | Independent | finite discrete alphabet | new |
| 15.1.3: Entropy of a Rare Bernoulli Event | Independent | small event probability;  binary event | new |
| 15.1.4: Perplexity Is Exponentiated Cross-Entropy | Independent | consistent log base;  same tokenization | new |
| 15.2.1: Mutual Information as a Dependence Screen | Independent | well defined joint distribution | new |
| 15.2.2: Processing Cannot Create Information About the Source | Independent | markov chain conditional independence | new |
| 15.2.3: Pinsker Converts KL Divergence to Probability Error | Independent | common measurable space;  natural log kl | new |
| 15.3.1: Gaussian Has Maximum Entropy at Fixed Variance | Independent | fixed finite variance;  continuous distribution | new |
| 15.3.2: Gaussian Channel Capacity Grows Logarithmically with SNR | Independent | memoryless channel;  specified noise model | new |
| 15.3.3: Binary Symmetric Channel Capacity | Independent | memoryless channel;  specified noise model | new |
| 15.3.4: Optimal Prefix Coding Is Within One Bit of Entropy | Independent | discrete memoryless source;  binary prefix code | new |
| 15.3.5: Bloom Filter Optimal Hash Count | Specialized | independent uniform hashes;  known item count | new |

For a completed row, record: **rule number / my new input / mathematical claim type / verified assumptions / calculation / check / valid use / rejected use / next step**. “Practised” means you worked an example. “Demonstrated” means you can explain a valid use, transfer it, and reject a misuse without the answer key.

## 7. Apply it to your own problem

Return to your opening question. Choose the smallest rule set that can settle it. Use the book's independent/workflow/specialized classification separately from the claim type (exact, approximate, bound, diagnostic, or heuristic).

- Selected rule number(s) and reason:
- Assumptions that hold, fail, or remain uncertain:
- Substitution with units:
- Result and error, uncertainty, or bound:
- Independent check or limiting case:
- Decision this supports:
- Stop here, do a named next calculation, or gather missing information:

## 8. Transfer and continue

Useful nearby chapters: Chapter 9: combinatorics; Chapter 12: probability; Chapter 24: signal processing. Bring the question, units, assumptions, result type, and uncertainty to the next chapter. Choose a bridge only when it supplies an operation you actually need.

**Completion check:** Explain this chapter's lab in your own words; solve one changed-input exercise; reject one invalid use; and produce a decision record for your own problem. If one check fails, revisit that part of the chapter rather than marking every rule complete.
