# Chapter 13 notebook: Statistics: From Data to Defensible Decisions

**Goal:** Plan precision and report uncertainty using the actual independent sampling units.

**Start with:** Rules 13.1.5, 13.1.8, 13.2.9. Use a calculator or the optional Python lab. Programming is optional for the workbook; the Jupyter version requires a Python 3 kernel. All lab numbers are constructed practice inputs, not observed data.

**Source:** [Finished Chapter 13](../skills/math-thumb-statistics/references/chapter.md) from *Mathematical Rules of Thumb*. Numbered rules, classifications, and the decision path come from the book. The lab, exercises, answer key, and worksheet prompts are companion additions.

## 1. Frame your decision

Write a question in which an answer would change something you do. Gather: Estimand and target population; sampling or assignment design; independent units; variation; precision target; multiplicity and stopping plan.

- My question and intended decision:
- Known inputs and units:
- Required accuracy or threshold:
- What I expect before calculating:
- What I still need to find out:

Use the lab as a worked starting point if you do not yet have your own problem. For notation or prerequisites, ask the chapter skill to explain only the concept blocking the next step.

## 2. Choose a route

1. **Name the target.** Is it a finite-population total, a future-population mean, a treatment contrast, a proportion, a correlation, a predictive loss, or a model coefficient? Do not calculate until units and target population are explicit.
2. **Map the observation structure.** Mark clusters, repeated measurements, time order, strata, weights, blocks, pairs, censoring, and finite sampling fractions. Decide what the independent units really are.
3. **Choose the scale and summary.** Use arithmetic methods for additive variation, logarithms or geometric means for multiplicative variation, and resistant summaries when tails can dominate.
4. **Plan precision.** Select a meaningful margin or effect, estimate nuisance variation conservatively, apply the correct allocation and design effect, add attrition, and round upward.
5. **Choose inference from the design.** Use \(t\) when scale is estimated, Wilson or plus-four for ordinary proportions, exact or simulated table methods when sparse, and sequential boundaries for repeated looks.
6. **Declare the claim family.** Decide whether the goal is one protected conclusion, familywise control, or discovery screening. Set equivalence margins and multiplicity rules before outcomes are inspected.
7. **Fit and compare models.** Use AIC, BIC, or cross-validation only among scientifically plausible candidates fitted on comparable data. Preserve preprocessing inside resampling.
8. **Stress-test the answer.** Examine VIF, leverage, Cook’s distance, events per parameter, bootstrap convergence, tail rank, and sensitivity to flagged observations or design assumptions.
9. **Report the decision range.** Give the estimate, interval, assumptions, denominator, design, adjustment family, and the values within the uncertainty range that would reverse the decision.

**Chapter-specific stop check:** Plan for the independent units and claim family. Screening thresholds do not automatically justify deletion, equivalence, model validity, or a universal sample-size guarantee.

## 3. Work the lab

Plan estimation of a population mean with anticipated standard deviation 12 units and a 95% normal-approximation margin of 3 units. n≈(1.96 × 12/3)^2=61.4656, rounded upward to 62 independent observations. If groups contain five observations and intraclass correlation is 0.1, the planning design effect is 1+(5-1) × 0.1=1.4. Inflate the unrounded requirement, then round to complete groups: ceil(61.4656 × 1.4/5) × 5=90 observations in 18 groups. This is a planning shortcut; final inference must account for the group design.

Predict the sign and scale before running the code. Then change one input and explain why the result moves. The code checks the constructed example; it does not prove the rule for every possible input. Code assertions may describe the example's chosen regime, so inspect them before changing that regime.

```python
import math
sd, margin, z = 12.0, 3.0, 1.96
group_size, intraclass_correlation = 5, 0.1
assert sd > 0 and margin > 0 and group_size >= 1 and 0 <= intraclass_correlation <= 1
base = (z*sd/margin)**2
design_effect = 1+(group_size-1)*intraclass_correlation
groups = math.ceil(base*design_effect/group_size)
print("Independent planning n:", math.ceil(base))
print(f"Design effect: {design_effect:.3f}; clustered plan: {groups} groups, {groups*group_size} observations")
assert groups*group_size >= base*design_effect
```

**My prediction, observed result, and explanation:**

_Record your work here._

## 4. Practise without the answers

### Exercise 1

No events are observed in 200 independent trials. Give the rule-of-three upper limit.

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

### Exercise 2

A two-sided confidence interval for an effect is [-1,4]. Does failure to reject zero show equivalence within [-0.5,0.5]?

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

### Exercise 3

A study records 100 repeated measurements on each of ten people and treats n=1000 as independent. What must change?

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

**Coaching prompt:** “Use the Chapter 13 skill to help me with Exercise 2. Ask for my attempt, give one useful hint if I need it, and help me check the assumptions before showing the answer.”

## 5. Answer key and reasoning

Read this after attempting the exercises, or use it immediately if you prefer a complete walkthrough. An answer is complete only when its assumptions and stopping point are clear.

### Answer 1

An approximate one-sided 95% upper limit is 3/200=0.015, or 1.5%. Zero observed events does not establish zero risk.

### Answer 2

No. The interval contains effects far outside the equivalence margins. Equivalence needs a prespecified margin and an appropriate procedure.

### Answer 3

Identify the ten people as clusters and model or otherwise account for within-person dependence. The naive root-n standard error overstates information; the raw row count is not the independent sample size.

## 6. Build the complete chapter toolkit

The new lab samples the chapter; the following checklist covers all 49 rules. Study one thematic group at a time. A large group can take several sessions.

- **Precision Before Significance:** work with rules 13.1.1, 13.1.2, 13.1.3, 13.1.4, 13.1.5, 13.1.6, 13.1.7, 13.1.8, 13.1.9, 13.1.10, 13.1.11, 13.1.12, 13.1.13.
- **Robust Summaries and Efficient Designs:** work with rules 13.2.1, 13.2.2, 13.2.3, 13.2.4, 13.2.5, 13.2.6, 13.2.7, 13.2.8, 13.2.9, 13.2.10, 13.2.11, 13.2.12, 13.2.13, 13.2.14, 13.2.15, 13.2.16, 13.2.17, 13.2.18, 13.2.19.
- **Guardrails for Claims and Models:** work with rules 13.3.1, 13.3.2, 13.3.3, 13.3.4, 13.3.5, 13.3.6, 13.3.7, 13.3.8, 13.3.9, 13.3.10, 13.3.11, 13.3.12, 13.3.13, 13.3.14, 13.3.15, 13.3.16, 13.3.17.

For each selected rule, read its equation, explanation, and worked use in the source. Reproduce that example; change one input; then change one assumption so the rule is no longer justified. Record the result and what check catches the failure. Historical examples remain labeled and qualified as in the source.

Read each complete numbered profile in the [chapter reference](../skills/math-thumb-statistics/references/chapter.md) before applying it. The cues below abbreviate the graph metadata; they are not complete conditions. Change the status only after doing the practice described below.

| Rule | Book role | First assumptions to inspect | Practice status |
|---|---|---|---|
| 13.1.1: Standard Error of a Mean | Independent | independent observations;  finite variance | new |
| 13.1.2: Standard Error of a Sample Proportion | Independent | independent bernoulli trials;  probability away from boundaries | new |
| 13.1.3: The 95% Interval Is Roughly Plus or Minus Two SE | Independent | approximately normal estimator;  valid standard error | new |
| 13.1.4: Minimum Detectable Effect Shrinks as One Over Root N | Independent | fixed design;  fixed noise | new |
| 13.1.5: Sample Size for a Mean Margin of Error | Independent | anticipated standard deviation;  approximately normal estimator | new |
| 13.1.6: Conservative Sample Size for a Proportion | Workflow | independent bernoulli sample;  normal margin formula | new |
| 13.1.7: Two-Arm 80% Power Sample-Size Shortcut | Independent | two equal independent normal groups;  equal variance | new |
| 13.1.8: Rule of Three for Zero Observed Events | Independent | zero observed events;  independent identical trials | new |
| 13.1.9: Prefer the Wilson Interval to the Wald Interval | Workflow | independent binomial sampling | new |
| 13.1.10: Plus-Four Interval for a Proportion | Independent | independent bernoulli trials;  approximately 95 percent coverage | new |
| 13.1.11: Use t Rather Than z When Sigma Is Estimated | Workflow | normal population or adequate sample;  unknown sigma | new |
| 13.1.12: Fisher z Interval for a Correlation | Workflow | independent pairs;  approximately bivariate normal | new |
| 13.1.13: Correlation Significance Is Roughly Two Over Root N | Independent | independent pairs;  bivariate normal null model | new |
| 13.2.1: Geometric Mean for Compound Growth | Independent | positive multiplicative factors;  consistent periods | new |
| 13.2.2: Log Transform for Multiplicative Variation | Workflow | strictly positive values;  multiplicative error structure | new |
| 13.2.3: Tukey's 1.5-IQR Outlier Fences | Independent | univariate numeric data;  meaningful quartiles | new |
| 13.2.4: Robust z-Scores from the Median and MAD | Independent | roughly symmetric contaminated distribution;  nonzero mad | new |
| 13.2.5: Freedman-Diaconis Histogram Bin Width | Independent | univariate numeric data;  nonzero iqr | new |
| 13.2.6: Silverman's Rule-of-Thumb KDE Bandwidth | Independent | one dimensional continuous data;  gaussian kernel | new |
| 13.2.7: AR(1) Effective Sample Size | Independent | long stationary ar1 series;  positive or moderate autocorrelation | new |
| 13.2.8: Kish Effective Sample Size for Unequal Weights | Independent | positive weights;  independent homoscedastic observations | new |
| 13.2.9: Cluster Sampling Design Effect | Workflow | equal cluster size;  exchangeable intraclass correlation | new |
| 13.2.10: Finite-Population Correction After a Noticeable Sampling Fraction | Workflow | simple random sampling without replacement;  known population size | new |
| 13.2.11: Measurement Error Attenuates Correlation | Independent | classical independent measurement error;  known reliabilities | new |
| 13.2.12: Pair When Within-Pair Correlation Is Positive | Workflow | positive within pair correlation;  valid pairing | new |
| 13.2.13: Block What You Can, Randomize What You Cannot | Workflow | known nuisance factors;  randomization feasible within blocks | new |
| 13.2.14: Use 50-50 Allocation When Per-Unit Costs and Variances Match | Workflow | equal group variances;  equal per unit costs | new |
| 13.2.15: Neyman Allocation for Unequal Stratum Variability | Independent | known stratum sizes;  known stratum standard deviations | new |
| 13.2.16: Use a Pilot to Update Variance, Not to Test the Final Claim | Workflow | pilot resembles main population;  prespecified update plan | new |
| 13.2.17: Add Replicated Center Points to Screen for Curvature | Workflow | quantitative factors;  two level factorial base | new |
| 13.2.18: Use Resolution IV or Better to Screen Main Effects | Workflow | two level factorial screen;  low order effect priority | new |
| 13.2.19: Use a Foldover to Break Key Fractional-Factorial Aliases | Workflow | two level fractional factorial;  known alias structure | new |
| 13.3.1: Pair Every Effect Estimate with Its Confidence Interval | Workflow | valid effect estimator;  valid confidence interval | new |
| 13.3.2: No Significant Difference Is Not Equivalence | Workflow | prespecified equivalence margin;  valid confidence interval | new |
| 13.3.3: Repeated Peeking Inflates False Positives | Workflow | repeated hypothesis tests;  nominal fixed sample alpha | new |
| 13.3.4: Bonferroni Familywise Error Control | Workflow | valid individual p values;  planned test family | new |
| 13.3.5: Benjamini-Hochberg False Discovery Rate Rule | Workflow | valid p values;  independence or positive dependence | new |
| 13.3.6: Expected Cell Counts Around Five for Chi-Square Approximations | Workflow | independent counts;  fixed categories | new |
| 13.3.7: Fisher's Exact Test for a Sparse 2-by-2 Table | Workflow | two by two table;  fixed margins or valid sampling model | new |
| 13.3.8: AIC Differences Under Two Are Usually Small | Workflow | same data and likelihood;  candidate models fitted | new |
| 13.3.9: BIC Charges Log n per Added Parameter | Workflow | same data and likelihood;  regular parametric models | new |
| 13.3.10: One-Standard-Error Rule for Model Selection | Workflow | comparable cross validation estimates;  ordered model complexity | new |
| 13.3.11: Variance Inflation Factor Around Five as a Warning | Workflow | fitted linear predictor matrix | new |
| 13.3.12: High-Leverage Cutoff Around Two p Over n | Workflow | fitted regression design matrix;  parameter count includes intercept | new |
| 13.3.13: Cook's Distance Screening Threshold | Workflow | fitted linear regression;  valid residual scale | new |
| 13.3.14: Ten Outcome Events per Parameter as a Starting Check | Workflow | ordinary logistic regression;  binary outcome | new |
| 13.3.15: Bootstrap Replicate Budget: Thousands, Not Dozens | Workflow | standard nonparametric bootstrap;  finite compute budget | new |
| 13.3.16: Check Bootstrap Tail Resolution by Expected Rank | Workflow | independent bootstrap replicates;  requested tail probability | new |
| 13.3.17: Jackknife Works Best for Smooth Statistics | Workflow | smooth statistical functional;  approximately iid sample | new |

For a completed row, record: **rule number / my new input / mathematical claim type / verified assumptions / calculation / check / valid use / rejected use / next step**. “Practised” means you worked an example. “Demonstrated” means you can explain a valid use, transfer it, and reject a misuse without the answer key.

## 7. Apply it to your own problem

Return to your opening question. Choose the smallest rule set that can settle it. Use the book's independent/workflow/specialized classification separately from the claim type (exact, approximate, bound, diagnostic, or heuristic).

- Selected rule number(s) and reason:
- Assumptions that hold, fail, or remain uncertain:
- Substitution with units:
- Result and error, uncertainty, or bound:
- Independent check or limiting case:
- Decision this supports:
- Stop here, do a named next calculation, or gather missing information:

## 8. Transfer and continue

Useful nearby chapters: Chapter 12: probability; Chapter 14: stochastic processes; Chapter 16: measurement. Bring the question, units, assumptions, result type, and uncertainty to the next chapter. Choose a bridge only when it supplies an operation you actually need.

**Completion check:** Explain this chapter's lab in your own words; solve one changed-input exercise; reject one invalid use; and produce a decision record for your own problem. If one check fails, revisit that part of the chapter rather than marking every rule complete.
