Mathematical Rules of Thumb, illustrated reader · Chapter 13

13Statistics

From Data to Defensible Decisions

4 demonstrations follow the chapter's rules. Choose a value, watch the figure and the numbers change, and check your prediction. Every choice is precomputed from the notebook calculations.

Ask the chapter skill

“Help me use Chapter 13 for my question. Choose a rule, check its assumptions, and show how the result changes when an input changes.”

Use math-thumb-statistics from the companion's skill package. The demonstrations below also work on their own.

Examples use constructed inputs or the book's own values, disclosed in each panel. A picture illustrates a rule; its assumptions set its scope.

1Demonstration 1 of 4

Count independent information rather than readings

Does collecting five readings on each unit equal five independent units?

The same nominal sample count can buy less precision when observations arrive in correlated clusters.

D=1+(m−1)ρ,neff≈n/D D=1+(m-1)\rho,\quad n_{eff}\approx n/D

Within-cluster correlation ρ. Cluster size m=5, stipulated σ=12, planning z=1.96, margin 3; equal clusters and an adequate model.

Predict first. Does collecting five readings on each unit equal five independent units?

Choose an example

Count independent information rather than readings. With clusters of five and within-cluster correlation ρ=0.1, the design effect (how much clustering inflates variance) is 1.4. The unrounded target 61.4656 times 1.4 is 86.0518, rounded up to whole clusters of five: 90. The displayed planning formula assumes equal clusters and a specified variance; repeated measurements do not automatically count as independent observations.
Within-cluster correlation ρ: 0.1
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Design effect
1.4
Independent target n (unrounded)
61.4656
Target × design effect
86.0518
Clustered observations needed
90

With clusters of five and within-cluster correlation ρ=0.1, the design effect (how much clustering inflates variance) is 1.4. The unrounded target 61.4656 times 1.4 is 86.0518, rounded up to whole clusters of five: 90. The displayed planning formula assumes equal clusters and a specified variance; repeated measurements do not automatically count as independent observations.

Use the idea

Use rule 13.2.9 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Cluster size m=5, stipulated σ=12, planning z=1.96, margin 3; equal clusters and an adequate model.

Check your understanding: Does collecting five readings on each unit equal five independent units?
Only under an independence model. Positive within-unit correlation reduces effective information.

Book source: Rule 13.2.9: Cluster Sampling Design Effect. Demonstration C13-D01. Worked illustration.

2Demonstration 2 of 4

Put an upper bound on zero observed events

Do 200 zero-event trials establish risk below 1% at this confidence level?

Zero observed events leaves a nonzero upper confidence bound. Compare the exact binomial expression with the shortcut.

pU=1−0.051/n≈3/n p_U=1-0.05^{1/n}\approx3/n

Independent trials n. One-sided 95% binomial upper bound with independent identical trials and zero events.

Predict first. Do 200 zero-event trials establish risk below 1% at this confidence level?

Choose an example

Put an upper bound on zero observed events. With zero events in 200 independent identical trials, the exact one-sided 95% upper bound is 0.014867. Rule 3/n is a close approximation, not a probability of safety.
Independent trials n: 200
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Zero-event trials
200
Exact upper bound
0.014867
Rule-of-three bound
0.015

With zero events in 200 independent identical trials, the exact one-sided 95% upper bound is 0.014867. Rule 3/n is a close approximation, not a probability of safety.

Use the idea

Use rule 13.1.8 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

One-sided 95% binomial upper bound with independent identical trials and zero events.

Check your understanding: Do 200 zero-event trials establish risk below 1% at this confidence level?
No. The exact upper bound is about 1.4867%, and the rule of three gives 1.5%.

Book source: Rule 13.1.8: Rule of Three for Zero Observed Events. Demonstration C13-D02. Worked illustration.

3Demonstration 3 of 4

Compare interval behavior near zero and one

What is wrong with the Wald interval at zero successes?

The Wilson interval remains nondegenerate at the extremes where the simple Wald expression can fail.

Wilson center=p̂+z2/(2n)1+z2/n \text{Wilson center}=\frac{\hat p+z^2/(2n)}{1+z^2/n}

Successes k out of 20. Independent identical binomial trials, z=1.96 and n=20; Wilson is approximate, not an exact-coverage guarantee.

Predict first. What is wrong with the Wald interval at zero successes?

Choose an example

Compare interval behavior near zero and one. For 2/20 successes, Wilson gives [0.0278659,0.301038]. Wald gives [-0.0315,0.231], with an impossible negative lower end.
Successes k out of 20: 2
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Successes
2
Trials
20
Wilson lower
0.0278659
Wilson upper
0.301038

For 2/20 successes, Wilson gives [0.0278659,0.301038]. Wald gives [-0.0315,0.231], with an impossible negative lower end.

Use the idea

Use rule 13.1.9 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Independent identical binomial trials, z=1.96 and n=20; Wilson is approximate, not an exact-coverage guarantee.

Check your understanding: What is wrong with the Wald interval at zero successes?
Its estimated standard error is zero, causing a zero-width interval even though substantial probability uncertainty remains.

Book source: Rule 13.1.9: Prefer the Wilson Interval to the Wald Interval. Demonstration C13-D03. Worked illustration.

4Demonstration 4 of 4

Count the cost of peeking at p-values

If you check results 20 times and stop at the first p<.05, is your error rate still 5%?

Simulate experiments with no real effect. Stop as soon as any interim look shows p<.05 and count how often that false alarm happens.

P(any look has |Z|>1.96∣no effect)>.05 P(\text{any look has }|Z|>1.96\mid\text{no effect})>.05

Equally spaced looks. 4000 null experiments of 1000 normal observations, seed 1333, two-sided z test at each look, no correction.

Predict first. If you check results 20 times and stop at the first p<.05, is your error rate still 5%?

Choose an example

Count the cost of peeking at p-values. With no true effect and 5 equally spaced looks, stopping at the first p<.05 declares a false discovery in 14.2% of 4000 simulated experiments. That is about 2.8 times the promised 5%. Each extra look is another chance for noise to cross the line. Fix the look schedule in advance or use a sequential correction.
Equally spaced looks: 5
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Looks
5
Simulated false-positive rate
0.14225
Nominal rate
0.05
Simulated null experiments
4000
Seed
1333

With no true effect and 5 equally spaced looks, stopping at the first p<.05 declares a false discovery in 14.2% of 4000 simulated experiments. That is about 2.8 times the promised 5%. Each extra look is another chance for noise to cross the line. Fix the look schedule in advance or use a sequential correction.

Use the idea

Use rule 13.3.3 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

4000 null experiments of 1000 normal observations, seed 1333, two-sided z test at each look, no correction.

Check your understanding: If you check results 20 times and stop at the first p<.05, is your error rate still 5%?
No. In this simulation it is about 24%, nearly five times the promised rate.

Book source: Rule 13.3.3: Repeated Peeking Inflates False Positives. Demonstration C13-D04. Worked illustration.

Bring the idea to a question of your own

Choose the relationship that answers your question, check its conditions, and compare the result with the accuracy or decision threshold you need.

The chapter skill can adapt these calculations to your inputs. It should name the assumptions, explain what the result supports, and say what still needs evidence. The chapter workbook adds a lab and three exercises with answers.