# Chapter 21 notebook: Scientific Computing: Precision, Performance, and Reproducibility

**Goal:** Distinguish arithmetic accuracy from speed and set a measurable computational contract.

**Start with:** Rules 21.1.2, 21.2.2, 21.3.2. Use a calculator or the optional Python lab. Programming is optional for the workbook; the Jupyter version requires a Python 3 kernel. All lab numbers are constructed practice inputs, not observed data.

**Source:** [Finished Chapter 21](../skills/math-thumb-scientific-computing/references/chapter.md) from *Mathematical Rules of Thumb*. Numbered rules, classifications, and the decision path come from the book. The lab, exercises, answer key, and worksheet prompts are companion additions.

## 1. Frame your decision

Write a question in which an answer would change something you do. Gather: Precision and data scale; permitted absolute/relative error; reproducibility contract; profile; data movement; processor count or failure spacing.

- My question and intended decision:
- Known inputs and units:
- Required accuracy or threshold:
- What I expect before calculating:
- What I still need to find out:

Use the lab as a worked starting point if you do not yet have your own problem. For notation or prerequisites, ask the chapter skill to explain only the concept blocking the next step.

## 2. Choose a route

- **Is the claimed accuracy close to machine precision?** Combine unit roundoff with conditioning and algorithmic error; do not equate `eps` with final accuracy.
- **Does an expression subtract nearly equal values?** Search for a stable equivalent such as `log1p`, `expm1`, compensated summation, or a scaled formulation.
- **Do parallel runs differ?** Quantify the discrepancy, rule out races and undefined behavior, then choose a tolerance, deterministic tree, or reproducible accumulator consistent with the contract.
- **Are tests flaky?** Derive absolute and relative tolerances from expected error and units, including explicit policies for nonfinite values and arrays.
- **Is the program too slow?** Profile the complete representative workload before changing code, then reprofile after each intervention.
- **Is a hot kernel limited by arithmetic or data motion?** Measure arithmetic intensity and hardware ceilings before choosing vectorization, blocking, fusion, or traffic reduction.
- **Will more processors help?** Measure strong scaling and use the effective serial fraction to bound the gain.
- **Can the run outlive reliable hardware?** Estimate checkpoint cost and job-level failure spacing, calculate a starting interval, and test actual restart.

**Chapter-specific stop check:** Do not promise bitwise reproducibility from a tolerance check. Avoid reading machine epsilon as guaranteed final accuracy or a theoretical ceiling as measured performance.

## 3. Work the lab

For x=1e-12, compare log(1+x) with log1p(x); forming 1+x loses relative information before the logarithm is evaluated. Separately, if 10% of work is serial, Amdahl's idealized speedup on eight processors is 1/(0.1+0.9/8)≈4.706, with an infinite-processor ceiling of 10. These are mathematical ceilings under the stated model, not measured speedups.

Predict the sign and scale before running the code. Then change one input and explain why the result moves. The code checks the constructed example; it does not prove the rule for every possible input. Code assertions may describe the example's chosen regime, so inspect them before changing that regime.

```python
import math
x = 1e-12
naive, stable = math.log(1+x), math.log1p(x)
serial_fraction, processors = 0.1, 8
assert 0 < serial_fraction <= 1 and processors >= 1
speedup = 1/(serial_fraction+(1-serial_fraction)/processors)
print(f"log(1+x): {naive:.17g}; log1p(x): {stable:.17g}")
print(f"Difference relative to log1p: {abs(naive-stable)/abs(stable):.6g}")
print(f"Ideal speedup: {speedup:.6f}; ceiling: {1/serial_fraction:.6f}")
assert speedup <= min(processors, 1/serial_fraction)+1e-14
```

**My prediction, observed result, and explanation:**

_Record your work here._

## 4. Practise without the answers

### Exercise 1

Using |a-b|<=atol+rtol × |b|, compare a=1e-10 and b=0 with atol=1e-9, rtol=1e-6.

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

### Exercise 2

If a fraction 0.5 of runtime can be made twice as fast, what is the ideal total speedup?

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

### Exercise 3

Two parallel reductions differ in their last few bits. Is that alone proof of a race?

**My approach, assumptions, calculation, and check:**

_Write your attempt here._

**Coaching prompt:** “Use the Chapter 21 skill to help me with Exercise 2. Ask for my attempt, give one useful hint if I need it, and help me check the assumptions before showing the answer.”

## 5. Answer key and reasoning

Read this after attempting the exercises, or use it immediately if you prefer a complete walkthrough. An answer is complete only when its assumptions and stopping point are clear.

### Answer 1

The difference is 1e-10 and the threshold is 1e-9, so the comparison passes. A purely relative criterion at zero would be unsuitable.

### Answer 2

1/((1-0.5)+0.5/2)=4/3≈1.333. Speeding up one component by two does not double end-to-end speed.

### Answer 3

No. Floating addition is order-dependent. Measure the discrepancy and compare with the error contract, while separately checking races and undefined behavior.

## 6. Build the complete chapter toolkit

The new lab samples the chapter; the following checklist covers all 8 rules. Study one thematic group at a time. A large group can take several sessions.

- **Spend a Finite Precision Budget Deliberately:** work with rules 21.1.1, 21.1.2.
- **Make Reproducibility an Explicit Contract:** work with rules 21.2.1, 21.2.2, 21.2.3.
- **Model Hardware Limits and Failure Costs:** work with rules 21.3.1, 21.3.2, 21.3.3.

For each selected rule, read its equation, explanation, and worked use in the source. Reproduce that example; change one input; then change one assumption so the rule is no longer justified. Record the result and what check catches the failure. Historical examples remain labeled and qualified as in the source.

Read each complete numbered profile in the [chapter reference](../skills/math-thumb-scientific-computing/references/chapter.md) before applying it. The cues below abbreviate the graph metadata; they are not complete conditions. Change the status only after doing the practice described below.

| Rule | Book role | First assumptions to inspect | Practice status |
|---|---|---|---|
| 21.1.1: Budget binary64 as about sixteen decimal digits | Independent | representative workload;  floating point execution | new |
| 21.1.2: Use log1p and expm1 for small arguments | Independent | representative workload;  floating point execution | new |
| 21.2.1: Expect parallel floating reductions to change low bits | Independent | representative workload;  floating point execution | new |
| 21.2.2: Compare floating results with absolute plus relative tolerance | Workflow | meaningful absolute scale;  meaningful relative scale | new |
| 21.2.3: Profile before optimizing scientific code | Workflow | representative workload;  representative build configuration | new |
| 21.3.1: Know whether a kernel is compute-bound or bandwidth-bound | Workflow | representative workload;  floating point execution | new |
| 21.3.2: Use Amdahl's law to bound strong scaling | Independent | representative workload;  floating point execution | new |
| 21.3.3: Set checkpoint interval from checkpoint cost and failure rate | Independent | independent exponential failures;  fixed checkpoint cost | new |

For a completed row, record: **rule number / my new input / mathematical claim type / verified assumptions / calculation / check / valid use / rejected use / next step**. “Practised” means you worked an example. “Demonstrated” means you can explain a valid use, transfer it, and reject a misuse without the answer key.

## 7. Apply it to your own problem

Return to your opening question. Choose the smallest rule set that can settle it. Use the book's independent/workflow/specialized classification separately from the claim type (exact, approximate, bound, diagnostic, or heuristic).

- Selected rule number(s) and reason:
- Assumptions that hold, fail, or remain uncertain:
- Substitution with units:
- Result and error, uncertainty, or bound:
- Independent check or limiting case:
- Decision this supports:
- Stop here, do a named next calculation, or gather missing information:

## 8. Transfer and continue

Useful nearby chapters: Chapter 7: linear algebra; Chapter 19: optimization; Chapter 20: numerical methods. Bring the question, units, assumptions, result type, and uncertainty to the next chapter. Choose a bridge only when it supplies an operation you actually need.

**Completion check:** Explain this chapter's lab in your own words; solve one changed-input exercise; reject one invalid use; and produce a decision record for your own problem. If one check fails, revisit that part of the chapter rather than marking every rule complete.
