Mathematical Rules of Thumb, illustrated reader · Chapter 21

21Scientific Computing

Precision, Performance, and Reproducibility

5 demonstrations follow the chapter's rules. Choose a value, watch the figure and the numbers change, and check your prediction. Every choice is precomputed from the notebook calculations.

Ask the chapter skill

“Help me use Chapter 21 for my question. Choose a rule, check its assumptions, and show how the result changes when an input changes.”

Use math-thumb-scientific-computing from the companion's skill package. The demonstrations below also work on their own.

Examples use constructed inputs or the book's own values, disclosed in each panel. A picture illustrates a rule; its assumptions set its scope.

1Demonstration 1 of 5

Preserve a small logarithmic change

Why can log(1+x) vanish at x=10⁻¹⁶?

Compare mathematically equivalent evaluations after floating-point rounding.

log1p⁡(x)=log⁡(1+x) \operatorname{log1p}(x)=\log(1+x)

Small x. Binary64 arithmetic, positive x; stable function does not validate the mathematical model.

Predict first. Why can log(1+x) vanish at x=10⁻¹⁶?

Choose an example

Preserve a small logarithmic change. At x=1e-12, 1+x keeps only part of x, so log(1+x) has relative error 8.9e-05. log1p never forms 1+x, so it keeps every digit of x. Use it whenever x can be tiny, such as a daily interest rate or a small probability.
Small x: 1e-12
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Direct log(1+x)
1.00009e-12
Stable log1p(x)
1e-12

At x=1e-12, 1+x keeps only part of x, so log(1+x) has relative error 8.9e-05. log1p never forms 1+x, so it keeps every digit of x. Use it whenever x can be tiny, such as a daily interest rate or a small probability.

Use the idea

Use rule 21.1.2 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Binary64 arithmetic, positive x; stable function does not validate the mathematical model.

Check your understanding: Why can log(1+x) vanish at x=10⁻¹⁶?
The addition can round to exactly one. log1p works without first losing that small increment.

Book source: Rule 21.1.2: Use log1p and expm1 for small arguments. Demonstration C21-D01. Worked illustration.

2Demonstration 2 of 5

Bound speedup before buying more processors

What is the ceiling when s=.1?

The ideal strong-scaling curve bends toward a serial-fraction ceiling.

S(P)=1s+(1−s)/P S(P)=\frac1{s+(1-s)/P}

Serial fraction s. Fixed-work Amdahl model; communication and contention omitted. This is not measured speedup.

Predict first. What is the ceiling when s=.1?

Choose an example

Bound speedup before buying more processors. Serial fraction 0.1 limits ideal speedup to 10. Communication, bandwidth and scheduling can reduce actual speedup further. This is a bound calculation, not a benchmark. Eight processors give at most 4.71 times the one-processor speed, so shrink the serial part before buying hardware.
Serial fraction s: 0.1
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Serial fraction
0.1
Ideal eight-processor speedup
4.70588
Ideal ceiling
10

Serial fraction 0.1 limits ideal speedup to 10. Communication, bandwidth and scheduling can reduce actual speedup further. This is a bound calculation, not a benchmark. Eight processors give at most 4.71 times the one-processor speed, so shrink the serial part before buying hardware.

Use the idea

Use rule 21.3.2 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Fixed-work Amdahl model; communication and contention omitted. This is not measured speedup.

Check your understanding: What is the ceiling when s=.1?
Ten, even with infinitely many processors under the ideal model.

Book source: Rule 21.3.2: Use Amdahl's law to bound strong scaling. Demonstration C21-D02. Worked illustration.

3Demonstration 3 of 5

Balance checkpoint overhead against lost work

If failure interval quadruples, how does the optimal interval change?

Frequent checkpoints waste save time; infrequent checkpoints expose more work to loss.

L(T)≈C/T+T/(2M),T*≈2CM L(T)\approx C/T+T/(2M),\quad T_*\approx\sqrt{2CM}

Mean failure interval M (s). C=10s; small-overhead, approximate memoryless-failure model, recovery cost omitted.

Predict first. If failure interval quadruples, how does the optimal interval change?

Choose an example

Balance checkpoint overhead against lost work. The leading model C/T+T/(2M) balances checkpoint cost C=10 s with assumed mean failure interval M=10000 s. Its optimum is 447.214 s; recovery cost and nonmemoryless failures require a richer model.
Mean failure interval M (s): 10000
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Checkpoint cost (s)
10
Mean failure interval (s)
10000
Approximate optimum (s)
447.214
Approximate minimum loss fraction
0.0447214

The leading model C/T+T/(2M) balances checkpoint cost C=10 s with assumed mean failure interval M=10000 s. Its optimum is 447.214 s; recovery cost and nonmemoryless failures require a richer model.

Use the idea

Use rule 21.3.3 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

C=10s; small-overhead, approximate memoryless-failure model, recovery cost omitted.

Check your understanding: If failure interval quadruples, how does the optimal interval change?
It doubles because the leading optimum scales as √M.

Book source: Rule 21.3.3: Set checkpoint interval from checkpoint cost and failure rate. Demonstration C21-D03. Worked illustration.

4Demonstration 4 of 5

See why parallel sums disagree in the last digits

Is a parallel result wrong if it differs from the serial result in the last digits?

Add the same 100,000 numbers in different groupings and compare each total with the exactly rounded sum.

(a+b)+c≠a+(b+c)in floating point (a+b)+c\neq a+(b+c)\ \text{in floating point}

Partial sums (threads). Binary64; values span eight orders of magnitude; each chunk is summed left to right, then the chunk totals are added. Fixed seed 21.

Predict first. Is a parallel result wrong if it differs from the serial result in the last digits?

Choose an example

See why parallel sums disagree in the last digits. Splitting 100,000 numbers into 16 partial sums gives -463159.67116400891; the exact sum is -463159.67116400972. They differ by 14 units in the last place, a relative error of 2e-15. A plain serial loop is off by 50 units in the last place. Each thread count adds in a different order, and floating addition is not associative. So a parallel run can disagree with a serial run in the last digits without either one being wrong. Compare results with a tolerance, not with ==.
Partial sums (threads): 16
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Values summed
100000
Partial sums
16
Error (last-place units)
14
Serial one-sum error (last-place units)
50
Relative error
1.75945e-15

Splitting 100,000 numbers into 16 partial sums gives -463159.67116400891; the exact sum is -463159.67116400972. They differ by 14 units in the last place, a relative error of 2e-15. A plain serial loop is off by 50 units in the last place. Each thread count adds in a different order, and floating addition is not associative. So a parallel run can disagree with a serial run in the last digits without either one being wrong. Compare results with a tolerance, not with ==.

Use the idea

Use rule 21.2.1 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Binary64; values span eight orders of magnitude; each chunk is summed left to right, then the chunk totals are added. Fixed seed 21.

Check your understanding: Is a parallel result wrong if it differs from the serial result in the last digits?
No. Both are rounded answers that took different paths. Compare them with a tolerance, not with exact equality.

Book source: Rule 21.2.1: Expect parallel floating reductions to change low bits. Demonstration C21-D04. Worked illustration.

5Demonstration 5 of 5

Find out whether memory or arithmetic sets the speed limit

A kernel does 1 FLOP per byte. Will a chip with twice the arithmetic peak make it faster?

Plot the roofline for a machine with a 1000 GFLOP/s arithmetic peak and 100 GB/s memory bandwidth. The kernel dot shows its ceiling.

P≤min⁡(Ppeak,BmemI) P\le\min\left(P_{peak},\,B_{mem}I\right)

Arithmetic intensity I (FLOPs/byte). Idealized roofline; caches, latency and instruction mix omitted. A ceiling, not a measured speed.

Predict first. A kernel does 1 FLOP per byte. Will a chip with twice the arithmetic peak make it faster?

Choose an example

Find out whether memory or arithmetic sets the speed limit. At 5 FLOPs per byte the kernel can reach at most 500 GFLOP/s, 50% of the arithmetic peak. It sits left of the ridge at 10, so memory bandwidth is the limit: faster arithmetic units would not help. Reuse data (blocking, fusion) to move right. The roofline is a ceiling, not a prediction; measure the kernel against it.
Arithmetic intensity I (FLOPs/byte): 5
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Arithmetic intensity (FLOPs/byte)
5
Peak arithmetic (GFLOP/s)
1000
Memory bandwidth (GB/s)
100
Ridge point (FLOPs/byte)
10
Attainable ceiling (GFLOP/s)
500
Share of peak (%)
50
Limited by
memory bandwidth

At 5 FLOPs per byte the kernel can reach at most 500 GFLOP/s, 50% of the arithmetic peak. It sits left of the ridge at 10, so memory bandwidth is the limit: faster arithmetic units would not help. Reuse data (blocking, fusion) to move right. The roofline is a ceiling, not a prediction; measure the kernel against it.

Use the idea

Use rule 21.3.1 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Idealized roofline; caches, latency and instruction mix omitted. A ceiling, not a measured speed.

Check your understanding: A kernel does 1 FLOP per byte. Will a chip with twice the arithmetic peak make it faster?
No. At 1 FLOP/byte it is capped at 100 GFLOP/s by memory bandwidth, far below either peak. Raise its data reuse or the bandwidth.

Book source: Rule 21.3.1: Know whether a kernel is compute-bound or bandwidth-bound. Demonstration C21-D05. Worked illustration.

Bring the idea to a question of your own

Choose the relationship that answers your question, check its conditions, and compare the result with the accuracy or decision threshold you need.

The chapter skill can adapt these calculations to your inputs. It should name the assumptions, explain what the result supports, and say what still needs evidence. The chapter workbook adds a lab and three exercises with answers.