1Demonstration 1 of 5
Preserve a small logarithmic change
Why can log(1+x) vanish at x=10⁻¹⁶?
Compare mathematically equivalent evaluations after floating-point rounding.
Small x. Binary64 arithmetic, positive x; stable function does not validate the mathematical model.
Predict first. Why can log(1+x) vanish at x=10⁻¹⁶?
Choose an example
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.
Calculated values
- Direct log(1+x)
- 1.00009e-12
- Stable log1p(x)
- 1e-12
At x=1e-12, 1+x keeps only part of x, so log(1+x) has relative error 8.9e-05. log1p never forms 1+x, so it keeps every digit of x. Use it whenever x can be tiny, such as a daily interest rate or a small probability.
Use the idea
Use rule 21.1.2 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.
Where the conclusion applies
Binary64 arithmetic, positive x; stable function does not validate the mathematical model.
Check your understanding: Why can log(1+x) vanish at x=10⁻¹⁶?
Book source: Rule 21.1.2: Use log1p and expm1 for small arguments. Demonstration C21-D01. Worked illustration.
2Demonstration 2 of 5
Bound speedup before buying more processors
What is the ceiling when s=.1?
The ideal strong-scaling curve bends toward a serial-fraction ceiling.
Serial fraction s. Fixed-work Amdahl model; communication and contention omitted. This is not measured speedup.
Predict first. What is the ceiling when s=.1?
Choose an example
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.
Calculated values
- Serial fraction
- 0.1
- Ideal eight-processor speedup
- 4.70588
- Ideal ceiling
- 10
Serial fraction 0.1 limits ideal speedup to 10. Communication, bandwidth and scheduling can reduce actual speedup further. This is a bound calculation, not a benchmark. Eight processors give at most 4.71 times the one-processor speed, so shrink the serial part before buying hardware.
Use the idea
Use rule 21.3.2 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.
Where the conclusion applies
Fixed-work Amdahl model; communication and contention omitted. This is not measured speedup.
Check your understanding: What is the ceiling when s=.1?
Book source: Rule 21.3.2: Use Amdahl's law to bound strong scaling. Demonstration C21-D02. Worked illustration.
3Demonstration 3 of 5
Balance checkpoint overhead against lost work
If failure interval quadruples, how does the optimal interval change?
Frequent checkpoints waste save time; infrequent checkpoints expose more work to loss.
Mean failure interval M (s). C=10s; small-overhead, approximate memoryless-failure model, recovery cost omitted.
Predict first. If failure interval quadruples, how does the optimal interval change?
Choose an example
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.
Calculated values
- Checkpoint cost (s)
- 10
- Mean failure interval (s)
- 10000
- Approximate optimum (s)
- 447.214
- Approximate minimum loss fraction
- 0.0447214
The leading model C/T+T/(2M) balances checkpoint cost C=10 s with assumed mean failure interval M=10000 s. Its optimum is 447.214 s; recovery cost and nonmemoryless failures require a richer model.
Use the idea
Use rule 21.3.3 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.
Where the conclusion applies
C=10s; small-overhead, approximate memoryless-failure model, recovery cost omitted.
Check your understanding: If failure interval quadruples, how does the optimal interval change?
Book source: Rule 21.3.3: Set checkpoint interval from checkpoint cost and failure rate. Demonstration C21-D03. Worked illustration.
4Demonstration 4 of 5
See why parallel sums disagree in the last digits
Is a parallel result wrong if it differs from the serial result in the last digits?
Add the same 100,000 numbers in different groupings and compare each total with the exactly rounded sum.
Partial sums (threads). Binary64; values span eight orders of magnitude; each chunk is summed left to right, then the chunk totals are added. Fixed seed 21.
Predict first. Is a parallel result wrong if it differs from the serial result in the last digits?
Choose an example
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.
Calculated values
- Values summed
- 100000
- Partial sums
- 16
- Error (last-place units)
- 14
- Serial one-sum error (last-place units)
- 50
- Relative error
- 1.75945e-15
Splitting 100,000 numbers into 16 partial sums gives -463159.67116400891; the exact sum is -463159.67116400972. They differ by 14 units in the last place, a relative error of 2e-15. A plain serial loop is off by 50 units in the last place. Each thread count adds in a different order, and floating addition is not associative. So a parallel run can disagree with a serial run in the last digits without either one being wrong. Compare results with a tolerance, not with ==.
Use the idea
Use rule 21.2.1 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.
Where the conclusion applies
Binary64; values span eight orders of magnitude; each chunk is summed left to right, then the chunk totals are added. Fixed seed 21.
Check your understanding: Is a parallel result wrong if it differs from the serial result in the last digits?
Book source: Rule 21.2.1: Expect parallel floating reductions to change low bits. Demonstration C21-D04. Worked illustration.
5Demonstration 5 of 5
Find out whether memory or arithmetic sets the speed limit
A kernel does 1 FLOP per byte. Will a chip with twice the arithmetic peak make it faster?
Plot the roofline for a machine with a 1000 GFLOP/s arithmetic peak and 100 GB/s memory bandwidth. The kernel dot shows its ceiling.
Arithmetic intensity I (FLOPs/byte). Idealized roofline; caches, latency and instruction mix omitted. A ceiling, not a measured speed.
Predict first. A kernel does 1 FLOP per byte. Will a chip with twice the arithmetic peak make it faster?
Choose an example
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.
Calculated values
- Arithmetic intensity (FLOPs/byte)
- 5
- Peak arithmetic (GFLOP/s)
- 1000
- Memory bandwidth (GB/s)
- 100
- Ridge point (FLOPs/byte)
- 10
- Attainable ceiling (GFLOP/s)
- 500
- Share of peak (%)
- 50
- Limited by
- memory bandwidth
At 5 FLOPs per byte the kernel can reach at most 500 GFLOP/s, 50% of the arithmetic peak. It sits left of the ridge at 10, so memory bandwidth is the limit: faster arithmetic units would not help. Reuse data (blocking, fusion) to move right. The roofline is a ceiling, not a prediction; measure the kernel against it.
Use the idea
Use rule 21.3.1 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.
Where the conclusion applies
Idealized roofline; caches, latency and instruction mix omitted. A ceiling, not a measured speed.
Check your understanding: A kernel does 1 FLOP per byte. Will a chip with twice the arithmetic peak make it faster?
Book source: Rule 21.3.1: Know whether a kernel is compute-bound or bandwidth-bound. Demonstration C21-D05. Worked illustration.
Bring the idea to a question of your own
Choose the relationship that answers your question, check its conditions, and compare the result with the accuracy or decision threshold you need.
The chapter skill can adapt these calculations to your inputs. It should name the assumptions, explain what the result supports, and say what still needs evidence. The chapter workbook adds a lab and three exercises with answers.