Mathematical Rules of Thumb, illustrated reader · Chapter 7

07Linear Algebra

Choose Stable Matrix Methods Before Computing

5 demonstrations follow the chapter's rules. Choose a value, watch the figure and the numbers change, and check your prediction. Every choice is precomputed from the notebook calculations.

Ask the chapter skill

“Help me use Chapter 7 for my question. Choose a rule, check its assumptions, and show how the result changes when an input changes.”

Use math-thumb-linear-algebra from the companion's skill package. The demonstrations below also work on their own.

Examples use constructed inputs or the book's own values, disclosed in each panel. A picture illustrates a rule; its assumptions set its scope.

1Demonstration 1 of 5

Solve a sensitive system and inspect both errors

Does a small residual guarantee accuracy relative to unperturbed inputs?

The condition number κ measures how much the matrix can magnify a small input error. Nudge b by one part in 10⁸ along the direction the matrix shrinks most, and the solution moves κ times as much. The residual still looks perfect, because it only checks the equation you actually solved.

A=diag⁡(1,1/κ),Ax=b,δb=(0,10−8) A=\operatorname{diag}(1,1/\kappa),\quad A x=b,\quad\delta b=(0,10^{-8})

Condition number κ. Binary64 solve of a constructed diagonal matrix. Worst-case amplification depends on perturbation direction.

Predict first. Does a small residual guarantee accuracy relative to unperturbed inputs?

Choose an example

Solve a sensitive system and inspect both errors. A=diag(1,1/κ), b=(1,0), δb=(0,10⁻⁸). The actual solve amplifies this chosen perturbation by κ=10000; its small residual does not undo sensitivity to input changes.
Condition number κ: 10000
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Condition number
10000
Relative input perturbation
1e-08
Relative solution change
0.0001
Perturbed solve residual
0

A=diag(1,1/κ), b=(1,0), δb=(0,10⁻⁸). The actual solve amplifies this chosen perturbation by κ=10000; its small residual does not undo sensitivity to input changes.

Use the idea

Use rule 7.1.1 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Binary64 solve of a constructed diagonal matrix. Worst-case amplification depends on perturbation direction.

Check your understanding: Does a small residual guarantee accuracy relative to unperturbed inputs?
No. It checks the equation actually solved. Conditioning determines how much altered inputs can change its solution.

Book source: Rule 7.1.1: Estimate digits lost from the condition number. Demonstration C07-D01. Worked illustration.

2Demonstration 2 of 5

Fit data and read the residual pattern

Can a better solver fix a model that omitted the intercept?

Compare the same constructed data with and without an intercept. The residual panel reveals model structure left unexplained.

minβ∥Aβ−b∥2,bi=2+3xi+0.1sin⁡(9xi) \min_\beta\|A\beta-b\|_2,\quad b_i=2+3x_i+0.1\sin(9x_i)

Intercept in model. 101 points on [0,1], stable NumPy least squares. Data are deterministic teaching values, not a sample for statistical inference.

Predict first. Can a better solver fix a model that omitted the intercept?

Choose an example

Fit data and read the residual pattern. The data are 2+3x+.1 sin(9x), not measurements. With an intercept the line recovers 2+3x, and the leftover wiggle of RMS size about .07 is the sin term no straight line can capture.
Intercept in model: include
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Fitted intercept
2.02052
Fitted slope
3.00138
Residual RMS
0.0686635

The data are 2+3x+.1 sin(9x), not measurements. With an intercept the line recovers 2+3x, and the leftover wiggle of RMS size about .07 is the sin term no straight line can capture.

Use the idea

Use rule 7.2.6 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

101 points on [0,1], stable NumPy least squares. Data are deterministic teaching values, not a sample for statistical inference.

Check your understanding: Can a better solver fix a model that omitted the intercept?
No. Solver stability and model adequacy are different issues; the omitted intercept changes the fitting problem.

Book source: Rule 7.2.6: Prefer orthogonal transformations for numerical stability. Demonstration C07-D02. Worked illustration.

3Demonstration 3 of 5

Reconstruct a matrix at different ranks

What spectral error remains at rank two?

Compute an SVD and reconstruct the truncated matrix. Compare its actual spectral error with the next singular value.

∥A−Ak∥2=σk+1 \|A-A_k\|_2=\sigma_{k+1}

Retained rank k. Constructed diagonal matrix with singular values 10,3,1,.1; best approximation in the spectral norm.

Predict first. What spectral error remains at rank two?

Choose an example

Reconstruct a matrix at different ranks. Keeping rank 2 gives spectral error 1, equal to the next singular value in this actual SVD reconstruction. Compression is justified by an error budget, not rank alone.
Retained rank k: 2
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Retained rank
2
Actual spectral error
1
Next singular value
1

Keeping rank 2 gives spectral error 1, equal to the next singular value in this actual SVD reconstruction. Compression is justified by an error budget, not rank alone.

Use the idea

Use rule 7.3.5 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Constructed diagonal matrix with singular values 10,3,1,.1; best approximation in the spectral norm.

Check your understanding: What spectral error remains at rank two?
The next singular value is 1, which equals the spectral norm of the omitted part.

Book source: Rule 7.3.5: Use the next singular value to price a low-rank approximation. Demonstration C07-D03. Worked illustration.

4Demonstration 4 of 5

Watch the eigenvalue ratio set the speed of power iteration

How many steps does ratio 0.9 need to reach 10⁻⁶?

Power iteration multiplies a vector by the matrix again and again. The top eigenvector wins, but only as fast as the runner-up eigenvalue fades. A ratio near one means a near tie and a long wait.

errork≈C|λ2/λ1|k \text{error}_k\approx C\,|\lambda_2/\lambda_1|^k

Ratio |λ₂/λ₁|. Constructed matrix diag(1,ratio,0.3), start (1,1,1)/√3, 60 normalized steps, target error 10⁻⁶.

Predict first. How many steps does ratio 0.9 need to reach 10⁻⁶?

Choose an example

Watch the eigenvalue ratio set the speed of power iteration. Matrix diag(1,0.9,0.3), start (1,1,1)/√3. Each step multiplies the error by about 0.9. After 60 steps the error is still 0.0018. A ratio this close to one needs about 132 steps, so use a shift or a different method.
Ratio |λ₂/λ₁|: 0.9
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Eigenvalue ratio |λ₂/λ₁|
0.9
Error after 60 steps
0.00179701
Steps to reach 10⁻⁶
More than 60

Matrix diag(1,0.9,0.3), start (1,1,1)/√3. Each step multiplies the error by about 0.9. After 60 steps the error is still 0.0018. A ratio this close to one needs about 132 steps, so use a shift or a different method.

Use the idea

Use rule 7.3.2 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Constructed matrix diag(1,ratio,0.3), start (1,1,1)/√3, 60 normalized steps, target error 10⁻⁶.

Check your understanding: How many steps does ratio 0.9 need to reach 10⁻⁶?
About 132. Each step only removes 10% of the error, so 60 steps leave it near 0.0018.

Book source: Rule 7.3.2: Power iteration speed is set by the eigenvalue ratio. Demonstration C07-D04. Worked illustration.

5Demonstration 5 of 5

See the normal equations square the condition number

With κ=10⁶, about how many correct digits can the normal equations keep?

Solve the same exact least-squares problem two ways. QR works with A itself; the normal equations form AᵀA first, which squares the condition number and doubles the digits lost.

κ(A𝖳A)=κ(A)2 \kappa(A^{\mathsf T}A)=\kappa(A)^2

Condition number κ of A. Constructed 50×2 matrix with chosen κ, seed 705, exact fit with true answer (1,1); binary64 arithmetic.

Predict first. With κ=10⁶, about how many correct digits can the normal equations keep?

Choose an example

See the normal equations square the condition number. A is a constructed 50×2 matrix with κ=10000 and an exact fit, so the true answer is (1,1). QR works with A directly and loses about log10 κ=4 digits; its error is 1e-13. The normal equations solve with AᵀA, whose condition number is κ²=1e+08. Their error is 1.9e-09, about 1.9e+04 times worse. Squaring κ doubles the digits lost.
Condition number κ of A: 10000
Constructed teaching inputs; calculations executed locally. Supported menu choices are precomputed.

Calculated values

Condition number of A
10000
Condition number of AᵀA
1e+08
QR relative error
1.01502e-13
Normal-equations relative error
1.94067e-09

A is a constructed 50×2 matrix with κ=10000 and an exact fit, so the true answer is (1,1). QR works with A directly and loses about log10 κ=4 digits; its error is 1e-13. The normal equations solve with AᵀA, whose condition number is κ²=1e+08. Their error is 1.9e-09, about 1.9e+04 times worse. Squaring κ doubles the digits lost.

Use the idea

Use rule 7.2.4 when its stated conditions fit. Compare the calculation with your own decision threshold; retain the relevant error or uncertainty.

Where the conclusion applies

Constructed 50×2 matrix with chosen κ, seed 705, exact fit with true answer (1,1); binary64 arithmetic.

Check your understanding: With κ=10⁶, about how many correct digits can the normal equations keep?
About 4. κ²=10¹² eats 12 of the roughly 16 digits, while QR still keeps about 10.

Book source: Rule 7.2.4: Prefer QR to normal equations for accurate least squares. Demonstration C07-D05. Worked illustration.

Bring the idea to a question of your own

Choose the relationship that answers your question, check its conditions, and compare the result with the accuracy or decision threshold you need.

The chapter skill can adapt these calculations to your inputs. It should name the assumptions, explain what the result supports, and say what still needs evidence. The chapter workbook adds a lab and three exercises with answers.