← Illustrated chapter

Chapter 1 , Algebra: See the Structure Before You Solve

Suppose one quantity rises from 10 to 20 and another rises from 100 to 110. Both have increased by 10. If subtraction is the only lens you use, the changes look identical. They are not. The first quantity doubled; the second rose by only 10 percent.

That contrast reveals algebra’s central habit: identify the structure before beginning the calculation. Algebra is often taught as a sequence of legal operations, expand, collect, divide, substitute, and solve, but experienced problem solvers look first. Is the change additive or multiplicative? Is there a small parameter that makes a local approximation trustworthy? Is a long sum controlled by its endpoints? Is a polynomial’s behavior already visible from its highest power? Will a sign check prevent an invalid transformation?

The seventeen rules in this chapter fall into three families. The first translates change and growth into the right scale. The second compresses calculations by preserving the structure that controls the answer. The third diagnoses algebraic form before a procedure is chosen. Together they support one governing habit: see the structure before you solve.

An exact answer reached through the wrong model, an unstable form, or an invalid sign operation is not rescued by the amount of algebra performed afterward. A short structural check can determine whether the next ten minutes of calculation will be useful.

1.1: Change, Growth, and Local Approximation

Algebra is most useful before the long calculation begins. The five rules in this section show how to choose the right scale for change, translate growth into time, and replace a difficult expression with a controlled local model.

1.1.1: Compare proportional change with ratios, not differences

History

Book V of Euclid’s Elements never once measures a change by subtracting one magnitude from another. Around 300 BCE in Alexandria, Euclid built an entire theory of proportion on the ordering and equality of ratios, letting him compare magnitudes without assigning them modern numerical coordinates at all. That habit survives in a plain modern complaint: two quantities can move by the same raw amount and still not have moved by comparable amounts.

Euclid did not write a percentage-change formula, and he never called ratio a diagnostic tool; the equation below is a modern operational restatement of his organizing idea. What his propositions bought later readers was a question to ask before any subtraction begins: compared to what.

The equation

For an old value y1y_1 and a new value y2y_2,

scale factor=y2y1,relative change=y2−y1y1. \text{scale factor}=\frac{y_2}{y_1}, \qquad \text{relative change}=\frac{y_2-y_1}{y_1}.

The two measurements are connected:

y2−y1y1=y2y1−1. \frac{y_2-y_1}{y_1}=\frac{y_2}{y_1}-1.

The scale factor says how many times as large as the old quantity the new one is. Relative change says how much of the old quantity was added or lost.

How to read it

A scale factor of 1.251.25 means the new value is one and a quarter times the old one; the matching relative change, 0.250.25, says a quarter of the old value was added. A number below 11 signals shrinkage: 0.80.8 means the new value is eighty percent of the old, a twenty percent drop.

Ratios strip away the unit, so the comparison survives a switch from dollars to cents or feet to meters, while a raw difference does not: five dollars more can be a rounding error or a windfall depending on what it is five more than. Near a zero baseline the fraction turns unstable or undefined, and even a well-behaved ratio never speaks to whether the change actually matters, only to its scale.

How to use it

A shopper tracking a weekend price update sees the same five-dollar increase on two items: a kitchen scale moves from 2020 to 2525, and a stand mixer moves from 100100 to 105105. Subtraction says both got five dollars more expensive and stops there. Ratios separate them:

2520=1.25,25−2020=0.25=25%; \frac{25}{20}=1.25,\qquad \frac{25-20}{20}=0.25=25\%;

105100=1.05,105−100100=0.05=5%. \frac{105}{100}=1.05,\qquad \frac{105-100}{100}=0.05=5\%.

The scale’s price rose a quarter over its old value; the mixer’s rose one twentieth. If a five percent monthly hike is the store’s normal drift but a twenty five percent hike signals something unusual, the shopper buys the scale today and can afford to wait on the mixer. The catch is the baseline itself: this comparison only works because both items had a nonzero starting price on the tag. An item marked down to nearly nothing beforehand would make the ratio swing wildly for a change of a few cents, exaggerating a trivial move. This is an Independent rule: once a meaningful nonzero baseline is available, it can answer the comparison directly across many applications.

1.1.2: Estimate doubling time from the exponential rate

History

A merchant needed to know, quickly, whether money would double within a working lifetime, not after grinding through a compound-interest table by hand. In 1494, Luca Pacioli’s Summa de arithmetica, printed in Venice, gathered exactly that kind of commercial arithmetic: compound-growth problems used by Renaissance merchants, including a shortcut traditionally associated with estimating doubling time. Speed was the point; a merchant needed a practical horizon, not an abstract rate sitting in a table.

The historical record secures the commercial context and the printed work; it would be anachronistic, though, to hand Pacioli the modern label “Rule of 70” or “Rule of 72.” Pacioli’s text supplies the commercial habit; the mnemonic names, and the derivation below, belong to arithmetic teaching that came after him.

The equation

For continuous exponential growth,

y=y0ekt. y=y_0e^{kt}.

Doubling means y=2y0y=2y_0, so

2=ekt2⇒t2=ln⁡2k≈0.693k. 2=e^{kt_2} \quad\Longrightarrow\quad t_2=\frac{\ln 2}{k}\approx\frac{0.693}{k}.

If the continuous rate is quoted as p%p\% per period, then k=p/100k=p/100 and

t2=100ln⁡2p≈69.3p≈70p periods. t_2=\frac{100\ln2}{p} \approx\frac{69.3}{p} \approx\frac{70}{p}\text{ periods}.

For discrete periodic growth,

yn=y0(1+p100)n,n2=ln⁡2ln⁡(1+p/100). y_n=y_0\left(1+\frac{p}{100}\right)^n, \qquad n_2=\frac{\ln2}{\ln(1+p/100)}.

When pp is modest, 70/p70/p or the highly divisible mental numerator 72/p72/p gives a quick estimate of this discrete doubling time.

How to read it

The rate kk must be written as a fraction of the whole, so a 5%5\% growth rate becomes 0.050.05 before it goes anywhere near the formula. Dividing 0.6930.693 (essentially ln⁡2\ln2) by that fraction turns “growth per year” into “years per doubling.” The relationship runs backward: double the growth rate and the doubling time is roughly cut in half, the way a savings balance growing twice as fast needs half as many birthdays to double. Swap in the whole-number percentage pp instead of the fraction and the same shortcut becomes 70/p70/p periods, easy to do without a calculator.

Whether the growth rate itself stays steady is not something this shortcut checks: it assumes one constant rate held the whole way, and it is silent on a rate that speeds up, slows down, or compounds only once a year rather than continuously.

How to use it

A bakery owner watching revenue grow at a steady 5%5\% a year wants to know whether sales will double before her fifteen-year lease runs out. The mental shortcut gives

t2≈705=14 years, t_2\approx\frac{70}{5}=14\text{ years},

just inside the lease. The continuous model checks that estimate:

t2=0.6930.05=13.86 years, t_2=\frac{0.693}{0.05}=13.86\text{ years},

so the shortcut is off by only about 0.140.14 year, comfortably inside the fifteen-year window. If growth instead compounds once a year rather than continuously, the exact discrete figure is

n2=ln⁡2ln⁡1.05≈14.21 years, n_2=\frac{\ln2}{\ln1.05}\approx14.21\text{ years},

and the alternate mental version, 72/5=14.472/5=14.4, brackets it just as closely. On the strength of either estimate, the owner plans to renew the lease rather than search for a bigger space now. The one thing the shortcut cannot see is a bad year: if growth stalls or reverses for even one season, the doubling clock resets from wherever revenue actually sits, and averaging a good year against a bad one is not the same as compounding through the sequence she actually lived. This is an Independent rule: under a steady-growth assumption, it directly turns a rate into a useful horizon.

1.1.3: Linearize a small binomial perturbation

History

Multiply five small errors together by hand and intuition usually fails: a chain of 2%2\% overages does not obviously compound to 10%10\%, or to anything a quick guess would land on. In the 1660s in Cambridge, Isaac Newton extended the binomial expansion beyond positive integer powers to fractional and negative exponents, turning a finite algebraic identity into a much broader infinite series and tying algebraic expressions to local approximation.

Newton built the series; the modern instruction to stop after its linear term because a perturbation is small is a later asymptotic reading, expressed in today’s notation and error language. What his expansion bought later users is a way to take just its first two terms and trust the trade: a fast estimate, plus a visible next term that says how much was left on the table.

The equation

For a small, dimensionless xx,

(1+x)a=1+ax+a(a−1)2x2+O(x3). (1+x)^a =1+ax+\frac{a(a-1)}{2}x^2+O(x^3).

Here O(x3)O(x^3) means that the remaining terms have cubic or smaller scale as xx approaches zero, with aa held fixed.

Dropping the quadratic and higher terms gives

(1+x)a≈1+ax. (1+x)^a\approx1+ax.

The first omitted contribution,

a(a−1)2x2, \frac{a(a-1)}{2}x^2,

is the quickest warning about the size and direction of the leading correction.

How to read it

Read 1+x1+x as a baseline of one multiplied by a small fractional bump every time a stage repeats, like a gear stage that speeds up its output by a hair with each mesh. Raising that repeated bump to the power aa multiplies the first-order disturbance by aa: a 2%2\% speed-up at one gear stage becomes roughly a 10%10\% speed-up once the same 2%2\% bump compounds through five stages in a row (a=5a=5). “Small” is not a formality; it means the terms being thrown away are small next to the precision the job actually needs.

Whether that discarded piece is safe to ignore is not something the shortcut checks for you: only comparing a(a−1)2x2\tfrac{a(a-1)}{2}x^2 with the accuracy the job needs will tell you that.

How to use it

A machinist is checking a five-stage gear train in which each stage speeds up the output shaft by 2%2\% over its input, and needs to know the overall speed-up before running the machine at full load. With x=0.02x=0.02 and a=5a=5,

(1.02)5≈1+5(0.02)=1.10, (1.02)^5\approx1+5(0.02)=1.10,

a projected 10%10\% speed-up through the finished train. Before signing off, the machinist checks the first ignored term:

5(4)2(0.02)2=10(0.0004)=0.004, \frac{5(4)}{2}(0.02)^2 =10(0.0004)=0.004,

so a sharper estimate is 1.1041.104, close to the exact value 1.104081.10408. Since the downstream equipment’s tolerance on final output speed is ±0.5%\pm0.5\%, a projected 10.4%10.4\% speed-up is well outside it, and the train gets flagged for a stage regrind before the machine runs a full shift rather than after hours of overspeed operation. The rule only holds because the speed-up was small and the same at every stage; a stage that slips differently from the others, or a single mesh far outside 2%2\%, would make the linear estimate misleading exactly when the train’s overall speed matters most. This is an Independent rule: once its small-perturbation condition is checked, it produces a direct estimate in many unrelated settings.

1.1.4: Replace a small exponential by one plus its exponent

History

Because Leonhard Euler folded exponentials into the same organized system as series and logarithms, what once needed a printed table can now be read off one line. His 1748 Introductio in analysin infinitorum, associated with Lausanne and Berlin, gathered functions, infinite series, logarithms, exponentials, and algebraic decompositions into a single computational language and explicitly worked out the power-series expansions behind them.

Euler used those expansions directly, but he wrote before modern convergence tests and floating-point error language existed, so calling this exact instruction, replace a small exponential by one plus its exponent, his own slogan would overstate the case. It is a modern first-pass reading of a series he built for broader purposes. What his systematizing bought later readers is this: near zero, the whole curve of the exponential collapses to a straight line cheap enough to compute by hand.

The equation

The exponential series is

ex=1+x+x22+x36+⋯. e^x=1+x+\frac{x^2}{2}+\frac{x^3}{6}+\cdots.

When xx is small and dimensionless,

ex≈1+x. e^x\approx1+x.

The first correction is x2/2x^2/2, which is positive whether xx is positive or negative.

How to read it

At zero, the exponential function equals 11 and has slope 11, so near zero its graph tracks the straight line 1+x1+x closely, the way a gently curving road looks almost straight for the first few steps onto it. A small positive xx pushes the true curve a little above that line; a small negative xx does too, because the next correction, x2/2x^2/2, is always positive whichever way xx points.

This is a local reading, not a new definition of the exponential: it only holds near zero, and it stops working once xx grows large, where the line and the curve pull apart fast.

How to use it

A nurse tracking a medication that clears the bloodstream at a steady 2%2\% per hour wants a fast read on how much remains after one hour, before consulting the full decay chart. With x=−0.02x=-0.02,

e−0.02≈1−0.02=0.98, e^{-0.02}\approx1-0.02=0.98,

so roughly 98%98\% of the dose remains. The correction term,

(−0.02)22=0.0002, \frac{(-0.02)^2}{2}=0.0002,

sharpens that to 0.98020.9802, close to the true value of about 0.9801990.980199. Both the direction (a small loss) and the scale (about two percent) check out, so the nurse can trust the bedside estimate for a routine one-hour reading. The estimate only holds because the hourly clearance rate is small and steady; over many compounded hours, or with a rate that is not genuinely constant, the gap between the straight-line estimate and the true exponential curve widens, and for a large enough negative xx the linear formula can even predict a negative concentration, which no real drug level can be. This is an Independent rule: when its small-input condition holds, it directly replaces a difficult exponential across many applications rather than serving only one fixed calculation.

Two curves meet at x=0 and y=1. The exponential bends above the straight line on either side.

Figure 1.1. For dimensionless x, the exact exponential lies above 1+x on both sides of zero. The gap grows away from zero; judge it against the required tolerance.

1.1.5: Replace log one plus x by x for small x

History

Is this Euler’s rule, or only Euler’s algebra read by later eyes? His 1748 Introductio in analysin infinitorum explicitly develops the series methods this rule depends on, working out the logarithm’s expansion term by term.

Euler wrote before today’s convergence tests and numerical-error vocabulary existed, so treating the compact modern instruction, replace log⁡(1+x)\log(1+x) by xx when xx is small, as his own phrasing would claim more than the record supports. The structure is his; the slogan is not. What that structure bought later readers is a logarithm collapsed to a single term whenever the input sits close enough to one to trust the shortcut.

The equation

For |x|<1|x|<1,

log⁡(1+x)=x−x22+x33−⋯, \log(1+x) =x-\frac{x^2}{2}+\frac{x^3}{3}-\cdots,

where log\log is the natural logarithm. At first order,

log⁡(1+x)≈x. \log(1+x)\approx x.

The leading correction is negative:

log⁡(1+x)≈x−x22. \log(1+x)\approx x-\frac{x^2}{2}.

How to read it

Think of 1+x1+x as a multiplicative factor sitting just above one, the kind of small step a reservoir’s inflow takes when it rises a few percent in a week. The logarithm turns that multiplicative step into an additive one close to the same size: a 3%3\% rise in inflow corresponds to a log change near 0.030.03. Unlike the exponential’s companion rule, this correction subtracts rather than adds, so the plain estimate always runs a touch high for a positive xx.

How many of these small steps can be chained before the errors pile up is left unanswered here: repeat the approximation across many weeks and the individually tiny quadratic corrections accumulate into something no longer safe to ignore.

How to use it

A hydrologist watching a reservoir’s weekly inflow rise by a steady 3%3\% wants the additive change on a log scale, to add it directly to other small weekly effects tracked in the same report. At first order,

log⁡(1.03)≈0.03. \log(1.03)\approx0.03.

Including the correction term gives

0.03−0.0322=0.03−0.00045=0.02955, 0.03-\frac{0.03^2}{2} =0.03-0.00045 =0.02955,

close to the true value 0.029560.02956. For a single week’s entry in the log, the first-order number is good enough to write down. The hydrologist reaches for the corrected figure only when several weeks of readings are being summed together, because using the plain xx for every one of many weeks would let the small overestimates from each week quietly compound into a report that overstates total inflow. This approximation converts multiplicative growth into an additive one that is easy to add up, but only as long as xx stays small; a week with an inflow spike far above 3%3\% needs the exact logarithm, not this shortcut. This is an Independent rule: with a small dimensionless perturbation, it converts multiplicative change into additive change in many settings.

1.2: Compress the Calculation and Expose Dominant Structure

Some algebraic expressions are difficult only because they are written at the wrong level of detail. A long list may be governed by its endpoints, an infinite tail by its first unseen term, and a complicated polynomial by one dominant power. The rules in this section teach a common move: preserve the structure that controls the answer and postpone everything else.

1.2.1: Average the endpoints of an arithmetic list

History

Adding a long list of evenly spaced numbers one term at a time is slow, and slow arithmetic invites mistakes exactly where a business cannot afford one. In the mathematical portion of the Aryabhatiya, composed in 499 CE at Kusumapura in India, Aryabhata gave a procedure for summing arithmetic progressions directly from the number of terms and the two endpoints, without adding each one in turn. The surviving text presents this in compact verse meant to carry a calculation, not in modern symbolic notation.

That supports the connection to evenly spaced lists and their sums; the endpoint-average formula below is a modern algebraic unpacking of his verse, not a quotation from it. What his procedure bought its user was a whole list collapsed into three numbers.

The equation

For an arithmetic progression

ak=a1+(k−1)d, a_k=a_1+(k-1)d,

with nn terms and final term ana_n,

Sn=a1+a2+⋯+an=n(a1+an2). S_n=a_1+a_2+\cdots+a_n =n\left(\frac{a_1+a_n}{2}\right).

In words: sum = number of terms × average of the endpoints.

How to read it

The formula says no term needs handling on its own: pair the first term with the last, the second with the next-to-last, and keep working inward, the way a worker might load a truck from both ends of a numbered rack toward the middle. Constant spacing guarantees every pair adds to the same total, first plus last; an odd count just leaves one middle term sitting at that same average, so nothing special has to be done for it.

The common difference itself never appears in the final formula: it only certifies that the pairing trick is valid. A broken list is invisible to the formula as written: skip an entry, or let the spacing wobble even once, and the pairing argument collapses even though the arithmetic still runs.

How to use it

A lumberyard worker is filling an order for planks cut to linearly increasing lengths, where plank kk measures kk board-feet, and the order runs from plank 3737 through plank 8282 inclusive; the worker needs the total board-feet without walking the rack and adding forty-six numbers by hand. Counting inclusively first,

n=82−37+1=46, n=82-37+1=46,

then averaging the endpoints and multiplying,

37+822=59.5,S=46(59.5)=2737, \frac{37+82}{2}=59.5, \qquad S=46(59.5)=2737,

gives the total directly. A rough check agrees: about fifty planks averaging about sixty should total somewhere near three thousand, so 27372737 looks right and goes straight onto the shipping ticket. The one place this shortcut can quietly fail is the count itself: forgetting that both endpoints are included would understate nn by one, and a gap somewhere in the numbering, say a missing plank 6060, would break the even spacing the whole pairing argument depends on even though the formula would still return a confident number. The worker walks the rack once more for gaps first, then logs the total.

This is an Independent rule: once the list is known to be arithmetic, it produces the desired sum directly rather than serving only as a step in a larger algorithm.

1.2.2: Estimate a geometric tail from its first omitted term

History

A parabolic segment, sliced again and again into ever-smaller triangles, is the artifact at the center of this rule. Around 250 BCE in Syracuse, Archimedes built exactly that construction in Quadrature of the Parabola: an initial triangle, then families of smaller triangles whose combined area shrank by a factor of one quarter at each stage, and he proved the whole infinite pile totals four thirds the area of the first triangle alone.

Archimedes never wrote a modern remainder formula. What his construction establishes is the underlying self-similar geometry: after any stopping point, the pieces still missing repeat the same shrinking pattern seen so far. Reading that leftover from its first missing piece is the modern rule-of-thumb reading of his geometry.

The equation

For a geometric series with |r|<1|r|<1, the tail beginning at index nn is

Rn=∑k=n∞ark=arn1−r. R_n=\sum_{k=n}^{\infty} ar^k =\frac{ar^n}{1-r}.

If the first omitted term is T=arnT=ar^n, this becomes simply

Rn=T1−r. R_n=\frac{T}{1-r}.

How to read it

The first term left out of a sum sets the scale of everything still missing; the factor 1/(1−r)1/(1-r) accounts for the whole shrinking tail that follows it, the way a single quarter-sized copy of a shrinking pile of scraps stands in for every scrap after it. When rr is small, the leftover tail is barely bigger than that first missing piece. When rr sits close to 11, the tail balloons because many later terms still matter almost as much as the first one. For a negative rr, the tail alternates sign, so its true size can be smaller than a bound built only from magnitudes would suggest.

The formula only measures a leftover that is genuinely geometric: it cannot bound a remainder whose ratio drifts from term to term.

How to use it

A quality-control engineer models a cumulative calibration error as a sum of contributions, each one third of the previous contribution. After recording the initial contribution and four further contributions, the engineer has kept the terms through k=4k=4 in

∑k=0∞(13)k. \sum_{k=0}^{\infty}\left(\frac13\right)^k.

The first contribution not recorded is k=5k=5, worth (1/3)5=1/243(1/3)^5=1/243 of the initial contribution, so the uncounted tail is

R5=1/2431−1/3=1162≈0.00617. R_5=\frac{1/243}{1-1/3} =\frac{1}{162} \approx0.00617.

That number is the total uncounted error, expressed as a fraction of the initial contribution. If the specification requires truncation error below 0.5%0.5\% of that contribution, stopping at k=4k=4 misses it, since 0.617%0.617\% remains uncounted. The engineer records one more contribution before signing off. The calculation assumes that the ratio stays exactly one third; if later contributions follow a different pattern, the geometric tail no longer gives the exact remainder and needs a new bound.

This is an Independent rule: under its convergence condition, it returns the entire remaining error in one calculation.

Both sequences decrease in parallel on a logarithmic vertical scale; the total tail is always above its first term.

Figure 1.2. For a=1 and r=1/3, the remainder starting at n is 1.5 times the first omitted term. At n=5 the remainder is 1/162. The logarithmic axis shows the constant ratio.

1.2.3: Let the leading term predict polynomial end behavior

History

Getting a large-scale prediction wrong by trusting a small piece of a formula can be expensive long before anyone finishes the exact calculation. Leonhard Euler routinely ordered expressions by their powers rather than treating every term as equally important, a habit running through his 1748 Introductio in analysin infinitorum.

Euler worked with power series and algebraic structure directly, but he wrote before today’s formal limit notation and convergence language, so it would overreach to claim he stated this classroom rule in its present form. What his power-ordering habit bought later readers is a shortcut: far from the origin, a polynomial can be read off its single highest power instead of its full expression.

The equation

If

p(x)=anxn+an−1xn−1+⋯+a0,an≠0, p(x)=a_nx^n+a_{n-1}x^{n-1}+\cdots+a_0, \qquad a_n\ne0,

then

p(x)anxn→1as|x|→∞. \frac{p(x)}{a_nx^n}\longrightarrow1 \quad\text{as}\quad |x|\longrightarrow\infty.

Equivalently, for sufficiently large |x||x|,

p(x)∼anxn. p(x)\sim a_nx^n.

How to read it

Divide every term of the polynomial by its leading term, anxna_nx^n. That term becomes exactly 11, and every lower-degree term shrinks toward zero because it carries a negative power of xx, the way a small side street stops mattering once you are far enough down the highway. The polynomial’s degree tells you whether the two far ends point the same way or opposite ways; the sign of the leading coefficient tells you which way.

Exactly where “sufficiently large” begins is left unstated. It can sit surprisingly close to the origin or surprisingly far from it, depending on the size of the lower coefficients, and the rule itself gives no warning of the crossover.

How to use it

Before quoting a client, a logistics analyst checks a shipping surcharge model, p(x)=−2x5+30x2−7p(x)=-2x^5+30x^2-7, where xx is order weight in some convenient unit, wanting to know whether a very large order’s surcharge will run positive or negative. The leading term, −2x5-2x^5, is odd degree with a negative coefficient, so the model predicts a surcharge falling further negative as order weight grows without bound. At x=10x=10,

p(10)=−2(10)5+30(10)2−7=−200,000+3,000−7=−197,007, p(10)=-2(10)^5+30(10)^2-7=-200{,}000+3{,}000-7=-197{,}007,

so the leading term already carries the right sign and roughly the right size, and the analyst trusts the formula’s direction for a genuinely large order. The catch shows up at moderate weight:

p(2)=−2(2)5+30(2)2−7=−64+120−7=49, p(2)=-2(2)^5+30(2)^2-7=-64+120-7=49,

even though the leading term alone is −64-64, because the 30x230x^2 term still dominates there. Quoting a midsize order using only the leading term would give a client a wildly wrong number and the wrong sign, so the analyst runs the full polynomial for anything short of a genuinely large weight and reserves the shortcut for extremes.

This is an Independent rule: it directly answers the end-behavior question without requiring the rest of a polynomial-solving procedure.

1.2.4: Rationalize a difference of nearby square roots

History

The conjugate trick for two nearby square roots is the artifact at the center of this rule: multiply through by it, and the cancellation trap that ruins a direct subtraction disappears. No dated page records who first reached for it that way. This rule has an evidence gap. Conjugate rationalization is old algebra, and the underlying identity is easy to find in general histories, but the research turned up no dated primary calculation, contemporary teaching text, or documented failure in which a historical actor used this exact nearby-root maneuver to stop digits from cancelling.

A broad history of radicals can show the component identity existed; it cannot show anyone used it this way, for this reason, on a given day. Closing the gap would need a located primary text, in an identifiable hand, applying the conjugate specifically to rescue a near-cancelling subtraction. Until then, the fair description is folk mathematics: a stabilizing trick that spread through calculation practice without a paper trail.

The equation

When both roots exist and their sum is nonzero, multiplying by the conjugate gives the exact identity

a+h−a=ha+h+a. \sqrt{a+h}-\sqrt a =\frac{h}{\sqrt{a+h}+\sqrt a}.

When a>0a>0 and |h|≪a|h|\ll a, read “the magnitude of hh is much smaller than aa”, the denominator is close to 2a2\sqrt a, so

a+h−a≈h2a. \sqrt{a+h}-\sqrt a \approx\frac{h}{2\sqrt a}.

How to read it

Subtracting two nearly equal square roots by hand throws away most of the meaningful digits before you even get an answer, the way rounding two almost-identical bills to the nearest dollar can erase the one cent that actually mattered. Multiplying by the conjugate moves the small quantity hh up into the numerator and divides it by a well-behaved sum instead. This step is exact, not approximate: it is simply the difference-of-squares identity

(a+h−a)(a+h+a)=h. (\sqrt{a+h}-\sqrt a)(\sqrt{a+h}+\sqrt a)=h.

Only the last step, replacing the denominator by 2a2\sqrt a, is an approximation, and it needs hh to be small next to aa.

A genuinely large gap between the two roots is not something the rationalized form fixes on its own. The identity stays exact, but the final shortcut denominator becomes the wrong one to trust.

How to use it

A land surveyor checking whether a foundation verified to be square matches its deed measures the plot’s corner coordinates and computes an area of 10,00110{,}001 square units where the deed specifies exactly 10,00010{,}000. The side length that would produce the deeded area is 100100, so the correction owed on the measured side is 10001−100\sqrt{10001}-100. Rather than rounding 10001\sqrt{10001} and subtracting 100100 by hand, which throws away the very digits that matter, the surveyor rationalizes:

10001−100=110001+100≈1200=0.005. \sqrt{10001}-100 =\frac{1}{\sqrt{10001}+100} \approx\frac{1}{200} =0.005.

The exact rationalized value is about 0.0049998750.004999875, so the shortcut recovers the scale and several digits immediately, and the surveyor logs the side length as roughly half a hundredth of a unit above the deeded length, well inside the crew’s tolerance. If the discrepancy had instead been large, say a plot closer to 12,00012{,}000 than 10,00010{,}000, replacing the denominator by a rounded 2a2\sqrt a would no longer track the true value closely, and the exact quotient, not the shortcut, would be the number to trust.

This is an Independent rule: for this recognizable form it directly supplies an exact stable representation, and often the desired estimate, even though the historical evidence for a specific originating episode remains open.

1.2.5: Use logarithms to tame products and powers

History

Trying to directly compare two enormous whole-number totals by multiplying them out by hand invites exactly the kind of arithmetic slip that changes which number is actually bigger. In 1614, John Napier published Mirifici Logarithmorum Canonis Descriptio in Edinburgh. Its tables and rules let astronomers and other human computers replace laborious multiplications, divisions, and root extractions with easier additive operations, a purpose the historical record documents directly.

Napier’s own construction and notation were not identical to a modern algebra textbook’s, so the equations below should not be read as his verbatim phrasing. His tables did the transforming; the notation below only records it: move a multiplicative problem onto an additive scale, do the easier work there, and translate back only if the original units are needed.

The equation

For positive real factors,

log⁡(∏i=1mai)=∑i=1mlog⁡ai, \log\left(\prod_{i=1}^{m}a_i\right) =\sum_{i=1}^{m}\log a_i,

and for x>0x>0,

log⁡(xp)=plog⁡x. \log(x^p)=p\log x.

Because the natural logarithm, and every real logarithm with base greater than 11, is strictly increasing, A>B>0A>B>0 exactly when log⁡A>log⁡B\log A>\log B.

How to read it

Logarithms change the grammar of a calculation: products become sums, quotients become differences, and exponents turn into ordinary multipliers out front, the way turning a stack of population figures into their logarithms turns a wall of huge numbers into a short column easy to compare by eye. Because the logarithm is strictly increasing, comparing two logged values is exactly the same as comparing the original quantities, so the originals never need to be written out in full.

Comparing two things safely once either one is zero or negative is outside what logarithms can do; the real logarithm simply has nothing to say there.

How to use it

A public health officer is comparing two rough outbreak projections to decide which needs emergency resources first: one model has cases tripling across 2020 generations, the other doubling across 3232 generations, and nobody wants to multiply either number out by hand under deadline pressure. Taking natural logs instead,

log⁡(320)=20log⁡3≈21.97,log⁡(232)=32log⁡2≈22.18, \log(3^{20})=20\log3\approx21.97, \qquad \log(2^{32})=32\log2\approx22.18,

and since 22.18>21.9722.18>21.97, the doubling model actually produces the larger case count, 232>3202^{32}>3^{20}, despite tripling sounding more alarming generation for generation. The log gap, about 0.2080.208, converts back to a ratio of e0.208≈1.23e^{0.208}\approx1.23, so the officer reports the doubling scenario as roughly a quarter larger and prioritizes resources there first. The comparison only holds because both case counts are positive; if a model instead tracked a net change that could go negative, taking a real logarithm of it would be invalid, and rounding on the log scale can itself become significant once the huge number is reconstructed, so the officer keeps a few extra digits of precision rather than rounding early.

This is a Workflow rule: it is highly transferable, but the logarithm usually creates an intermediate representation for a later comparison, equation solve, optimization, asymptotic analysis, or numerical evaluation.

1.3: Diagnose Structure Before Choosing a Procedure

The fastest algebra often happens before the formal solution begins. A small change of form can expose the roots, geometry, admissible methods, or hidden hazards of a problem. The seven rules in this section are therefore less about executing an algorithm than about deciding what kind of problem is actually in front of you.

1.3.1: Make a quadratic monic before reading its structure

History

A quadratic scaled up by an arbitrary leading coefficient hides its own root pattern until that scale is stripped away. Dividing it through by that coefficient turns two of the remaining coefficients directly into the sum and the product of its roots. Around 820 in Baghdad, al-Khwarizmi classified quadratic equations into verbal types and solved them by operations he called restoration and balancing; for a type equivalent to x2+bx=cx^2+bx=c, his geometric demonstration added pieces to complete a square and made the unknown readable as a length. He never wrote ax2+bx+c=0ax^2+bx+c=0, nor did he state dividing by the leading coefficient first; that connection is a modern reconstruction that exposes why a unit-scale squared term makes a quadratic’s structure visible.

The equation

For a≠0a\ne0, divide every term by aa:

[ ax^2+bx+c=0 x^2+ba x+ca=0. ]

If the roots are r1r_1 and r2r_2, the monic form immediately gives

[ r_1+r_2=-ba, r_1r_2=ca. ]

“Monic” simply means that the coefficient of the highest power is 11.

How to read it

Normalizing separates two kinds of information: the original aa sets an overall vertical scale, while the ratios b/ab/a and c/ac/a control the root relationships. Once aa is divided away, the coefficient of xx is the negative root sum and the constant is the root product, so sign and size patterns become visible at a glance.

Not every feature of the original function survives this division: rescaling f(x)f(x) by aa changes its vertical scale and flips it upside down when aa is negative, even though the zeros themselves stay exactly where they were.

How to use it

A homeowner has a contractor’s equation for a garden path’s edge positions, 2x2−10x+12=02x^2-10x+12=0, and wants the two edge locations without hunting for a formula. Dividing every term by 22,

[ x^2-5x+6=0, ]

the roots must sum to 55 and multiply to 66. The pair 22 and 33 fits, so

[ x^2-5x+6=(x-2)(x-3), ]

and the path’s edges sit at x=2x=2 and x=3x=3. The homeowner marks the stakes without computing a discriminant. The move is exact, but if the leading coefficient is itself a rough field measurement rather than a clean 22, dividing by it can magnify whatever error was already there, so the crew checks the original measurement first and only then trusts the divided-through equation.

This is a Workflow rule: it rarely finishes the calculation, but it improves several procedures that may follow.

1.3.2: Complete the square to read a quadratic’s geometry

History

Was al-Khwarizmi solving an equation, or building a shape? In the same Baghdad treatise, around 820, for equations equivalent to x2+bx=cx^2+bx=c, he arranged rectangular pieces around a square and supplied the missing corner pieces so the whole figure became a larger, completed square, letting the unknown be read off as a length. His documented act stops at the completed square. Coordinate geometry did the rest, turning that same figure into a graph’s minimum, maximum, or vertex with notation he never had.

The equation

For a≠0a\ne0,

[ ax^2+bx+c =a(x+)^2 +c-. ]

Equivalently, in vertex form,

[ ax2+bx+c=a(x-h)2+k, h=-, k=c-. ]

The point (h,k)(h,k) is the parabola’s vertex.

How to read it

Vertex form turns three coefficients into three geometric facts: hh marks the axis of symmetry, kk is the extreme value reached there, and the sign of aa says which way the parabola opens and whether kk is a minimum or a maximum. Move an equal distance either way from hh and the squared term returns the same value both times.

The form also previews real zeros: for a>0a>0, k>0k>0 rules them out, k=0k=0 gives one repeated zero, and k<0k<0 allows two. Whether the vertex answers the actual question is separate from finding it: on a restricted domain, the true extreme value can sit at a boundary instead.

How to use it

A track coach has modeled a sprinter’s split-time variance across repeated intervals as 2x2−8x+112x^2-8x+11, where xx counts the interval number, and wants the least variance the model predicts before scheduling a taper. Factoring 22 from the first two terms, then completing the square:

[ 2x2−8x+11=2(x2−4x)+11=2[(x−2)2−4]+11=2(x−2)2+3.\begin{aligned} 2x^2-8x+11 &=2(x^2-4x)+11\\ &=2\bigl[(x-2)^2-4\bigr]+11\\ &=2(x-2)^2+3. \end{aligned}

]

Since 2(x−2)2≥02(x-2)^2\ge0, the modeled variance never drops below 33, reached at interval x=2x=2, so the taper is scheduled there without plotting another point. The vertex only answers the question if interval 22 falls inside the session actually run: a workout that stops at interval 11 has its true minimum at whatever endpoint it stopped at, not at the untested x=2x=2.

This is a Workflow rule: the transformation is widely reusable, but its purpose is to enable a later geometric, optimization, sign, or root judgment.

1.3.3: Factor before reaching for the quadratic formula

History

Two techniques for solving the same equation can each have a long documented past without proving which one a solver reached for first. This rule has an evidence gap. Factoring and general quadratic-solving methods both have deep histories, but the research found no dated primary calculation, contemporary teaching text, or documented failure pinning “try factoring first” down as a lesson some historical figure actually taught. A broad history showing both techniques existed side by side does not prove any solver used this exact method-selection habit. Attaching the advice to a famous algebraist for color would turn a plausible guess into an invented story; the honest position is that it is a modern workflow convention whose mathematics is sound even though its origin story is not yet documented.

The equation

Factoring rewrites a sum as a product. Two common patterns are

[ ab+ac=a(b+c) ]

and

[ x^2-(r+s)x+rs=(x-r)(x-s). ]

Once an equation has the form

[ U(x)V(x)=0, ]

the zero-product property says that U(x)=0U(x)=0 or V(x)=0V(x)=0.

How to read it

A factor sitting in plain sight is compressed information about where the solutions are, the way spotting a shared tool already out on the workbench beats digging through the whole toolbox. A common factor exposes one whole branch of answers, a difference of squares exposes a symmetric pair, and a pair of integers matching the coefficients’ sum and product can hand over both roots at once. “Factor first” does not mean guess forever. Run one quick scan for a greatest common factor, a difference of squares, or an obvious integer pair, then move on.

A failed scan proves nothing about existence: plenty of quadratics with perfectly good real or complex roots simply do not factor over the integers.

How to use it

A tutoring center coordinator is walking a student through 6x2−15x=06x^2-15x=0 and wants the fastest path to the answer rather than the quadratic formula by default. Both terms share the factor 3x3x, so

[ 3x(2x-5)=0, ]

giving x=0x=0 or x=5/2x=5/2. The quadratic formula reaches the same answers, but only after a discriminant and a fraction are simplified along the way. The coordinator warns against dividing the equation by xx to simplify it, since that silently erases the solution x=0x=0; keeping the equation at zero and using the zero-product property protects both roots. When a quick scan turns up nothing, the student moves straight to the discriminant instead of hunting for a factor that may not exist.

This is a Workflow rule because it selects a cheaper route when one is available; it is not an independent solution method for every quadratic.

1.3.4: Use the discriminant to triage quadratic roots

History

Solving an equation all the way through only to discover the answer was never real is wasted work that a smaller check could have caught first. In 1545, Girolamo Cardano’s Ars Magna confronted the different algebraic cases that appear as coefficients change, presenting systematic solutions of cubic and quartic equations, some passing through quantities now interpreted as negative and complex numbers. Cardano never used a symbol DD, nor did he lay out today’s three-line quadratic triage. His documented contribution is the sustained encounter with changing root cases; folding that same case distinction for a quadratic into the sign of one coefficient expression is a modern shorthand for his struggle, not his own notation.

The equation

For

[ ax^2+bx+c=0, a, ]

the discriminant is

[ D=b^2-4ac. ]

For real coefficients,

[ {D>0:two distinct real roots,D=0:one repeated real root,D<0:a complex-conjugate pair.\begin{cases} D>0 &: \text{two distinct real roots},\\ D=0 &: \text{one repeated real root},\\ D<0 &: \text{a complex-conjugate pair}. \end{cases}

]

It is the quantity beneath the square root in x=(−b±D)/(2a)x=(-b\pm\sqrt D)/(2a).

How to read it

The discriminant is a classifier before it is anything else: its sign says whether a parabola crosses the horizontal axis twice, just grazes it once, or never reaches it. With rational coefficients, a positive perfect-square discriminant points to rational roots, a positive non-square to irrational ones.

Values near zero deserve caution: they mark the boundary between two roots and none, so a small coefficient change nearby can flip the classification.

How to use it

A bridge inspector has a quadratic model, 3x2+2x+5=03x^2+2x+5=0, for where a stress curve might cross zero along a beam, and wants to know whether to expect a crossing before running a full numeric solve. Computing the discriminant,

[ D=2^2-4(3)(5)=4-60=-56, ]

the negative result says immediately that the stress curve never actually reaches zero anywhere along the beam, so the inspector stops there rather than launching a root search. Because the coefficients came from rounded field measurements, the inspector also checks how close DD sits to zero relative to that measurement uncertainty; −56-56 is safely negative even after reasonable rounding, but a value near zero means the classification waits on tighter data first.

This is a Workflow rule: it is a self-contained diagnosis, yet its chief role is to select the next procedure, or show that no further procedure is needed for the question asked.

Three upward parabolas cross the horizontal axis twice, touch it once, or remain above it.

Figure 1.3. The constructed family y=x²+c has two, one repeated, or no real roots as c is negative, zero, or positive. Near discriminant zero, small coefficient changes can alter the classification.

1.3.5: Use the rational-root theorem as a shortlist

History

Searching every possible rational root of a polynomial by trial is hopeless; the candidates are infinite until something narrows the field. Carl Friedrich Gauss’s Disquisitiones Arithmeticae, tied to Leipzig in 1801, systematized congruences, divisibility, and the arithmetic of polynomial forms, and its structural treatment of integer factors made finite divisibility tests a natural first screen before more general algebraic methods. Gauss never stated the present classroom advice in these words; his documented contribution is the structural treatment of divisibility itself. The rule-of-thumb connection modernizes that structure for a narrower job: shrinking infinitely many possible rational roots to a short, testable list.

The equation

For an integer-coefficient polynomial

[ P(x)=a_nx^n++a_1x+a_0, ]

if a rational root is written in lowest terms as p/qp/q, then

[ pa_0, qa_n. ]

If a0=0a_0=0, first factor out the largest power of xx, record zero as a root, and apply the theorem to the remaining polynomial. Otherwise, every rational root must appear among

[ , ]

after duplicate fractions are reduced.

How to read it

The theorem hands over a necessary condition, not a promise, the way a short list of known combinations narrows a stuck padlock without opening it: a candidate missing from the list cannot be a rational root, but one on the list is only eligible and still needs testing by substitution or synthetic division. The constant term controls possible numerators, since it is what remains at x=0x=0; the leading coefficient controls possible denominators, since clearing powers of qq forces divisibility at the other end.

A short list is not guaranteed: when both outer coefficients carry many divisors, the candidate list can grow too long to be practical.

How to use it

A tank designer has a cubic volume-balance equation, 2x3−3x2−8x+12=02x^3-3x^2-8x+12=0, for a design parameter xx, and needs exact balance points rather than a numerical guess. Possible reduced rational roots, from factors of 1212 over factors of 22, are

[ , , , , , ,  , . ]

Testing x=2x=2 gives 16−12−16+12=016-12-16+12=0, exposing x−2x-2 as a factor; synthetic division gives

[ 2x3-3x2-8x+12=(x-2)(2x-3)(x+2), ]

so the balance points are x=2x=2, 3/23/2, and −2-2. Only the positive values matter physically, so the designer chooses between 22 and 3/23/2 using the design’s other constraint. The theorem only screens rational candidates: had none of the sixteen candidates worked, that would rule out a rational answer but say nothing about a real or complex one, leaving a numerical solver as the next step rather than a conclusion that no design exists.

This is a Workflow rule: it narrows a search and verifies promising exact factors, but it does not finish the problem until the candidates are tested.

1.3.6: Let denominator factors dictate partial fractions

History

Get the setup for a rational function’s decomposition wrong, and every integral, transform, or system solution built on it inherits the mistake. Leonhard Euler’s 1748 Introductio in analysin infinitorum treats rational functions directly through the factors of their own denominators, unlike the series expansions used elsewhere in this chapter: this is direct historical use of the governing structure, not a later reading layered onto work built for something else. Today’s concise template packages that same practice; what it buys is a decomposition whose shape is dictated before a single coefficient is computed.

The equation

For a proper rational function with distinct linear factors,

[ =+. ]

A repeated factor (x−a)m(x-a)^m requires every power,

[ + ++, ]

while an irreducible quadratic factor q(x)mq(x)^m requires a linear numerator (Bkx+Ck)/q(x)k(B_kx+C_k)/q(x)^k for each k=1,…,mk=1,\ldots,m.

How to read it

Each factor in a denominator earns its own family of terms above the line, the way sorting mail into boxes gives each address its own simpler pile instead of one tangled bag: how many times a factor repeats decides how many powers must appear, and its degree, the highest power present in that factor, decides the numerator’s degree over each one. An irreducible quadratic factor, one that cannot be split into real linear factors, needs a linear numerator instead of a constant. The unknown coefficients never change this template; they only pick out which member of it equals the function at hand.

What the template will not warn you about is a repeated or irreducible-quadratic factor hiding in the denominator: forcing the same distinct-linear pattern onto either one leaves coefficients with no consistent solution.

How to use it

Reducing a vibration model to the transfer function 1/[(x−1)(x+2)]1/[(x-1)(x+2)], a control-systems engineer needs it split before applying a standard inverse-transform table. The denominator’s two distinct linear factors dictate the template:

[ =+. ]

Clearing denominators gives 1=A(x+2)+B(x−1)1=A(x+2)+B(x-1); x=1x=1 gives A=1/3A=1/3, and x=−2x=-2 gives B=−1/3B=-1/3, so

[ =-. ]

Each piece is then looked up separately in the transform table. The setup holds only at the domain’s exclusions, x≠1,−2x\ne1,-2; if a later model revision turns one factor into a repeated one, this same two-term template would misrepresent the system’s response at that pole.

This is a Workflow rule: denominator-driven setup is one step in integration, inverse transforms, recurrence solving, and system analysis, but it transfers across all of those larger calculations.

1.3.7: Cross-multiply inequalities only after checking signs

History

A warning that is mathematically airtight can still have no verified birthday. This rule has a documented evidence gap. Sources on ancient proportion theory document plenty of ratio reasoning, but the research did not locate a primary episode where a consequential mistake was traced specifically to cross-multiplying by an expression of unknown sign. Treating any old ratio calculation as this rule’s origin would confuse the long history of proportion with the narrower history of this particular safety warning. What would close the gap is a dated calculation, a contemporary teaching text, or a recorded failure showing the warning itself at work. Until then, transparency about the missing anecdote is more honest than inventing one.

The equation

For nonzero bb and dd, begin with

[ ab<cd. ]

Multiplication by bdbd gives two different cases:

[ bd>0  ad<bc, bd<0  ad>bc. ]

The same rule applies to ≤,>,≥\le,>,\ge: multiplication by a positive quantity preserves the comparison, while multiplication by a negative quantity reverses it.

How to read it

Cross-multiplying is multiplication by the product of two denominators, so whether it is even legal depends on the domain, and which direction the inequality points afterward depends on that product’s sign. A constant positive denominator is easy; a variable denominator changes sign at its own zeros and splits the number line into intervals worth checking separately.

A denominator equal to zero is never part of the solution. Tracking a denominator that changes sign partway through the domain is exactly what one global cross-multiplication cannot do; only splitting at its zeros can.

How to use it

A bookkeeper is solving a scaling inequality, 1/x<21/x<2, for a signed adjustment factor xx that a ledger entry is divided by, so the entry’s effective multiplier is 1/x1/x (positive for a credit correction, negative for a debit one); company policy caps any effective multiplier at 22, and the bookkeeper needs the full valid range before reconciliation software accepts an entry. The domain excludes x=0x=0. For x>0x>0, multiplying by xx preserves the sign:

[ 1<2xx>. ]

For x<0x<0, multiplying flips it:

[ 1>2x, ]

true for every negative xx, so the valid range is

[ x(-,0)(,). ]

Both pieces go into the reconciliation entry, unlike the single positive interval a careless cross-multiplication would give, which would have discarded a whole valid branch of legitimate debit corrections.

This is a Workflow rule, specifically a guardrail. It does not solve rational inequalities by itself, but it prevents an otherwise valid transformation from reversing the answer or admitting an undefined point.

Chapter Synthesis: A Structural First Pass

The rules in this chapter do not form one long algorithm. They form a set of questions to ask before an algorithm is chosen.

Start with scale: decide whether the comparison is additive or multiplicative, and whether a genuinely small dimensionless quantity permits a local model. Then look for compression: symmetry, repeated ratios, dominant terms, and changes of representation can replace long arithmetic with a shorter equivalent question. Finally, diagnose structure before selecting a procedure. Normalization, factor patterns, coefficient tests, denominator structure, and sign partitions reveal what an exact method must preserve.

The deeper habit is to separate four questions:

  1. What representation exposes the controlling structure?
  2. Is the proposed step exact, approximate, or merely diagnostic?
  3. What assumption makes it legal?
  4. Does it answer the question, or does it prepare the next calculation?

That last distinction prevents two opposite errors. An independent estimate should not be buried under unnecessary computation. A workflow guardrail should not be mistaken for a finished answer.

One-Page Algebra Toolkit

Recognition cue Rule to try What it gives Role
Equal differences but different baselines Compare with ratios Scale factor or relative change Independent
Steady exponential rate Divide () by the rate Doubling horizon Independent
Small change inside a power ((1+x)^a+ax) First-order power-law response Independent
Small exponent (e^x+x) Local growth or attenuation Independent
Factor near one inside a logarithm ((1+x)x) Additive form of multiplicative change Independent
Evenly spaced finite list Average the endpoints Exact sum Independent
Fixed-ratio convergent tail First omitted term over (1-r) Exact remainder or error scale Independent
Large-input polynomial behavior Keep the leading term Tail direction and scale Independent
Difference of nearby radicals Multiply by the conjugate Exact stable form and local estimate Independent
Huge product or power Take logarithms Additive comparison scale Workflow
Quadratic obscured by common scale Make it monic Root-sum and root-product structure Workflow
Quadratic geometry or extremum Complete the square Vertex, bound, and sign information Workflow
Visible common or integer factor Factor briefly before general methods Exact solution branches Workflow
Need root type before root values Compute (b^2-4ac) Real/repeated/complex classification Workflow
Integer polynomial may have rational factors Build the rational-root shortlist Finite candidate set Workflow
Rational function with factored denominator Mirror its factors in the template Partial-fraction setup Workflow
Variable denominators in an inequality Split by denominator zeros and signs Valid order-preserving transformation Workflow guardrail

Decision Path

Transfer Problems

1. Same gain, different meaning

A small organization adds 200 customers, growing from 400 to 600. A large organization adds the same 200, growing from 20,000 to 20,200. Compare the changes on the appropriate scale. Then state why a raw difference answers a different question.

2. Build an approximation budget

Estimate each quantity without first using a calculator:

[ (1.01)^8,e^{-0.04},(1.04). ]

For each estimate, write the leading omitted correction and decide whether a requested accuracy of (10^{-3}) is plausible.

3. Choose the procedure before solving

For each expression, name the first structural rule you would apply and explain why:

[ 4x^2-20x+24=0, , . ]

Do not finish the calculations until you have stated the relevant domain exclusions, factor structure, or root diagnostic.

Where These Ideas Reappear

Historical Notes and Sources

Sources for the historical accounts in this chapter follow. The evidence-gap entries support the mathematics of their rules without claiming an unverified origin event.