← Illustrated chapter

Chapter 4 , Calculus: Turn Local Change into Global Control

Without evaluating a square root, estimate 4.04\sqrt{4.04}. The nearby value 4=2\sqrt4=2 is known, and the slope of x\sqrt x at 44 is 1/41/4. An input increase of 0.040.04 should therefore produce an output increase of about (1/4)(0.04)=0.01(1/4)(0.04)=0.01, giving 2.012.01. The actual value is about 2.0099752.009975.

That tiny calculation captures calculus at its most reusable. A derivative turns nearby change into a prediction. A derivative bound turns local slope information into a guarantee across an interval. Convexity traps a function between simple lines. An integral converts local contributions into accumulated effect, while comparison and symmetry can control that accumulation without an antiderivative.

The seventeen rules in this chapter fall into three families. The first builds local models and sensitivity estimates. The second reads global shape from derivative signs, convexity, and boundary growth. The third controls integrals, series, and method choice. Together they support one governing habit: use local information to make global decisions.

Calculus is not merely a catalogue of operations. Its deeper value is knowing what a slope, bound, comparison, or first omitted term licenses, and what it does not.

4.1: Build Useful Local Models

A curve viewed closely often behaves like its tangent. The four rules in this section turn that local resemblance into estimates, root corrections, percentage sensitivities, and hard change bounds.

4.1.1: Linearize near a point you already understand

History

Freeze a changing quantity at an old value, and an estimate quietly stops tracking reality the moment the input moves on. Between 1665 and 1671, working at Woolsthorpe and Cambridge, Isaac Newton developed methods of fluxions and infinite series connecting changing quantities, tangents, quadratures, and local approximation. Much of it circulated in manuscript before later publication, yet it became one foundation of differential calculus. Newton used tangent and series methods explicitly to read nearby behavior from a point already understood; the compact linearization formula below, and its error language, are later, more explicit expressions of that habit, not his own notation.

The equation

If ff is differentiable at aa, then for a small increment hh,

f(a+h)=f(a)+f′(a)h+o(h). f(a+h)=f(a)+f'(a)h+o(h).

The linear approximation is therefore

f(a+h)≈f(a)+f′(a)h. f(a+h)\approx f(a)+f'(a)h.

If a reliable second derivative is available, the leading correction is often f″(a)h2/2f''(a)h^2/2.

How to read it

The known value f(a)f(a) is the answer at a point you already trust, and the slope f′(a)f'(a) converts a small step hh into a predicted output step, f′(a)hf'(a)h. The leftover term o(h)o(h) shrinks faster than hh itself as hh shrinks toward zero. What the formula does not say is how large a step still counts as close enough for the accuracy a given job needs; differentiability guarantees the trend, not a distance.

Pick the expansion point partly for convenience: a nearby perfect square, a familiar angle, or a tabulated value can make both f(a)f(a) and f′(a)f'(a) mental arithmetic.

A quick units check catches assembly mistakes: f′(a)f'(a) carries output units per input unit, so f′(a)hf'(a)h must carry the same units as f(a)f(a). A mismatch means the local model was built wrong before any number was plugged in.

How to use it

A quality inspector on a machining line measures a shaft blank’s verified square cross-sectional area at 4.044.04 square centimeters, against a nominal 44, and needs the equivalent side length before checking the shop’s tolerance chart. This is the chapter’s opening estimate at work: with f(x)=xf(x)=\sqrt x, a=4a=4, h=0.04h=0.04,

f(4)=2,f′(4)=124=14,4.04≈2+14(0.04)=2.01 cm. f(4)=2, \qquad f'(4)=\frac{1}{2\sqrt4}=\frac14, \qquad \sqrt{4.04}\approx2+\frac14(0.04)=2.01\text{ cm}.

The chart allows a side length up to 2.0152.015 cm, so the inspector passes the blank without running an exact square root. The next correction,

12(−132)(0.04)2=−0.000025, \frac12\left(-\frac1{32}\right)(0.04)^2=-0.000025,

predicts 2.0099752.009975 cm, confirming the tangent estimate runs slightly high.

Use linearization for mental arithmetic, calibration checks, and first-order uncertainty propagation, measured on the function’s own scale: a numerically tiny step can still cross a corner or region of rapid curvature. Had the tolerance been tighter than that quadratic correction, the quick estimate alone would not have been trustworthy. This is an Independent rule: once a trustworthy nearby point is chosen, it directly produces an estimate across many applications.

4.1.2: Interpret a Newton step as the tangent-line root

History

One tangent-line correction can land within thousandths of a true root, cutting many algebraic passes down to one. Newton produced that shortcut in De analysi, written in 1669 at Cambridge: an iterative method for improving an approximate root by substituting a correction term, keeping the part that controlled the answer, and dropping the rest as negligible. Joseph Raphson later recast the process in a form closer to the update taught today. Newton’s correction process is equivalent in substance to local linearization, though the derivative quotient and the geometric tangent-root picture are later standardizations, not his own algebra.

The equation

At a current estimate xkx_k, linearize

f(x)≈f(xk)+f′(xk)(x−xk). f(x)\approx f(x_k)+f'(x_k)(x-x_k).

Setting the right side to zero gives Newton’s update:

xk+1=xk−f(xk)f′(xk). x_{k+1}=x_k-\frac{f(x_k)}{f'(x_k)}.

Near a simple root and under suitable smoothness, the error typically becomes proportional to the square of the previous error.

How to read it

The quotient f(xk)/f′(xk)f(x_k)/f'(x_k) carries the units of xx: it estimates how far the current guess sits, horizontally, from where the curve actually crosses zero. Subtracting it moves the guess to where the local tangent line crosses the axis instead. A steep slope gives a modest correction, while a nearly flat one can fling the next guess far from where it started.

What the step actually solves is the tangent’s equation, not the original curve’s equation: a step can look small and still be a poor guide to the real root once the tangent has drifted from the curve. Its speed comes from rebuilding that tangent after every step, never from touching the nonlinear equation directly.

How to use it

A venture analyst wants the annual growth factor xx that would double a portfolio company’s revenue over exactly two years, the positive root of x2−2=0x^2-2=0. Starting from a guess x0=1.5x_0=1.5, with f(x)=x2−2f(x)=x^2-2 and f′(x)=2xf'(x)=2x,

x1=1.5−1.52−22(1.5)=1.5−0.253=1.416666…. x_1 =1.5-\frac{1.5^2-2}{2(1.5)} =1.5-\frac{0.25}{3} =1.416666\ldots.

A second step gives about 1.41421571.4142157, close enough to quote as a required growth rate of about 41.42%41.42\% per year. Before trusting it, the analyst checks the residual f(x1)f(x_1) rather than just the size of the correction: a small correction can be misleading when the derivative is poorly scaled, while a small residual confirms the equation itself is nearly satisfied. A nearly flat slope on an early guess, or a tangent pointing outside a bracket already believed to contain the answer, is reason enough to reject a step rather than accept it blindly.

Newton’s method earns its use once a credible starting bracket exists. Poor starts, flat derivatives, or multiple solutions in a more complicated growth model can cause slow convergence or divergence, so in decision-critical work the analyst keeps a bracket and rejects any step that leaves it. This is a Workflow rule: the tangent-root interpretation selects and diagnoses one step inside a larger iterative solution.

The function crosses zero near 1.414; its tangent at x=2 crosses at 1.5, to the right of the function’s root.

Figure 4.1. For f(x)=x²-2, the tangent at x=2 reaches zero at x=1.5. That improves the estimate but still differs from sqrt(2); later steps and a stopping check remain necessary.

4.1.3: Use relative differentials for power-law sensitivity

History

Notation is supposed to be neutral bookkeeping, a record of a calculation already understood elsewhere. Leibniz’s own symbols argued otherwise. His differential calculus appeared in 1684 in a Leipzig journal, with an account of integral calculus following in 1686, and his notation made the differential structure of products, composites, powers, and accumulated sums visible and directly manipulable. Leibniz supplied that differential language; turning a product of powers into a compact percentage-sensitivity checklist is a modern operational use of it, with the first-order assumptions stated explicitly. The payoff is practical: a design with several changing dimensions reduces to which single measurement matters most.

The equation

For positive reference values in

y=Cxazb, y=Cx^a z^b,

take logarithms and differentiate:

log⁡y=log⁡C+alog⁡x+blog⁡z, \log y=\log C+a\log x+b\log z,

dyy≈adxx+bdzz. \frac{dy}{y} \approx a\frac{dx}{x}+b\frac{dz}{z}.

With more factors, add one exponent-weighted relative change for each factor.

How to read it

Each exponent acts like a percentage multiplier tuned to its own input. If a=2a=2, a small 1%1\% nudge in xx contributes roughly 2%2\% to yy; a negative exponent flips that contribution’s direction. Because logarithms turn a product into a sum, contributions from several inputs simply add at this first-order level, one term per factor.

What this gives is signed sensitivity, not a worst-case bound: it flags which input’s error matters most, not the largest error several inputs could jointly produce. For finite fractional changes εx\varepsilon_x and εz\varepsilon_z, the exact scale factor is

y′y=(1+εx)a(1+εz)b, \frac{y'}y=(1+\varepsilon_x)^a(1+\varepsilon_z)^b,

where y′y' is the new value of yy, not a derivative, showing precisely which squared and cross terms the differential version leaves out.

The exponent-weighted coefficients are elasticities, percentage-change ratios rather than raw units: ranking them by size flags which input deserves the tightest measurement before any change is supplied.

How to use it

A concrete contractor is pouring a cylindrical silo column, and the formwork crew reports the radius will likely run 1%1\% over spec while the height runs 2%2\% short. Volume follows V=πr2hV=\pi r^2h, so

dVV≈2drr+dhh=2(1%)+(−2%)=0%. \frac{dV}{V}\approx2\frac{dr}{r}+\frac{dh}{h}=2(1\%)+(-2\%)=0\%.

The contractor reads this as the order being fine as specified, since the radius error carries twice the weight of the height error and the two nearly cancel. The exact scale factor,

(1.01)2(0.98)=0.999698, (1.01)^2(0.98)=0.999698,

confirms a decrease of only about 0.0302%0.0302\%, close enough to zero that no extra truck gets called.

Relative differentials answer exactly this kind of tolerance question, and flag which measurement, here the radius, deserves the tightest control on site. The formula assumes small fractional changes and a nonzero reference size; it does not say when “small” has stopped being small. If a form panel bows and the radius error grows to 10%10\%, the first-order cancellation can no longer be trusted, and the contractor should compute the exact ratio, since the very quadratic terms this rule discards can then become the whole story. This is an Independent rule: it directly estimates percentage response for power-law models across science, engineering, and finance.

4.1.4: Turn a derivative bound into a change bound

History

How much can a function possibly change, given only a ceiling on its slope? Newton’s tangent-and-series manuscripts from the 1660s already used local slopes to infer how a changing quantity behaved nearby, establishing the local idea long before anyone wrote today’s polished inequality. That inequality, which turns a single slope ceiling into a guarantee across an entire stretch, belongs to the later, rigorous account built around the mean value theorem. The payoff is a certificate that needs no formula at all: bound the slope everywhere on an interval, and no secant across it can exceed the same rate.

The equation

If ff is continuous on the closed interval between aa and bb, differentiable inside it, and

|f′(x)|≤M |f'(x)|\le M

throughout that interval, then

|f(b)−f(a)|≤M|b−a|. |f(b)-f(a)|\le M|b-a|.

How to read it

The mean value theorem guarantees some intermediate point cc between aa and bb where the secant slope equals the derivative there,

f(b)−f(a)b−a=f′(c). \frac{f(b)-f(a)}{b-a}=f'(c).

Because that unknown cc is still covered by the ceiling MM, bounding every slope on the interval bounds the total change. Where linearization predicts behavior near one point, this inequality is a certificate across the whole stretch between aa and bb, whenever its hypotheses hold.

A smaller valid MM tightens the result, but the bound never reveals the sign or exact size of f(b)−f(a)f(b)-f(a), only a ceiling on it; the real value still needs ff evaluated directly.

The inequality also reads as a stability statement: two nearby inputs cannot produce outputs farther apart than MM times their separation, the same idea that becomes a Lipschitz bound in Chapter 5.

For example, f(x)=log⁡xf(x)=\log x has |f′(x)|=1/x≤1/2|f'(x)|=1/x\le1/2 on [2,3][2,3], so

|log⁡2.1−log⁡2|≤12(0.1)=0.05. |\log2.1-\log2|\le\frac12(0.1)=0.05.

The interval matters as much as the starting point.

How to use it

A solar-monitoring technician models a panel’s relative output as proportional to sin⁡θ\sin\theta, where θ\theta is the sun’s elevation angle in radians, and needs to know how far a small angle-sensor error could shift the reading. Because |cos⁡θ|≤1|\cos\theta|\le1 for every angle,

|sin⁡(1.01)−sin⁡(1)|≤1(0.01)=0.01, |\sin(1.01)-\sin(1)|\le1(0.01)=0.01,

so a sensor drift of 0.010.01 radian cannot move the modeled output by more than 0.010.01 on that same scale, with no trigonometric evaluation required.

Apply this over the whole angle range the uncertainty could reach rather than at the nominal reading alone: if a passing cloud makes the sensor’s output jump rather than vary smoothly, the continuity this argument relies on breaks. A loose but honest guarantee still beats an unsupported number that merely looks precise. This is an Independent rule: it directly converts a local rate ceiling into a transferable global change guarantee.

4.2: Read Shape from Derivatives and Convexity

Local slopes organize an entire graph when their signs or ordering are known. The three rules in this section use derivatives to locate rises and falls, convexity to build two-sided bounds, and boundary rates to estimate thin geometric layers.

4.2.1: Use the derivative sign before solving for extrema

History

Newton’s manuscript on fluxions, circulating privately between 1665 and 1671, already linked a rate of change to a curve’s tangents, its quadratures, and its behavior near a point. That body of work, published only later, helped found differential calculus. Reading a rate’s sign as a verdict on where a curve climbs, falls, or turns is the modern extension of that habit, formalized today through the mean value theorem into a systematic sign chart.

The equation

On an interval where the required hypotheses hold,

f′(x)>0⇒f is increasing, f'(x)>0\quad\Longrightarrow\quad f\text{ is increasing},

f′(x)<0⇒f is decreasing. f'(x)<0\quad\Longrightarrow\quad f\text{ is decreasing}.

Partition the domain at points where f′=0f'=0, where f′f' does not exist, and at relevant endpoints.

How to read it

A single, unchanging derivative sign across an interval forces every secant slope there to share it, through the mean value theorem: positive throughout means the function only climbs, negative throughout means it only falls. An extremum shows up as a change of sign, positive to negative for a local maximum, negative to positive for a local minimum.

A zero derivative by itself is only a candidate, not a verdict: x3x^3 has derivative zero at the origin yet keeps climbing straight through it. Corners and cusps can also hide an extremum where no derivative exists, so those points need checking directly.

Repeated factors inside f′f' often predict the sign pattern without extra work: an odd-multiplicity factor usually flips sign as its root is crossed, an even-multiplicity factor usually does not, though a test point on each interval remains the final check.

How to use it

A logistics manager tracks a delivery route’s efficiency index as f(x)=x3−3xf(x)=x^3-3x, where xx counts stops added to or cut from the route’s baseline, negative for cuts and positive for additions. Rather than solve the model directly, the manager reads

f′(x)=3(x2−1)=3(x−1)(x+1) f'(x)=3(x^2-1)=3(x-1)(x+1)

for its sign: positive below x=−1x=-1, negative between −1-1 and 11, positive above x=1x=1, so the index rises, falls, then rises, peaking after cutting one stop at f(−1)=2f(-1)=2 and bottoming out after adding one at f(1)=−2f(1)=-2.

The sign chart only ranks nearby options. Before locking in a schedule, the manager still checks the index at the shift’s true endpoints, the most stops the trucks can carry and the fewest the contract allows, since an unchecked endpoint can beat both local extremes. Points where the model itself breaks down, such as a stop needing an extra vehicle, belong in the partition too. This is an Independent rule: it directly determines qualitative shape and local extrema from reusable sign information.

4.2.2: Bracket a convex function between tangents and secants

History

A safety margin built on a bound facing the wrong direction is a false promise: a quantity assumed to stay below some ceiling can already have crossed it. Johan Jensen closed that kind of gap, working in Copenhagen in 1905 and 1906 on a systematic account of convex functions and the inequalities between mean values: a convex graph lies below every chord drawn between two of its points, while any supporting tangent line lies below the graph itself. Jensen connected convexity to inequalities between averages explicitly; today’s tangent-secant “bracket” operationalizes that structure to estimate a convex function without evaluating it exactly.

The equation

For differentiable convex ff, a tangent at aa gives

f(x)≥f(a)+f′(a)(x−a). f(x)\ge f(a)+f'(a)(x-a).

For x=(1−t)a+tbx=(1-t)a+tb with 0≤t≤10\le t\le1, the chord gives

f(x)≤(1−t)f(a)+tf(b). f(x)\le(1-t)f(a)+tf(b).

For a concave function, the geometric inequalities reverse.

How to read it

Convexity means slopes never decrease as you move right along the graph. A tangent line uses just one local slope and can never rise above the curve; a chord connecting two known points averages the slopes between them and can never dip below the curve there. Together the two lines become a lower and an upper certificate, with no evaluation needed in between.

The tangent bound travels as far as convexity and differentiability hold; the chord bound only applies between its own two endpoints, since its weight tt must stay between zero and one. What neither line tells you is how tight the gap between them is: a curve that hugs close to straight leaves tangent and chord nearly touching, while a sharply bending curve can leave a wide space between them.

For a point xx between aa and bb, the chord weight is t=(x−a)/(b−a)t=(x-a)/(b-a); moving either the chord’s endpoints or the tangent point closer to xx tightens the corresponding bound, so several known points can build a tighter, piecewise bracket.

How to use it

An epidemiologist facing an outbreak whose case count grows like exe^x over xx weeks needs a fast bound on the growth factor at x=0.4x=0.4 weeks, without waiting on the full exponential model during a briefing. The tangent at zero gives

ex≥1+x, e^x\ge1+x,

so at x=0.4x=0.4, 1+0.4=1.41+0.4=1.4, and growth is at least 1.41.4-fold. The chord through (0,1)(0,1) and (1,e)(1,e) over the first week gives

ex≤1+(e−1)x, e^x\le1+(e-1)x,

so at x=0.4x=0.4, 1+(e−1)(0.4)=1+0.687=1.6871+(e-1)(0.4)=1+0.687=1.687, and growth is at most about 1.6871.687-fold. The true factor, about 1.49181.4918, sits inside that bracket, which the epidemiologist can quote before the exact model finishes running.

The bracket holds only because the case-growth model is convex; the epidemiologist checks that with f″≥0f''\ge0 where available, and would split the timeline into separate windows once curvature changed sign, for instance after interventions started bending the curve down, since one global picture would no longer bound the new regime. For a concave process, such as cases approaching a saturation ceiling, the roles reverse: tangents lie above and chords below. This is an Independent rule: convexity directly supplies portable lower and upper bounds without exact evaluation.

The quadratic stays between the lower tangent 2x-1 and upper secant 2x over the plotted interval.

Figure 4.2. On [0,2], x² lies above its tangent at x=1 and below the secant joining the endpoints. The upper bound is restricted to the interval spanned by that secant.

4.2.3: Estimate a thin shell as boundary measure times thickness

History

An eyeballed estimate of a thin coating drifts further from the true volume the more curved the surface underneath. Around 250 BCE in Syracuse, Archimedes compared planar slices and solid sections through geometric balance arguments, preserved in The Method and related works, using infinitesimal-style decompositions heuristically before supplying geometric proofs acceptable by ancient standards. Archimedes did not write a modern shell differential; his documented slice-and-balance reasoning is the historical practice, and reading a thin boundary layer as boundary measure times thickness is the modern calculus translation of it. What that translation buys is a fast, first-order answer for any thin uniform layer, without rebuilding the whole solid from scratch.

The equation

For a circle of radius rr and small outward thickness drdr,

dA≈2πrdr. dA\approx2\pi r\,dr.

For a sphere,

dV≈4πr2dr. dV\approx4\pi r^2\,dr.

More generally, a sufficiently thin uniform layer of thickness tt along surface area SS has leading volume

ΔV≈St. \Delta V\approx St.

How to read it

The boundary measure is the rate at which interior measure grows when the boundary itself is nudged outward. Differentiating A=πr2A=\pi r^2 gives the circle’s circumference 2πr2\pi r; differentiating V=4πr3/3V=4\pi r^3/3 gives the sphere’s surface area 4πr24\pi r^2. Multiplying that rate by a thickness gives the first-order layer.

What this shortcut leaves out is curvature: the omitted terms grow with how sharply the surface bends, so the estimate is only as good as the thickness is small compared with every relevant curvature radius.

The exact spherical expansion makes those omitted terms explicit:

4π3[(r+t)3−r3]=4πr2t+4πrt2+4π3t3. \frac{4\pi}{3}\bigl[(r+t)^3-r^3\bigr] =4\pi r^2t+4\pi rt^2+\frac{4\pi}{3}t^3.

Surface area times thickness is only the first term; curvature supplies the rest.

A units check confirms the setup before any number is plugged in: boundary area times length gives volume, and circumference times length gives planar area. A mismatch means the wrong boundary measure was used.

How to use it

A manufacturing engineer is specifying a radial allowance around a circular sheet-metal end cap, radius 1010 cm. The allowance extends the cap’s radius by 0.10.1 cm. The quick estimate for the added planar area is

ΔA≈2π(10)(0.1)=2π≈6.283 cm2, \Delta A\approx2\pi(10)(0.1)=2\pi\approx6.283\text{ cm}^2,

while the exact annulus area is

π(10.12−102)=2.01π≈6.315 cm2. \pi(10.1^2-10^2)=2.01\pi\approx6.315\text{ cm}^2.

The missing 0.01π0.01\pi is exactly the quadratic thickness correction, small enough here that the engineer budgets the additional sheet-metal area by the quick estimate and rounds up.

Boundary times thickness works for coatings, machining allowances, and other thin annuli or shells, but it assumes a small, roughly uniform layer and a well-behaved offset. On this same tank, a corner where the cap meets the cylindrical wall, or a spot where two coating passes overlap, would need a separate, exact calculation, since the local formula cannot see those topological changes or self-intersections. This is an Independent rule: it directly estimates thin-layer measure in many geometries without constructing a full volume integral.

4.3: Control Integrals, Series, and Method Choice

Accumulation problems rarely begin with antiderivatives. The ten rules in this section first ask what range, symmetry, endpoint motion, benchmark, or structural pattern can settle the question, or at least choose the correct procedure.

4.3.1: Check an integral average against the function’s range

History

A number that looks reasonable can still be wrong in a way no re-checked arithmetic will catch, unless it is tested against the one thing it cannot violate: the range of the function it came from. Augustin-Louis Cauchy built the tools for that kind of test. Working in Paris, his Cours d’analyse of 1821 and the calculus lessons that followed rebuilt differentiation, integration, and infinite series around limits, using inequalities and comparison to control convergence where earlier methods had often been used only formally. Cauchy treated limits, integrals, and series comparison explicitly; the compact “check the average against the range” wording here selects one of those mechanisms as a fast, first-pass audit.

The equation

For a continuous function on [a,b][a,b] with a<ba<b, define

favg=1b−a∫abf(x)dx. f_{\mathrm{avg}} =\frac{1}{b-a}\int_a^b f(x)\,dx.

If m≤f(x)≤Mm\le f(x)\le M on the interval, then

m≤favg≤M. m\le f_{\mathrm{avg}}\le M.

Continuity also guarantees some c∈[a,b]c\in[a,b] with f(c)=favgf(c)=f_{\mathrm{avg}}.

How to read it

Multiply the inequality m≤f(x)≤Mm\le f(x)\le M through by the interval’s positive length and integrate, and the total signed area gets trapped between m(b−a)m(b-a) and M(b−a)M(b-a); dividing back by b−ab-a returns everything to the function’s own scale. An average landing outside that range cannot be some unusual feature of the curve. It has to be an arithmetic, sign, unit, or normalization mistake. Passing the check is not the same as being correct: a wrong value can still land inside the range and slip through undetected.

The check is about average value, not the raw integral itself: on an interval whose length is not one, those two quantities live on different scales. If xx measures time and ff measures power, the integral carries units of energy; dividing by elapsed time restores power.

How to use it

A city water-utility analyst computes the average flow rate through a meter over a one-hour test as ∫01f(x)dx\int_0^1 f(x)\,dx, where ff, in liters per second, ranges from 00 to 11 across the hour. With f(x)=x2f(x)=x^2,

favg=∫01x2dx=13, f_{\mathrm{avg}} =\int_0^1x^2\,dx =\frac13,

which passes the range check at once, since 1/31/3 sits between 00 and 11. Over a longer two-hour test the same flow model gives an integral of 8/38/3 and an average of

12⋅83=43, \frac12\cdot\frac83=\frac43,

which still lies between the endpoints 00 and 44, so the check passes here too, even though 8/38/3 is the raw integral and not an average; mistaking one for the other would have gone undetected by this same check.

Run this test right after any hand or software integration, before quoting a flow rate as representative of the period. The same range bound holds for a normalized average with nonnegative weights and positive total weight. Signed weights do not guarantee it, and an unbounded monitoring window requires a separately defined averaging procedure. This is an Independent rule: it directly validates the scale of an accumulated result across many applications.

4.3.2: Exploit even and odd symmetry before integrating

History

Mechanical intuition and mathematical proof are not the same thing, and Archimedes kept the two separate. Writing in Syracuse in the third century BCE, in the treatise now surviving as The Method, he set matching cross-sections of a figure on an imagined scale, letting paired pieces on opposite sides of a center cancel or add, then followed that mechanical reasoning with a separate geometric proof built to satisfy ancient standards of rigor. Archimedes never wrote a signed integral or an even-odd identity; the symmetry rule used here is a later expression of that same instinct for weighing matched pieces against each other.

The equation

If f(−x)=f(x)f(-x)=f(x) and the integral exists, then

∫−aaf(x)dx=2∫0af(x)dx. \int_{-a}^{a}f(x)\,dx =2\int_0^a f(x)\,dx.

If f(−x)=−f(x)f(-x)=-f(x), then

∫−aaf(x)dx=0. \int_{-a}^{a}f(x)\,dx=0.

How to read it

Reflecting an even function across the vertical axis leaves it looking unchanged, so its two half-interval pieces reinforce each other when integrated. Rotating an odd function’s graph through the origin flips its sign, so equal contributions on opposite sides cancel in the signed sum.

That zero for an odd function is a statement about net accumulation, not about area: the positive and negative lobes on either side can each be large even while their signed total comes out to nothing.

Any function on a symmetric domain splits into an even part and an odd part,

fe(x)=f(x)+f(−x)2,fo(x)=f(x)−f(−x)2, f_e(x)=\frac{f(x)+f(-x)}2, \qquad f_o(x)=\frac{f(x)-f(-x)}2,

and since the odd piece always integrates to zero, only the even piece’s half-interval integral, doubled, survives:

∫−aaf(x)dx=2∫0afe(x)dx. \int_{-a}^{a}f(x)\,dx =2\int_0^a f_e(x)\,dx.

This extends the shortcut beyond functions that are purely even or purely odd. In numerical work, integrating only half a symmetric interval can nearly halve the computation, and pairing opposite sample points makes an odd component’s cancellation exact instead of leaving it to rounding.

How to use it

A sports biomechanics analyst models the net torque of a golfer’s downswing about its midpoint as g(x)=x3+xcos⁡xg(x)=x^3+x\cos x over x∈[−2,2]x\in[-2,2] seconds, and needs the net rotational impulse over that symmetric window without integrating term by term. Since x3x^3 is odd and cos⁡x\cos x is even, making xcos⁡xx\cos x odd, the integrand is odd on a symmetric interval, so

∫−22(x3+xcosx)dx=0 \int_{-2}^{2}\left(x^3+x\cos x\right)dx=0

without finding any antiderivative: torque in the first half of the swing exactly cancels torque in the second half.

Run this symmetry check before substitution, integration by parts, or numerical quadrature. A net of zero does not mean the torque itself was small: each half can carry a large, opposite torque that a coach still needs broken out, and a shifted window, say [−1,2][-1,2] instead of one centered on the midpoint, would destroy the cancellation entirely. Taking the absolute value of an odd integrand also changes the conclusion, since |f||f| is even where ff is odd, and the integral must actually exist for informal cancellation of two divergent halves to mean anything. This is an Independent rule: on a symmetric domain, parity can directly finish the integral across many settings.

4.3.3: Differentiate an accumulated integral at its moving endpoint

History

An accumulation with no closed formula for its running total can still be differentiated, once the endpoint doing the accumulating is treated as the real variable in play. Leibniz’s publications of 1684 and 1686 made composite change visible through notation, the same symbols that expressed differentiation rearranging to recognize accumulated change in a new variable. Leibniz established that calculus structure; scanning an integrand’s moving endpoint for the rate it is adding, rather than solving for the whole accumulated amount, is a modern reading built on it, with continuity and chain-rule hypotheses made explicit.

The equation

For continuous ff,

ddx∫ag(x)f(t)dt=f(g(x))g′(x). \frac{d}{dx}\int_a^{g(x)}f(t)\,dt =f(g(x))g'(x).

With both limits moving,

ddx∫u(x)v(x)f(t)dt=f(v(x))v′(x)−f(u(x))u′(x). \frac{d}{dx}\int_{u(x)}^{v(x)}f(t)\,dt =f(v(x))v'(x)-f(u(x))u'(x).

How to read it

A small outward shift of the upper endpoint adds a thin slice of accumulation: its height is roughly the integrand’s value there, its width the endpoint’s own displacement, and the chain factor g′(x)g'(x) converts a step in xx into that displacement. A moving lower endpoint removes a slice instead, which is where the minus sign comes from.

What this derivative gives is the accumulation’s rate at a single instant, not the accumulated total itself; finding the running sum still requires the integral, closed form or not. In the simplest case, A(x)=∫axf(t)dtA(x)=\int_a^x f(t)\,dt has A′(x)=f(x)A'(x)=f(x): moving the endpoint one small step adds that much area. A continuous probability density pp gives its cumulative function the same property, F′(x)=p(x)F'(x)=p(x).

How to use it

A simulation analyst studies the accumulated model response F(x)=∫0x2cos⁡(t2)dtF(x)=\int_0^{x^2}\cos(t^2)\,dt, where xx is a dimensionless control parameter and tt is a dimensionless integration coordinate. The integrand is a signed contribution density, so the analyst needs the response’s rate of change with respect to xx. Differentiating the accumulation directly gives

F′(x)=cos⁡((x2)2)⋅2x=2xcos⁡(x4), F'(x) =\cos\left((x^2)^2\right)\cdot2x =2x\cos(x^4),

with the dummy variable tt disappearing only after evaluation at the endpoint; replacing it by xx too soon would obscure the chain. At x=1x=1, F′(1)=2(1)cos⁡(1)≈2(0.5403)=1.0806F'(1)=2(1)\cos(1)\approx2(0.5403)=1.0806, so the accumulated response is locally increasing as the control parameter increases. As a two-endpoint check, if G(x)=∫xx2f(t)dtG(x)=\int_x^{x^2}f(t)\,dt describes accumulation between two moving coordinates, G′(x)=2xf(x2)−f(x)G'(x)=2xf(x^2)-f(x), and each term records one endpoint’s contribution.

This rule works for net infusion rates, distance traveled, or any parameter-dependent accumulation, given the fundamental theorem’s regularity; a discontinuous absorption rate needs a more careful, almost-everywhere version. This is an Independent rule: moving-endpoint structure directly yields a derivative even when the accumulated function cannot be written elementarily.

4.3.4: Memorize the p-integral convergence threshold

History

This threshold traces to one document: Cauchy’s 1821 lecture program, which treated improper integrals and convergent series through explicit comparison and limiting arguments, giving a rigorous setting in which a vanishing term could finally be told apart from a finite accumulated tail. That framework is documented history; treating one power-law exponent as a universal pass or fail line is a present-day shorthand built on it.

The equation

At infinity,

∫1∞x−pdx{converges,p>1,diverges,p≤1. \int_1^\infty x^{-p}\,dx \begin{cases} \text{converges},&p>1,\\ \text{diverges},&p\le1. \end{cases}

Near zero, the direction reverses:

∫01x−pdx converges exactly when p<1. \int_0^1x^{-p}\,dx\text{ converges exactly when }p<1.

How to read it

For p≠1p\ne1, the antiderivative contains x1−p/(1−p)x^{1-p}/(1-p), which settles to a finite value at infinity only when the exponent 1−p1-p is negative, that is, when p>1p>1. At the boundary p=1p=1 itself, the antiderivative is log⁡x\log x, whose unbounded growth creates the harmonic threshold.

An integrand shrinking to zero is not sufficient for its tail to converge: 1/x1/x decays forever and still accumulates infinite area.

For p>1p>1, the leftover area beyond a cutoff BB is explicit,

∫B∞x−pdx=B1−pp−1, \int_B^\infty x^{-p}\,dx =\frac{B^{1-p}}{p-1},

which explains why convergence near p=1p=1 can be painfully slow: the denominator p−1p-1 is small and B1−pB^{1-p} decays weakly. The threshold classifies convergence, while this explicit formula also sizes a pure power-law tail. Logarithmic factors require a separate comparison.

How to use it

An environmental engineer is tracking a contaminant’s concentration downstream of a spill, which decays like x−1.1x^{-1.1} per unit distance xx past the outfall, and needs to know whether total exposure integrated out to infinity is even finite before running a full transport model. Since

∫1∞x−1.1dx=11.1−1=10, \int_1^\infty x^{-1.1}\,dx =\frac{1}{1.1-1}=10,

the tail converges, though to a large finite value, while a nearby model with x−1x^{-1} decay,

∫1∞dxx, \int_1^\infty\frac{dx}{x},

diverges outright: the tiny exponent difference crosses a real qualitative line, and both concentrations trending toward zero is not enough to tell the two cases apart.

The engineer first identifies which end of the domain is improper. Every fixed power is integrable on a finite stretch away from zero; the threshold only bites because one endpoint is singular or infinite, here the infinite downstream limit. A logarithmic correction in the real decay curve could also decide a borderline case where the exponent lands exactly at p=1p=1, so an exponent that close needs a sharper benchmark before the engineer signs off on bounded total exposure. This is an Independent rule: it directly classifies the model power integral and anchors convergence decisions throughout analysis.

4.3.5: Compare positive integrands with a known benchmark

History

A tail declared finite without a real certificate can leave a budget or a warranty resting on nothing. Cauchy grounded calculus in the limit concept, and his rigorous treatment of series and integrals let a positive, ever-shrinking quantity be checked against a known benchmark instead of summed or integrated outright. That comparison-under-limits program is the historical mathematics; the concise direction-check used here is a modern first-pass formulation of it.

The equation

For eventually nonnegative functions on a tail,

0≤f(x)≤g(x),∫g<∞⇒∫f<∞. 0\le f(x)\le g(x), \quad \int g<\infty \quad\Longrightarrow\quad \int f<\infty.

Conversely,

0≤g(x)≤f(x),∫g=∞⇒∫f=∞. 0\le g(x)\le f(x), \quad \int g=\infty \quad\Longrightarrow\quad \int f=\infty.

How to read it

For nonnegative functions, a taller graph accumulates more area: a function that never exceeds a convergent benchmark cannot itself blow up, and a function that never falls below a divergent one cannot itself stay finite. Only the direction that supports the conclusion is useful; running the comparison backward proves nothing.

Only the tail has to obey the inequality: a bounded initial stretch contributes a finite amount regardless and can be set aside. What the comparison never delivers is the integral’s own value, only whether it stays finite, and it works only as well as the benchmark in hand.

How to use it

A logistics network planner is estimating whether the total expected cost of increasingly rare, increasingly long shipment delays stays bounded, modeling the cost density of a delay of length xx days as 1/(x2+1)1/(x^2+1). For x≥1x\ge1,

0≤1x2+1≤1x2, 0\le\frac{1}{x^2+1}\le\frac1{x^2},

and since the p=2p=2 benchmark converges, so does the planner’s tail cost, meaning a fixed reserve can in principle cover it. A rival model with cost density 1/(x+1)1/(x+1) instead gives, for x≥1x\ge1,

1x+1≥12x, \frac1{x+1}\ge\frac1{2x},

a constant multiple of the divergent p=1p=1 benchmark, so that tail cost has no finite bound at all, no matter how large a reserve is set aside.

When the comparison isn’t obvious by inspection, the planner can check the ratio of the candidate density to a benchmark as xx grows large: if it settles on some positive, finite constant, the two tails share the same convergence behavior, though that ratio is only a discovery tool, and the actual inequality still has to be shown to certify the answer. This works only while the cost density stays nonnegative on the tail; a density that can go negative would need an absolute-value or alternating-series treatment instead. This is an Independent rule: it directly transfers a known convergence verdict to broad families of positive integrands.

4.3.6: Bracket a decreasing series tail with integrals

History

Sum a slowly converging series by hand past a few hundred terms, and patience runs out long before the tail does. Cauchy’s rigorous, limit-based treatment of calculus and series turned that stall into an auditable claim, using inequalities and comparison to control convergence instead of demanding every term be summed by hand. That comparison discipline is documented history; the bracket below, pairing positive decreasing terms with neighboring areas, is the modern form of that discipline.

The equation

If ff is positive and decreasing and ak=f(k)a_k=f(k), then the tail after nn satisfies

∫n+1∞f(x)dx≤∑k=n+1∞f(k)≤∫n∞f(x)dx. \int_{n+1}^{\infty}f(x)\,dx \le \sum_{k=n+1}^{\infty}f(k) \le \int_n^{\infty}f(x)\,dx.

The upper integral is a guaranteed truncation-error bound.

How to read it

Picture decreasing rectangles of width one stacked along the curve. A right-endpoint rectangle sits below the curve over the interval just before it, while the matching left-endpoint rectangle sits above the curve over the interval just after. Summing those comparisons traps every unseen rectangle between two neighboring improper integrals.

This bracket sizes an actual remainder, beyond confirming convergence, yet never evaluates the sum itself; if monotonicity lapses even briefly, the picture stops trapping anything. For a pp-series with p>1p>1, the general bracket becomes

(n+1)1−pp−1≤∑k=n+1∞1kp≤n1−pp−1, \frac{(n+1)^{1-p}}{p-1} \le \sum_{k=n+1}^{\infty}\frac1{k^p} \le \frac{n^{1-p}}{p-1},

which turns a target error tolerance directly into a term budget.

How to use it

A financial analyst values the tail of a structured settlement whose remaining payments shrink like 1/k21/k^2 for payment number kk, and needs a bound on the value left after the first 100100 payments without summing hundreds of terms by hand. Since

∫101∞dxx2≤R100≤∫100∞dxx2, \int_{101}^{\infty}\frac{dx}{x^2} \le R_{100}\le \int_{100}^{\infty}\frac{dx}{x^2},

1101≤R100≤1100, \frac1{101}\le R_{100}\le\frac1{100},

the remaining value is about 0.010.01 in the payment’s own units, with the midpoint near 0.00995050.0099505, a bound obtained without summing the full schedule. If a client instead needs the value within 0.0010.001, the upper bound 1/n1/n alone shows that n>1000n>1000 payments already suffice.

The bracket applies only once payments are positive and genuinely decreasing; an irregular first few, negotiated separately, can be pulled out and handled on their own. A stream that oscillates or occasionally increases calls for an alternating-series or absolute-convergence tool, and indexing matters: the first payment left out is n+1n+1, which sets both integral limits. This is a Workflow rule: it controls truncation inside a larger series computation and tells you when to stop.

The exact series tail falls between 1/(N+1) and 1/N; all three decrease with N.

Figure 4.3. For the tail sum of 1/k² after k=N, integrating from N+1 gives the lower bound and from N the upper bound. The exact tail here is pi²/6 minus the finite partial sum.

4.3.7: Look for an inner function and its derivative

History

A composite integral that looks intractable can collapse to one line once its inner expression and that expression’s derivative are both spotted inside it. Leibniz’s published calculus of the mid-1680s let a compound expression be read and rearranged symbol by symbol, rather than treated as a single opaque block. That documented notation is the historical structure; the habit of searching an integral for the chain rule run backward came later, as a workflow reading of it.

The equation

If F′=fF'=f, then

∫f(g(x))g′(x)dx=F(g(x))+C. \int f(g(x))g'(x)\,dx =F(g(x))+C.

With the substitution

u=g(x),du=g′(x)dx, u=g(x), \qquad du=g'(x)\,dx,

the composite integral becomes ∫f(u)du\int f(u)\,du.

How to read it

Substitution runs the chain rule backward: the inner function g(x)g(x) becomes the new coordinate, and g′(x)dxg'(x)\,dx converts a small step in xx into a small step in uu. A constant mismatch is harmless once corrected exactly; a variable mismatch is not.

Spotting the pattern must come before expanding anything: multiplying out a composite power or product can destroy the structure that would have collapsed the integral. What the rule does not do is say which inner function to try when more than one looks plausible; only trial and a differentiation check settle that.

Common signatures include g′(x)/g(x)g'(x)/g(x), suggesting log⁡|g(x)|\log|g(x)|; g′(x)eg(x)g'(x)e^{g(x)}, suggesting eg(x)e^{g(x)}; and g′(x)[g(x)]mg'(x)[g(x)]^m, suggesting a power of gg.

How to use it

A mechanical engineer computing the work done by a variable restoring force during a compression stroke needs ∫2xcos⁡(x2)dx\int2x\cos(x^2)\,dx, where xx is piston displacement in centimeters. Spotting x2x^2 and its derivative 2x2x together, the engineer sets u=x2u=x^2, so

∫2xcos⁡(x2)dx=∫cos⁡udu=sin⁡u+C=sin⁡(x2)+C. \int2x\cos(x^2)\,dx =\int\cos u\,du =\sin u+C =\sin(x^2)+C.

Differentiating the answer confirms both the outer derivative and inner factor return, catching a dropped constant before it reaches a design report.

For a full stroke from x=0x=0 to 11, transforming a related model’s limits via u=1+x2u=1+x^2 turns ∫012x/(1+x2)dx\int_0^1 2x/(1+x^2)\,dx into ∫12du/u=log⁡2\int_1^2du/u=\log2, with no need to substitute back to xx. The scan works for powers, exponentials, logarithms, and rational expressions with a repeated inner form, but a derivative only resembling the needed factor, not matching it up to an exact constant, is not a match. For definite integrals, transform both limits into uu or back-substitute fully, never mixing an xx-limit with a uu-integrand. This is a Workflow rule: it chooses a representation and simplifies one stage of a broader integration problem.

4.3.8: Choose integration by parts to move complexity

History

Is a derivative something you assign, or something already latent in a product, waiting to be transferred? Leibniz’s published calculus treated products, sums, and composites as one shared symbolic system, and rearranging his product rule supplied the structural basis for shifting a derivative from one factor onto another inside an integral. Modern selection mnemonics and the phrase “move complexity” are later workflow guidance, not quotations from Leibniz.

The equation

From d(uv)=udv+vdud(uv)=u\,dv+v\,du,

∫udv=uv−∫vdu. \int u\,dv=uv-\int v\,du.

A common selection order for uu is logarithmic, inverse trigonometric, algebraic, trigonometric, exponential, often remembered as LIATE, but simplification is the real criterion.

How to read it

Integration by parts moves the derivative onto whichever factor you call uu, chosen because its derivative gets simpler, while the factor assigned to dvdv needs a manageable antiderivative of its own. The mnemonic ranks common cases; it does not prove any particular choice optimal.

Sometimes two applications return the original integral. Let I=∫excos⁡xdxI=\int e^x\cos x\,dx. Parts gives I=excos⁡x+∫exsin⁡xdxI=e^x\cos x+\int e^x\sin x\,dx, and a second application turns the new integral into exsin⁡x−Ie^x\sin x-I, so

2I=ex(sin⁡x+cos⁡x), 2I=e^x(\sin x+\cos x),

and I=ex(sin⁡x+cos⁡x)/2+CI=e^x(\sin x+\cos x)/2+C. Recurrence can be exactly the intended simplification, and repeated parts against a polynomial factor is especially predictable: each application lowers its degree until it disappears.

How to use it

Redevelopment pressure along a corridor grows like xexxe^x for distance xx in kilometers from the waterfront toward downtown, and a city planner needs ∫xexdx\int xe^x\,dx to total that pressure over a stretch. Taking u=xu=x and dv=exdxdv=e^x\,dx gives du=dxdu=dx and v=exv=e^x, so

∫xexdx=xex−∫exdx=ex(x−1)+C. \int xe^x\,dx =xe^x-\int e^x\,dx =e^x(x-1)+C.

Differentiating confirms xexxe^x; the reverse assignment would integrate xx while differentiating exe^x, producing a higher-degree polynomial and no simplification.

Over the corridor’s first kilometer,

∫01xexdx=[ex(x−1)]01=0+1=1 \int_0^1xe^x\,dx =\bigl[e^x(x-1)\bigr]_0^1 =0+1=1

unit of pressure, above the 0.50.5 review threshold, so the planner flags that block for review. Parts suits a polynomial paired with an exponential, logarithm, inverse function, or repeated trigonometric-exponential term; a poor choice of uu can reproduce or worsen the integral, which the planner then collects algebraically rather than abandoning it. This is a Workflow rule: it reallocates complexity within an integration procedure rather than independently answering every integral.

4.3.9: Use L’Hôpital only after confirming an indeterminate form

History

A ratio that vanishes on both top and bottom looks unsolvable until each part is differentiated separately, a shortcut whose credited author did not fully derive it himself. In Paris in 1696, the Marquis de l’Hôpital published Analyse des infiniment petits, the first textbook on differential calculus, including the limit rule bearing his name, based on results Johann Bernoulli supplied under their documented arrangement. The book presented the rule for certain vanishing ratios; the modern guardrail, verifying a genuine 0/00/0 or ∞/∞\infty/\infty before differentiating, is the precise operational lesson drawn from it.

The equation

Under L’Hôpital’s rule hypotheses, if a quotient has the indeterminate form 0/00/0 or ∞/∞\infty/\infty and the derivative quotient has a limit, then

limx→af(x)g(x)=limx→af′(x)g′(x). \lim_{x\to a}\frac{f(x)}{g(x)} = \lim_{x\to a}\frac{f'(x)}{g'(x)}.

The statement also has appropriate one-sided and infinite-limit forms.

How to read it

What the theorem transfers is a single limit under specific conditions, not a claim that a quotient and its derivative quotient are equal as functions; the link runs through a mean-value relationship between changes in numerator and denominator. Direct substitution must show indeterminacy first; an already-determined ratio should simply be evaluated as is.

Forms such as 0⋅∞0\cdot\infty, ∞−∞\infty-\infty, and 1∞1^\infty are not quotients of the required type until rewritten as one. For example, lim⁡x→0+xlog⁡x\lim_{x\to0^+}x\log x has form 0⋅(−∞)0\cdot(-\infty); rewritten as log⁡x/(1/x)\log x/(1/x), an (−∞)/∞(-\infty)/\infty quotient, differentiating gives (1/x)/(−1/x2)=−x→0(1/x)/(-1/x^2)=-x\to0. The algebraic rewrite is part of the method, not a cosmetic prelude.

For powers such as 1∞1^\infty, take a logarithm, convert the resulting product into a quotient if needed, evaluate, and exponentiate back, provided the domain permits the logarithm.

How to use it

A calculus tutor watches a student reach for L’Hopital’s rule on every limit in a problem set, including lim⁡x→0(x2+1)/(x+1)\lim_{x\to0}(x^2+1)/(x+1), already the determined form 1/1=11/1=1 by direct substitution; differentiating anyway gives 2x/1→02x/1\to0, the wrong answer, because the rule never applied there. On the actual indeterminate case in the same set,

limx→0ex−1x, \lim_{x\to0}\frac{e^x-1}{x},

substitution gives 0/00/0, and differentiating numerator and denominator correctly yields lim⁡x→0ex/1=1\lim_{x\to0}e^x/1=1.

The tutor’s fix is a habit, not a formula: check substitution first, and reach for the rule only once a limit is genuinely indeterminate, since repeated use needs the hypotheses re-verified at each stage, and domain restrictions or zeros of the denominator’s derivative near the limit can still derail it. This is a Workflow rule, specifically a guardrail that selects a legitimate limit method and blocks a seductive invalid manipulation.

4.3.10: Use the first omitted Taylor term as an error scale

History

A single book title, Methodus incrementorum directa et inversa, is the documented artifact behind this expansion: Brook Taylor published it in London in 1715. Later analysts, especially Joseph-Louis Lagrange in Paris, developed systematic remainder forms separating a finite approximation from what had been left out. Taylor supplied the expansion; Lagrange’s era supplied the controlled remainder language. Treating the next term as an error scale is a modern rule still requiring regularity checks.

The equation

Taylor’s theorem writes

f(a+h)=∑k=0nf(k)(a)k!hk+Rn. f(a+h) =\sum_{k=0}^{n}\frac{f^{(k)}(a)}{k!}h^k+R_n.

If |f(n+1)(x)|≤M|f^{(n+1)}(x)|\le M between aa and a+ha+h, the Lagrange bound gives

|Rn|≤M|h|n+1(n+1)!. |R_n|\le\frac{M|h|^{n+1}}{(n+1)!}.

How to read it

The first omitted term carries the next power of the small increment, making it the natural error scale when derivatives and coefficients behave regularly. That scale is not automatically a bound: a rigorous certificate still needs a derivative bound or, for a suitable alternating decreasing series, the alternating-series remainder theorem.

The rule also helps choose the expansion order: keep terms until the next credible contribution falls below the required tolerance. Choosing nn so that

M|h|n+1(n+1)!≤tolerance \frac{M|h|^{n+1}}{(n+1)!}\le\text{tolerance}

separates two decisions, where to center the expansion, which sets hh and MM, and how many terms to keep.

For an alternating series with decreasing terms, the first omitted term is itself a rigorous bound and predicts the error’s sign too.

How to use it

A financial analyst approximates the growth factor for a 10%10\% continuously compounded return, e0.1e^{0.1}, using terms through the quadratic:

e0.1≈1+0.1+0.122=1.105. e^{0.1}\approx1+0.1+\frac{0.1^2}{2}=1.105.

The first omitted term,

0.136≈0.0001667, \frac{0.1^3}{6}\approx0.0001667,

predicts the error’s scale; the actual error is about 0.00017090.0001709, close but slightly larger than the term itself. A rigorous Lagrange bound, using M=e0.1M=e^{0.1} on [0,0.1][0,0.1], gives an error below about 0.00018420.0001842, the figure the analyst would actually cite in a filing.

Use the first omitted term for order selection, quick error budgeting, and interpreting local approximations from Chapter 1, but verify the expansion point is not near a singularity and that the needed derivatives stay controlled. If successive terms are not shrinking, the next term may be a poor guide rather than a reliable one. This is a Workflow rule: it determines when a finite approximation is adequate inside a larger calculation.

Chapter Synthesis: From Local Information to Reliable Decisions

Calculus becomes portable when its operations are read as statements of control.

A derivative can make a prediction, produce a root correction, convert percentage inputs into percentage outputs, or impose a hard limit on total change. Its sign organizes rises and falls; the ordering of slopes creates convexity brackets; and the derivative of area or volume turns a boundary into the leading measure of a thin layer.

Integrals and series are controlled by equally structural questions. An average must remain inside the function’s range. Symmetry can reinforce or cancel contributions. A moving endpoint turns accumulated area back into a local rate. Power thresholds and positive comparison decide convergence without exact antiderivatives, while adjacent integrals quantify a series tail.

Finally, method rules protect formal manipulation. Substitution reverses the chain rule, integration by parts rearranges the product rule, L’Hôpital requires a verified indeterminate quotient, and Taylor truncation requires evidence about the omitted remainder.

The deeper habit is to separate four questions:

  1. Is the result an estimate, a bound, a classification, or a method choice?
  2. What interval or neighborhood must the hypothesis cover?
  3. What is the first neglected effect, and is it below the decision tolerance?
  4. Does local information genuinely control the larger object?

Those questions turn symbolic technique into mathematical judgment.

One-Page Calculus Toolkit

Recognition cue Rule to try What it gives Role
Nearby easy input value Use the tangent line Local numerical estimate Independent
Approximate nonlinear root Take the tangent-line root Rapid correction step Workflow
Product of powers with small percentage changes Add exponent-weighted relative changes Sensitivity estimate Independent
Derivative bounded on a whole interval Multiply slope bound by input distance Hard change bound Independent
Need rises, falls, or extrema Make a derivative sign chart Qualitative shape and candidates Independent
Convex or concave function Use tangents and chords Two-sided function bounds Independent
Thin uniform geometric layer Multiply boundary measure by thickness Leading area or volume change Independent
Integral average Compare with integrand range Scale audit Independent
Symmetric interval Test even or odd parity Doubling or cancellation Independent
Integral with moving endpoint Evaluate at each endpoint and multiply by its derivative; subtract the lower contribution Accumulation derivative Independent
Power-law improper integral Compare exponent with one Convergence threshold Independent
Positive tail resembling a benchmark Compare above or below in the useful direction Convergence transfer Independent
Positive decreasing series tail Use neighboring improper integrals Remainder bracket Workflow
Composite integrand plus inner derivative Substitute the inner expression Chain-rule reversal Workflow
Product with a simplifying derivative Integrate by parts Complexity transfer Workflow
Quotient limit giving 0/00/0 or ∞/∞\infty/\infty Check L’Hôpital’s hypotheses Valid derivative-quotient route Workflow guardrail
Truncated Taylor expansion Inspect and bound the next contribution Error scale or certificate Workflow

Decision Path

Transfer Problems

1. Predict, then certify

Estimate e0.03e^{0.03} with a linear model. Then use a Taylor remainder bound to certify the error and decide whether the approximation is trustworthy to three decimal places.

2. Read a graph without plotting it

For f(x)=x4−4x2f(x)=x^4-4x^2 on [−3,3][-3,3], build a derivative sign chart, locate local and absolute extrema, and identify intervals where a tangent–secant convexity bracket can be used without crossing an inflection point.

3. Choose structure before procedure

For each object, name the first calculus rule you would use and justify its hypotheses:

∫−22xex2dx,∫1∞dxx2+x,∑k=101∞1k2. \int_{-2}^{2}xe^{x^2}\,dx, \qquad \int_1^\infty\frac{dx}{x^2+x}, \qquad \sum_{k=101}^{\infty}\frac1{k^2}.

For the first, compare symmetry with substitution; for the other two, give a convergence argument or numerical tail bracket without seeking unnecessary exact forms.

Where These Ideas Reappear

Historical Notes and Sources

Sources for the historical accounts in this chapter follow. Every calculus rule has a verified historical connection; the concise decision language and modern hypotheses remain later interpretations.