Module 1 of 5 · Lesson 2 of 3

The Derivative

A limit of difference quotients, and the three places it fails to exist.

What you will be able to do

Given a function built from powers, products, quotients and compositions, the learner can compute its derivative by selecting the rule its structure demands, evaluate a derivative from the limit definition when asked, decide where a function fails to be differentiable and say which failure it is, and read the derivative as a slope, a rate and a linear approximation.

Orientation

A rate at an instant

Average speed over an interval is a division: distance by time. Speed at an instant is the same division with a zero denominator, which is not a calculation at all.

The derivative is the resolution. Rather than dividing at the instant, it watches the average over shorter and shorter intervals and takes the value those averages approach. For f ( x ) = x 2 at x = 3 , the average rates over intervals of width 1 , 0.5 , 0.1 and 0.01 are 7 , 6.5 , 6.1 and 6.01 , heading for 6, which the algebra confirms exactly.

What makes this work rather than merely sound plausible is that the quotient simplifies before any limit is taken. The h in the denominator cancels, leaving 6 + h , and the limit of that is not in doubt.

This unit establishes the definition, derives the rules for powers, products, quotients and compositions as consequences of it, and identifies the three ways the limit can fail, because knowing where a derivative does not exist is as much a part of using one as computing it.

Intuition

The line that fits best

The canonical intuition gives two readings, rate and linearisation. The second deserves working out numerically, because it is what makes the derivative useful rather than merely defined.

How good is the tangent? For f ( x ) = x 2 at a = 3 , the tangent is L ( x ) = 9 + 6 ( x − 3 ) = 6 x − 9 . Comparing:

x f ( x ) = x 2 L ( x ) = 6 x − 9 error
2.5 6.25 6.0 0.25
2.9 8.41 8.4 0.01
3.1 9.61 9.6 0.01
3.5 12.25 12.0 0.25

The error is exactly ( x − 3 ) 2 , not approximately, exactly, since x 2 − ( 6 x − 9 ) = ( x − 3 ) 2 . Halving the distance from 3 quarters the error, which is the signature of a first-order approximation.

Why that is the right standard. Any line through ( 3 , 9 ) agrees with f at the point. What distinguishes the tangent is that its error vanishes faster than linearly: a line with the wrong slope has error growing like a constant times x − 3 , which dominates ( x − 3 ) 2 near 3. So the tangent is not merely a good line, it is the unique line whose error is o ( x − a ) , and the derivative is its slope.

Where the two readings meet. Saying f ′ ( 3 ) = 6 says both that f is increasing at 6 units of output per unit of input at that instant, and that near x = 3 the function is indistinguishable from 6 x − 9 to first order. Every linear approximation in applied work, Newton's method, error propagation, marginal analysis in economics, is one of these two sentences applied somewhere.

The limitation is built in. The approximation is local. At x = 3.5 the error is already 0.25 , about 2% of the value, and it grows quadratically from there. A derivative tells you about the neighbourhood of a point and nothing about behaviour far away, which is exactly why a single derivative cannot certify a global maximum.

Simulation

Secant slopes approaching the tangent slope at x = 3

the difference quotient is 6 + h, from either side

Drag h through zero, from either side. The red line joins ( 3 , 9 ) to ( 3 + h , ( 3 + h ) 2 ) , so its slope is the difference quotient

( 3 + h ) 2 − 9 h = 6 + h .

From the right, h = 1 gives 7 , h = 0.5 gives 6.5 , h = 0.1 gives 6.1 . From the left, h = − 1 gives 5 , h = − 0.5 gives 5.5 , h = − 0.1 gives 5.9 . Both sides approach 6 , and that agreement is what the two-sided limit requires: the derivative exists only when the left and right approaches give the same number.

The secant never becomes the tangent — at h = 0 there is no second point and no line through one point — but the slopes approach 6 , and that limit is what f ′ ( 3 ) names.

Watch what stays fixed: the point ( 3 , 9 ) , which every secant passes through. The derivative is a statement about behaviour arbitrarily close to that point, from both directions, not at any particular h . Later, at a corner, the two sides disagree and no derivative exists, which the same control would show.

Definition

What the definition requires, term by term

The canonical definition states the limit and the rules. What follows is why each piece is phrased as it is.

Why a limit and not a value. At h = 0 the quotient reads 0 / 0 , which is not a number. The limit asks a different question, what the quotient approaches, and that question has an answer whenever the numerator vanishes at the same rate as the denominator. The cancellation of h is not a trick for avoiding the division; it is the demonstration that the rates match.

Why both one-sided limits must agree. lim h → 0 means h approaching from either side. Requiring agreement is what makes the tangent well defined: at a corner the secants from the left and right settle on different lines, and there is no single slope to report.

Why differentiability implies continuity, and not the reverse. The implication is one line: f ( a + h ) − f ( a ) = h ⋅ f ( a + h ) − f ( a ) h , and as h → 0 the right side tends to 0 ⋅ f ′ ( a ) = 0 . The reverse fails because continuity constrains only the size of the change, while differentiability constrains its ratio to h . A strictly stronger demand, and | x | satisfies the first and not the second.

Why the product rule has two terms. Write the change in f g over a step h and add and subtract one cross term:

f ( a + h ) g ( a + h ) − f ( a ) g ( a ) = [ f ( a + h ) − f ( a ) ] g ( a + h ) ⏟ f changes + f ( a ) [ g ( a + h ) − g ( a ) ] ⏟ g changes .

Dividing by h and letting h → 0 gives f ′ ( a ) g ( a ) + f ( a ) g ′ ( a ) , using continuity of g to send g ( a + h ) → g ( a ) . Two terms appear because there are two ways the product can change, and neither is f ′ g ′ , that expression counts the contribution of both factors changing simultaneously, which is second order and vanishes in the limit.

Why the chain rule multiplies. If g changes at rate g ′ ( x ) and f changes at rate f ′ per unit of its input, then composing scales one rate by the other. The factor g ′ ( x ) is the conversion between the two input scales, and omitting it is the most common error in the rule.

Why the quotient rule has g 2 underneath. It follows from the product rule applied to f ⋅ g − 1 together with the chain rule on g − 1 , whose derivative is − g ′ / g 2 . The minus sign in the numerator is inherited from there, which is why the order f ′ g − f g ′ matters and the reversed version is wrong by a sign.

Derivation

The power rule from the definition

The case f ( x ) = x 2 . From the definition at a general point a :

( a + h ) 2 − a 2 h = a 2 + 2 a h + h 2 − a 2 h = 2 a h + h 2 h = 2 a + h .

The cancellation is exact, and lim h → 0 ( 2 a + h ) = 2 a . So f ′ ( a ) = 2 a , giving f ′ ( 3 ) = 6 as the numerical table suggested. Checking at a = 3 , h = 0.01 : the quotient is 6.01 and 2 a + h = 6.01 .

The general positive integer case. Expand ( a + h ) n by the binomial theorem:

( a + h ) n = a n + n a n − 1 h + ( n 2 ) a n − 2 h 2 + ⋯ + h n .

Subtracting a n removes the leading term, and every surviving term carries at least one factor of h :

( a + h ) n − a n h = n a n − 1 + ( n 2 ) a n − 2 h + ⋯ + h n − 1 .

Every term after the first still contains h , so all of them vanish as h → 0 , leaving

f ′ ( a ) = n a n − 1 .

In the argument, the first-order term of the expansion is the derivative, and everything of higher order is discarded by the limit. That is the same separation the tangent-line approximation makes.

Verification at three points. d d x x 3 at x = 2 gives 3 ⋅ 4 = 12 ; d d x x 5 at x = 1 gives 5 ⋅ 1 = 5 ; d d x x − 2 at x = 4 gives − 2 ⋅ 4 − 3 = − 2 / 64 = − 1 / 32 . Symmetric difference quotients at h = 10 − 6 return 12.00000000 , 5.00000000 and − 0.03125000 , and − 1 / 32 = − 0.03125 .

It assumes n is a positive integer, since otherwise ( a + h ) n has no finite expansion. The rule nonetheless holds for every real exponent, the − 2 case above is an instance, but establishing that requires either the quotient rule for negative integers, implicit differentiation for rationals, or the logarithmic route for arbitrary reals. Applying the rule to x 1 / 3 or x π is legitimate; claiming the binomial proof covers them is not.

Example

Derivatives of the functions that keep appearing

Each of these is worth knowing on sight, and each says something the formula alone does not.

A constant, f ( x ) = 7 . The difference quotient is 7 − 7 h = 0 for every h , so f ′ ( x ) = 0 everywhere. A constant function has no rate of change, which is the base case the sum rule leans on whenever a constant term is dropped.

A line, f ( x ) = 3 x + 2 . The quotient is 3 ( x + h ) + 2 − 3 x − 2 h = 3 h h = 3 , with no limit needed. It is already free of h . The derivative is the slope, constant everywhere, and the tangent to a line is the line itself. This is the case where the linear approximation is exact rather than merely good nearby.

A square, f ( x ) = x 2 . f ′ ( x ) = 2 x . The derivative is negative for x < 0 , zero at the vertex, positive for x > 0 . The parabola falling, levelling, then rising. Reading the sign of f ′ off the graph's shape, and the shape off the sign, is the habit this unit is building.

A reciprocal, f ( x ) = 1 / x = x − 1 . The power rule gives f ′ ( x ) = − x − 2 = − 1 / x 2 , negative for every x ≠ 0 : the function decreases on both branches. At x = 4 , f ′ ( 4 ) = − 1 / 16 , a gentle slope; at x = 0.1 it is − 100 , extremely steep. The derivative grows without bound near the origin, which is the analytic form of the vertical asymptote.

A root, f ( x ) = x = x 1 / 2 . The power rule applies with a fractional exponent: f ′ ( x ) = 1 2 x − 1 / 2 = 1 2 x . At x = 4 that is 1 4 ; at x = 100 it is 1 20 . The slope decreases as x grows, the curve flattens, and as x → 0 + it diverges, giving the vertical tangent at the origin. Note that f is defined at 0 but not differentiable there.

A cubic with a flat spot, f ( x ) = x 3 . f ′ ( x ) = 3 x 2 ≥ 0 everywhere, and zero only at x = 0 . The function is increasing throughout, yet its tangent at the origin is horizontal. This is the standard counterexample to the belief that a vanishing derivative signals a turning point.

Four of the six are instances of the single power rule, with exponents 0 , 1 , 2 , − 1 , 1 2 and 3 . One formula covering constants, lines, curves, reciprocals and roots. The two that behave unusually, 1 / x near 0 and x at 0, are unusual for the same reason: the derivative diverges where the graph turns vertical.

Worked example

Four rules, and a curve's shape

1. From the definition: f ( x ) = x 2 at a = 3 .

f ( 3 + h ) − f ( 3 ) h = ( 3 + h ) 2 − 9 h = 9 + 6 h + h 2 − 9 h = 6 h + h 2 h = 6 + h .

The limit as h → 0 is 6 . Note the order of operations: cancel first, then take the limit. Substituting h = 0 before cancelling gives 0 / 0 and no information.

2. Product rule: f ( x ) = x 2 , g ( x ) = x 3 , at x = 2 .

( f g ) ′ ( 2 ) = f ′ ( 2 ) g ( 2 ) + f ( 2 ) g ′ ( 2 ) = 4 ⋅ 8 + 4 ⋅ 12 = 32 + 48 = 80 .

Check by simplifying first: f g = x 5 , so ( f g ) ′ = 5 x 4 and at x = 2 that is 5 ⋅ 16 = 80 .

What the wrong rule gives: f ′ ( 2 ) g ′ ( 2 ) = 4 ⋅ 12 = 48 , which is not 80. The two agree only in contrived cases, and the availability of the independent check makes this the easiest rule to self-verify.

3. Quotient rule: y = x 2 + 1 x − 1 at x = 3 .

With u = x 2 + 1 , u ′ = 2 x , v = x − 1 , v ′ = 1 :

y ′ = u ′ v − u v ′ v 2 = 2 x ( x − 1 ) − ( x 2 + 1 ) ( x − 1 ) 2 .

At x = 3 : numerator = 6 ⋅ 2 − 10 = 2 , denominator = 4 , so y ′ ( 3 ) = 2 4 = 1 2 .

Check numerically: the symmetric difference quotient at h = 10 − 7 returns.

4. Chain rule: y = ( 3 x 2 + 1 ) 4 at x = 1 .

Outer function u 4 , inner function u = 3 x 2 + 1 with u ′ = 6 x :

y ′ = 4 u 3 ⋅ u ′ = 4 ( 3 x 2 + 1 ) 3 ⋅ 6 x .

At x = 1 : u = 4 , so y ′ = 4 ⋅ 64 ⋅ 6 = 1536 .

Check numerically: the symmetric difference quotient returns. Omitting the factor u ′ = 6 would give 256, which is the standard chain-rule error.

5. Reading a curve's shape: f ( x ) = x 3 − 3 x .

f ′ ( x ) = 3 x 2 − 3 = 3 ( x 2 − 1 ) = 0 ⟹ x = ± 1 .

Both are critical points. Classify with f ″ ( x ) = 6 x :

x f ( x ) f ″ ( x ) verdict
− 1 2 − 6 local maximum
+ 1 − 2 + 6 local minimum

Confirming by the sign of f ′ : f ′ ( − 2 ) = 9 > 0 (increasing), f ′ ( 0 ) = − 3 < 0 (decreasing), f ′ ( 2 ) = 9 > 0 (increasing). Rising then falling at x = − 1 is a maximum; falling then rising at x = + 1 is a minimum.

Neither is a global extremum: x 3 − 3 x is unbounded in both directions. A derivative reports local behaviour only, which is the limitation the intuition block's error analysis already showed.

Procedure

Choosing the rule

Differentiation is mechanical once the function's outermost structure is identified. The error is almost never in applying a rule; it is in applying the wrong one because the structure was misread.

Step 1 — simplify if simplification is free. x 2 ⋅ x 3 is x 5 ; differentiate that directly rather than invoking the product rule. Likewise x 3 + x x = x 2 + 1 needs no quotient rule. This step removes most opportunities for error.

Step 2 — identify the outermost operation. Ask what is done last when evaluating the expression at a number:

Outermost operationRule
a sum or differencedifferentiate term by term
a constant times somethingpull the constant out
two functions multipliedproduct rule
one function divided by anotherquotient rule
a function applied to an expressionchain rule
a bare power of x power rule

Step 3 — apply the rule, leaving inner derivatives unexpanded at first. Write 4 ( 3 x 2 + 1 ) 3 ⋅ d d x ( 3 x 2 + 1 ) before evaluating the second factor. Writing the structure first is what prevents the dropped inner derivative.

Step 4 — evaluate inner derivatives, then simplify.

Step 5 — check. Three cheap checks, in decreasing order of availability:

  • Independent route. If the function can be simplified into another form, differentiate that too. x 2 ⋅ x 3 by the product rule must match 5 x 4 .
  • Numerical spot check. Evaluate f ( a + h ) − f ( a − h ) 2 h at h = 10 − 6 and compare. This catches sign errors and dropped factors immediately.
  • Degree check. Differentiating a polynomial of degree n gives degree n − 1 . A derivative that gained degree is wrong.

Nested compositions. The chain rule applies repeatedly, outermost first. For ( x 2 + 1 ) 3 , the outermost operation is the square root, then the cube, then the sum, so the derivative carries three factors, and each layer contributes one.

The common failures. Dropping the inner derivative in the chain rule. Reversing the numerator of the quotient rule, which flips the sign. And using f ′ g ′ for a product, which the non-example block treats in detail.

Non-example

Where the derivative fails to exist

Three failures of the limit.

A corner: f ( x ) = | x | at x = 0 . The difference quotient is | h | h , which equals + 1 for every h > 0 and − 1 for every h < 0 . Both one-sided limits exist and they disagree, so the two-sided limit does not. The function is continuous at 0. This is the standard demonstration that continuity does not imply differentiability.

A vertical tangent: f ( x ) = x 1 / 3 at x = 0 . The quotient is h 1 / 3 h = h − 2 / 3 , which grows without bound: 100 at h = 10 − 3 , 10,000 at h = 10 − 6 . Here the one-sided limits agree, both are + ∞ , but agreeing on ∞ is not having a limit, since no real number is approached. The tangent line exists geometrically and is vertical, which has no finite slope.

A discontinuity. Any f discontinuous at a fails there, by the contrapositive of "differentiable implies continuous". A jump is the clearest case: the numerator f ( a + h ) − f ( a ) does not tend to 0, so the quotient diverges.

Rules that do not hold.

( f g ) ′ ≠ f ′ g ′ . With f = x 2 and g = x 3 at x = 2 : the product rule gives 4 ⋅ 8 + 4 ⋅ 12 = 80 , and the simplification f g = x 5 confirms 5 ⋅ 16 = 80 . The false rule gives 4 ⋅ 12 = 48 . The error is not small and does not vanish for large x .

( f g ) ′ ≠ f ′ g ′ . For x 2 + 1 x − 1 at x = 3 the quotient rule gives 1 2 , while f ′ g ′ = 6 1 = 6 , wrong by a factor of twelve.

The chain rule's inner factor is not optional. For ( 3 x 2 + 1 ) 4 at x = 1 , the correct value is 4 ⋅ 64 ⋅ 6 = 1536 ; omitting u ′ = 6 x gives 256.

Inferences the derivative does not license.

f ′ ( c ) = 0 does not make c an extremum. f ( x ) = x 3 has f ′ ( 0 ) = 0 , yet x 3 is strictly increasing everywhere and 0 is neither a maximum nor a minimum. It is an inflection point with a horizontal tangent.

f ″ ( c ) = 0 decides nothing. At x = 0 the functions x 3 , x 4 and − x 4 all have vanishing first and second derivatives, and have respectively no extremum, a minimum and a maximum. The second-derivative test is silent here, and the sign of f ′ on either side must be used instead.

A local extremum is not a global one. x 3 − 3 x has a local maximum at x = − 1 with value 2, while the function exceeds 2 for all x > 2 , indeed f ( 3 ) = 18 . Nothing about a derivative at a point constrains behaviour far from it.

Optional enrichment (1)

Application

What a derivative is used for

Optimisation. Finding a maximum or minimum starts by solving f ′ ( x ) = 0 , because an interior extremum of a differentiable function must occur at a critical point. The condition is necessary and not sufficient, so candidates must be classified, by the sign of f ′ around each, or by f ″ where it is nonzero, and endpoints checked separately, since an extremum on a closed interval may sit where the derivative is not zero at all.

This is the single-variable case of what the quadratic forms unit does in several variables, where the Hessian's definiteness replaces the sign of f ″ .

Linear approximation and error propagation. Near a , f ( x ) ≈ f ( a ) + f ′ ( a ) ( x − a ) . If a measured input carries uncertainty Δ x , the induced uncertainty in the output is approximately | f ′ ( a ) | Δ x , so the derivative is the amplification factor for measurement error. A large derivative means a sensitive quantity, which is why the condition number of a numerical problem is built from derivatives.

Newton's method. To solve f ( x ) = 0 , replace f by its tangent at the current guess and solve the linear equation instead:

x n + 1 = x n − f ( x n ) f ′ ( x n ) .

The method converges quadratically near a simple root, because the tangent's error is second order. The same ( x − a ) 2 behaviour the intuition block measured. It fails where f ′ ( x n ) is near zero, since a nearly horizontal tangent meets the axis far away.

Rates in applied settings. Velocity is the derivative of position and acceleration the derivative of velocity. Marginal cost in economics is the derivative of total cost, and the standard rule "produce while marginal revenue exceeds marginal cost" is the first-order condition for maximising profit. A reaction rate in chemistry and a growth rate in population models are derivatives of concentration and of population.

In statistics and optimisation. Maximum likelihood estimation sets the derivative of the log-likelihood to zero, which is why the estimator's variance involves the second derivative through the Fisher information. Gradient-based optimisation, including every method used to fit a modern statistical model, is the multivariable extension of "move in the direction the derivative indicates".

The unifying point. Each of these replaces a hard nonlinear question with an easy linear one that is accurate near a point. That substitution is what the derivative licenses, and the quadratic error term is what limits how far from the point the answer can be trusted.

Next step

Practice The Derivative

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to L'Hôpital's Rule

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.