Module 1 of 5 · Lesson 2 of 3
The Derivative
A limit of difference quotients, and the three places it fails to exist.
What you will be able to do
Given a function built from powers, products, quotients and compositions, the learner can compute its derivative by selecting the rule its structure demands, evaluate a derivative from the limit definition when asked, decide where a function fails to be differentiable and say which failure it is, and read the derivative as a slope, a rate and a linear approximation.
Orientation
A rate at an instant
Average speed over an interval is a division: distance by time. Speed at an instant is the same division with a zero denominator, which is not a calculation at all.
The derivative is the resolution. Rather than dividing at the instant, it watches the average over shorter and shorter intervals and takes the value those averages approach. For
What makes this work rather than merely sound plausible is that the quotient simplifies before any limit is taken. The
This unit establishes the definition, derives the rules for powers, products, quotients and compositions as consequences of it, and identifies the three ways the limit can fail, because knowing where a derivative does not exist is as much a part of using one as computing it.
Intuition
The line that fits best
The canonical intuition gives two readings, rate and linearisation. The second deserves working out numerically, because it is what makes the derivative useful rather than merely defined.
How good is the tangent? For
| error | |||
|---|---|---|---|
The error is exactly
Why that is the right standard. Any line through
Where the two readings meet. Saying
The limitation is built in. The approximation is local. At
Simulation
Secant slopes approaching the tangent slope at x = 3
Drag
From the right,
The secant never becomes the tangent — at
Watch what stays fixed: the point
Definition
What the definition requires, term by term
The canonical definition states the limit and the rules. What follows is why each piece is phrased as it is.
Why a limit and not a value. At
Why both one-sided limits must agree.
Why differentiability implies continuity, and not the reverse. The implication is one line:
Why the product rule has two terms. Write the change in
Dividing by
Why the chain rule multiplies. If
Why the quotient rule has
Derivation
The power rule from the definition
The case
The cancellation is exact, and
The general positive integer case. Expand
Subtracting
Every term after the first still contains
In the argument, the first-order term of the expansion is the derivative, and everything of higher order is discarded by the limit. That is the same separation the tangent-line approximation makes.
Verification at three points.
It assumes
Example
Derivatives of the functions that keep appearing
Each of these is worth knowing on sight, and each says something the formula alone does not.
A constant,
A line,
A square,
A reciprocal,
A root,
A cubic with a flat spot,
Four of the six are instances of the single power rule, with exponents
Worked example
Four rules, and a curve's shape
1. From the definition:
The limit as
2. Product rule:
Check by simplifying first:
What the wrong rule gives:
3. Quotient rule:
With
At
Check numerically: the symmetric difference quotient at
4. Chain rule:
Outer function
At
Check numerically: the symmetric difference quotient returns. Omitting the factor
5. Reading a curve's shape:
Both are critical points. Classify with
| verdict | |||
|---|---|---|---|
| local maximum | |||
| local minimum |
Confirming by the sign of
Neither is a global extremum:
Procedure
Choosing the rule
Differentiation is mechanical once the function's outermost structure is identified. The error is almost never in applying a rule; it is in applying the wrong one because the structure was misread.
Step 1 — simplify if simplification is free.
Step 2 — identify the outermost operation. Ask what is done last when evaluating the expression at a number:
| Outermost operation | Rule |
|---|---|
| a sum or difference | differentiate term by term |
| a constant times something | pull the constant out |
| two functions multiplied | product rule |
| one function divided by another | quotient rule |
| a function applied to an expression | chain rule |
| a bare power of | power rule |
Step 3 — apply the rule, leaving inner derivatives unexpanded at first. Write
Step 4 — evaluate inner derivatives, then simplify.
Step 5 — check. Three cheap checks, in decreasing order of availability:
- Independent route. If the function can be simplified into another form, differentiate that too.
by the product rule must match . - Numerical spot check. Evaluate
at and compare. This catches sign errors and dropped factors immediately. - Degree check. Differentiating a polynomial of degree
gives degree . A derivative that gained degree is wrong.
Nested compositions. The chain rule applies repeatedly, outermost first. For
The common failures. Dropping the inner derivative in the chain rule. Reversing the numerator of the quotient rule, which flips the sign. And using
Non-example
Where the derivative fails to exist
Three failures of the limit.
A corner:
A vertical tangent:
A discontinuity. Any
Rules that do not hold.
The chain rule's inner factor is not optional. For
Inferences the derivative does not license.
A local extremum is not a global one.
Optional enrichment (1)
Application
What a derivative is used for
Optimisation. Finding a maximum or minimum starts by solving
This is the single-variable case of what the quadratic forms unit does in several variables, where the Hessian's definiteness replaces the sign of
Linear approximation and error propagation. Near
Newton's method. To solve
The method converges quadratically near a simple root, because the tangent's error is second order. The same
Rates in applied settings. Velocity is the derivative of position and acceleration the derivative of velocity. Marginal cost in economics is the derivative of total cost, and the standard rule "produce while marginal revenue exceeds marginal cost" is the first-order condition for maximising profit. A reaction rate in chemistry and a growth rate in population models are derivatives of concentration and of population.
In statistics and optimisation. Maximum likelihood estimation sets the derivative of the log-likelihood to zero, which is why the estimator's variance involves the second derivative through the Fisher information. Gradient-based optimisation, including every method used to fit a modern statistical model, is the multivariable extension of "move in the direction the derivative indicates".
The unifying point. Each of these replaces a hard nonlinear question with an easy linear one that is accurate near a point. That substitution is what the derivative licenses, and the quadratic error term is what limits how far from the point the answer can be trusted.