Module 4 of 5 · Lesson 1 of 2

Partial Derivatives, the Gradient and Critical Points

Partial derivatives, the gradient, and the saddle that one variable cannot produce.

What you will be able to do

Given a function of two variables, the learner can compute its partial derivatives, assemble the gradient, evaluate a directional derivative along a normalised direction, identify the direction of steepest increase and the level-curve direction, and locate and classify critical points.

Orientation

A surface has no single slope

Every derivative so far answered one question: how fast does f change as x changes. With two inputs the question is incomplete, because there is no longer one way to move.

Standing on a hillside, walking east might climb steeply, walking north gently, walking north-east somewhere between, and walking along the contour not at all. All four are rates of change of the same function at the same point. Asking for "the" slope has no answer.

The resolution is to fix a direction first. The partial derivatives answer for the two coordinate directions, differentiate in x while holding y , and conversely, and each is an ordinary single-variable derivative, so no new technique is needed to compute them.

What is new is that those two numbers determine all the others. For f ( x , y ) = x 2 y + 3 y 2 at ( 2 , 1 ) the partials are 4 and 10, and the rate in the direction ( 3 , 4 ) / 5 is 4 ( 0.6 ) + 10 ( 0.8 ) = 10.4 , obtained by a dot product, with no further limit taken. Packaging the partials as the gradient ( 4 , 10 ) makes that computation the definition rather than a trick.

The gradient also answers the geometric questions at once: it points the steepest way up, its length | ∇ f | ≈ 10.77 is that steepest rate, and directions perpendicular to it give zero, which is what a contour line is.

One genuinely new phenomenon appears. In one variable a vanishing derivative leaves a maximum, a minimum or an inflection. In two, a point can rise along one axis and fall along another: a saddle, with no single-variable counterpart, and the reason a matrix rather than a single number decides the classification.

Definition

Why each definition is shaped as it is

The canonical definition states the partials, the gradient and the directional derivative. What follows is why each carries its particular conditions.

Why ∂ rather than d . The symbol records that something is being held fixed. ∂ f ∂ x is a rate along one axis with the other inputs frozen; d f d x would claim f depends on x alone. The distinction matters once a chain of dependencies appears, where x and y may themselves vary together.

Why the directional derivative needs a unit vector. D u f = ∇ f ⋅ u measures change per unit distance travelled. Using an unnormalised v multiplies the answer by | v | : for the gradient ( 4 , 10 ) and direction ( 3 , 4 ) , the dot product gives 40, while the correct rate along that direction is 10.4 . A factor of exactly | ( 3 , 4 ) | = 5 too large. The number 40 is not a rate in any direction; it is a rate scaled by an arbitrary choice of vector length.

Why the gradient points steepest uphill. From ∇ f ⋅ u = | ∇ f | | u | cos ⁡ θ with | u | = 1 :

θ D u f meaning
0 | ∇ f | steepest increase
π / 2 0 along a level curve
π − | ∇ f | steepest decrease

None of this is extra geometry. It is the cosine, read off. The gradient's length being the maximum rate is the same statement.

Why Clairaut's theorem needs continuity. The proof compares two ways of taking a second difference over a small rectangle, and shows both tend to the same limit. Continuity of the mixed partials is what makes those limits agree. Without it the two orders can genuinely differ at a point, which is why the theorem is stated with a hypothesis rather than as an identity.

Why partial derivatives existing is weaker than differentiability. In one variable, f ′ ( a ) existing forces continuity at a . In two, both partials can exist at a point where f is not even continuous. They only probe two lines through the point, and a function can behave badly along every other direction. Differentiability in several variables demands a genuine linear approximation in all directions, which is a stronger requirement than the existence of the two partials.

Why the Hessian rather than one second derivative. A critical point needs classifying against every direction. The second-order behaviour along direction u is u T H u , so the question is whether that quadratic form is always positive, always negative, or takes both signs, which is exactly definiteness. A single number cannot express a condition over all directions, and the sign-changing case is the saddle.

Intuition

Two numbers, and what they do not determine

The gradient's claim is that two numbers settle every directional rate at a point. Worth testing how far that goes.

What it does determine. At ( 2 , 1 ) for f = x 2 y + 3 y 2 , knowing f x = 4 and f y = 10 gives the rate in any direction by one dot product, no further limits. Sixty-two thousand directions scanned at a hundredth of a degree produce no rate above | ∇ f | = 10.770 , and the perpendicular direction returns zero to machine precision. The two numbers genuinely exhaust the first-order information.

What it does not determine: whether the function is differentiable at all. Both partials probe exactly two lines through the point. A function can be well-behaved along both axes and discontinuous along every other direction, so f x and f y can exist where f is not even continuous. In one variable this cannot happen, because there is only one direction to check and the derivative existing forces continuity. The gradient assembled from two existing partials is therefore a vector that may not describe any linear approximation.

What it does not determine: the classification of a critical point. ∇ f = 0 says the surface is level and stops there. The paraboloid, the inverted bowl and the saddle all satisfy it, and separating them needs second-order information in every direction, which is what the quadratic form u T H u supplies and a single second derivative cannot.

The pattern. Each step upward in dimension costs a condition that was free before. One variable: the derivative existing gives continuity, and one second derivative classifies. Two: neither holds, and both repairs, differentiability as a genuine linear approximation, definiteness as a statement over all directions, are demands quantified over directions rather than checks at a point.

Representation

A function as a surface and as a gradient field

The level curves of f and the gradient arrows crossing them

Take f ( x , y ) = x 2 y + 3 y 2 and read it two ways: as a shape rising above the x y -plane, and as a vector field on that plane. The figure draws both in the domain: the level curves of f , and the gradient arrows ∇ f = ( 2 x y ,   x 2 + 6 y ) , which at ( 2 , 1 ) is ( 4 , 10 ) .

Which space each object lives in. For f : R 2 → R the gradient at a point is a vector of R 2 , drawn in the domain plane at that point. It is not a tangent vector to the surface at ( x , y , f ( x , y ) ) ; it does not point uphill along the surface but names the uphill direction as seen from above. The chain is:

( x , y ) ⟶ height  f ( x , y ) ⟶ the contour through  ( x , y ) ⟶ ∇ f ( x , y ) ⟂ that contour .
QuestionSurface answersGradient answers
rate along ( 3 , 4 ) / 5 measure a slope on the drawing 4 ( 0.6 ) + 10 ( 0.8 ) = 10.4
steepest directioneyeball the contoursalong ( 4 , 10 ) itself
that steepest rateestimate | ( 4 , 10 ) | = 116 ≈ 10.770
zero-change directiontangent to the contourperpendicular: ( − 10 , 4 ) / 116
is this a critical point?is the surface level?is the vector zero?

Every row above reads at a regular point, where ∇ f ≠ 0 — as at ( 2 , 1 ) , the marked point. There the level curve has a tangent line and the gradient is perpendicular to it. At a critical point, where ∇ f = 0 , every directional derivative is zero, no direction is steepest, and the level set need not be a curve with a tangent at all. The perpendicularity statement is about regular points.

Why the steepest direction is the gradient. For a unit vector u the directional derivative is ∇ f ⋅ u , and Cauchy–Schwarz gives ∇ f ⋅ u ≤ | ∇ f | with equality exactly when u points along ∇ f . So the maximum rate is | ∇ f | = 116 , attained in the gradient direction, and it is zero precisely for u perpendicular to ∇ f — which is the tangent to the contour. No search over directions is needed; the inequality settles every direction at once.

What each form is for. The gradient computes: every directional rate is a dot product, and optimisation steps along ± ∇ f without drawing anything. The shape explains: it is what distinguishes a maximum, a minimum and a saddle, which a vanishing gradient cannot.

That last point is why both are kept. ∇ f = 0 says the surface is level and nothing more. Whether the level point is a peak, a basin or a pass requires the second derivatives, or one glance at whether the contours close around the point or cross through it.

Theorem

The gradient's direction, and the second-derivative test

Theorem (what the gradient maximises). Let f be differentiable at p with ∇ f ( p ) ≠ 0 . Then over all unit vectors u , the directional derivative D u f ( p ) = ∇ f ( p ) ⋅ u attains

  • its maximum + | ∇ f ( p ) | when u points along ∇ f ( p ) ;
  • its minimum − | ∇ f ( p ) | when u points along − ∇ f ( p ) ;
  • the value 0 exactly when u ⟂ ∇ f ( p ) .

Proof. For a unit u , ∇ f ⋅ u = | ∇ f | | u | cos ⁡ θ = | ∇ f | cos ⁡ θ , where θ is the angle between them. Since cos ⁡ θ ranges over [ − 1 , 1 ] , taking its extremes at θ = 0 and θ = π and vanishing at θ = π / 2 , the three statements follow. ◼

The proof is one line because the geometry is entirely in the dot product. Note where | u | = 1 was used, without it the factor | u | survives and the "maximum" could be made arbitrarily large by lengthening u .

Verification. For f = x 2 y + 3 y 2 at ( 2 , 1 ) , ∇ f = ( 4 , 10 ) with | ∇ f | = 116 ≈ 10.770 . Scanning D u f over directions at a hundredth of a degree gives a maximum of 10.770330 , matching | ∇ f | = 10.770330 , and the perpendicular direction returns − 4.4 × 10 − 16 .

Corollary (level curves). The gradient is perpendicular to the level curve through the point. Moving along a contour, f is constant, so its rate of change is zero, and zero directional derivative means perpendicular to ∇ f .

Theorem (second-derivative test). Let f have continuous second partials near a critical point ( a , b ) , where ∇ f ( a , b ) = 0 . Put

D = f x x f y y − f x y 2 evaluated at  ( a , b ) .
Verdict
D > 0 and f x x > 0 local minimum
D > 0 and f x x < 0 local maximum
D < 0 saddle point
D = 0 test is inconclusive

D is the determinant of the Hessian ( f x x f x y f x y f y y ) , and the three cases are its definiteness: positive definite, negative definite, indefinite. That connection is why the quadratic forms unit's classification applies here directly.

Why D < 0 means a saddle. A negative determinant makes the Hessian's eigenvalues opposite in sign, so the quadratic form u T H u is positive along one eigenvector and negative along the other. The surface rises in one direction and falls in another.

Verification on three cases.

f critical point D f x x verdict
x 2 + 3 y 2 − 4 x + 6 y ( 2 , − 1 ) 12 2 minimum, value − 7
− x 2 − y 2 + 2 x + 4 y ( 1 , 2 ) 4 − 2 maximum, value 5
x 2 − y 2 ( 0 , 0 ) − 4 2 saddle

For the first, 20,000 random points within 0.5 found none below − 7 ; for the second, none above 5 . For the third, h ( ± 0.1 , 0 ) = + 0.01 while h ( 0 , ± 0.1 ) = − 0.01 . Both signs within any neighbourhood, so neither a maximum nor a minimum.

Why D = 0 is genuinely undecided. As with f ″ ( c ) = 0 in one variable, the second-order terms vanish in some direction and the behaviour is settled by higher-order terms the test does not examine.

Worked example

Partial derivatives, a directional derivative, and two critical points

1. Partial derivatives of f ( x , y ) = x 2 y + 3 y 2 at ( 2 , 1 ) .

For f x , treat y as a constant. The term x 2 y differentiates to 2 x y ; the term 3 y 2 has no x and differentiates to 0:

f x = 2 x y ⟹ f x ( 2 , 1 ) = 4 .

For f y , treat x as a constant. Now x 2 y differentiates to x 2 and 3 y 2 to 6 y :

f y = x 2 + 6 y ⟹ f y ( 2 , 1 ) = 4 + 6 = 10 .

Check numerically: symmetric difference quotients at h = 10 − 6 return 4.00000000 .

So ∇ f ( 2 , 1 ) = ( 4 , 10 ) .

2. A directional derivative along ( 3 , 4 ) .

The vector ( 3 , 4 ) has length 5, so it must be normalised first:

u = ( 3 5 , 4 5 ) = ( 0.6 , 0.8 ) .
D u f ( 2 , 1 ) = ∇ f ⋅ u = 4 ( 0.6 ) + 10 ( 0.8 ) = 2.4 + 8 = 10.4 .

Check: f ( p + t u ) − f ( p − t u ) 2 t at t = 10 − 6 .

Skipping the normalisation would give 4 ( 3 ) + 10 ( 4 ) = 52 , five times too large, since | ( 3 , 4 ) | = 5 . That is not the rate in any direction.

3. Steepest increase and the level direction.

The steepest increase is along ∇ f = ( 4 , 10 ) itself, at rate

| ∇ f | = 16 + 100 = 116 ≈ 10.770 .

Normalising, the direction is ( 0.3714 , 0.9285 ) , and dotting the gradient with it returns 10.770 , equal to the magnitude, as it must be.

Perpendicular to that, ( − 0.9285 , 0.3714 ) , the dot product is − 4.4 × 10 − 16 : zero to machine precision. That is the tangent to the level curve, along which f does not change.

4. Classifying critical points of g ( x , y ) = x 2 + 3 y 2 − 4 x + 6 y .

Set both partials to zero:

g x = 2 x − 4 = 0 ⇒ x = 2 , g y = 6 y + 6 = 0 ⇒ y = − 1 .

The only critical point is ( 2 , − 1 ) , where g = 4 + 3 − 8 − 6 = − 7 .

Second derivatives: g x x = 2 , g y y = 6 , g x y = 0 . The Hessian determinant is

D = g x x g y y − g x y 2 = 12 − 0 = 12 > 0 ,

and g x x = 2 > 0 , so ( 2 , − 1 ) is a local minimum.

Check: 20,000 random points within 0.5 of ( 2 , − 1 ) produced no value below − 7 .

5. A saddle: h ( x , y ) = x 2 − y 2 at the origin.

h x = 2 x and h y = − 2 y vanish only at ( 0 , 0 ) . Here h x x = 2 , h y y = − 2 , h x y = 0 , so

D = ( 2 ) ( − 2 ) − 0 = − 4 < 0 ,

which makes it a saddle. The behaviour is visible directly: along the x -axis h ( ± 0.1 , 0 ) = + 0.01 , rising both ways; along the y -axis h ( 0 , ± 0.1 ) = − 0.01 , falling both ways. The point is a minimum in one direction and a maximum in another, which is why no single second derivative can classify it.

Procedure

Working with partial derivatives and the gradient

To compute a partial derivative. Decide which variable is live, treat every other as a constant, and differentiate by the ordinary single-variable rules. A term containing none of the live variable differentiates to 0. Nothing from the derivative unit changes. The only new discipline is keeping track of which symbol is currently a number.

To compute a directional derivative.

  1. Compute ∇ f = ( f x , f y ) at the point.
  2. Normalise the direction: u = v / | v | . Skip this and the answer is scaled by | v | .
  3. Take the dot product D u f = ∇ f ⋅ u .

Check: the result must lie in [ − | ∇ f | ,   | ∇ f | ] . A value outside that range means the direction was not normalised.

To answer the geometric questions. With ∇ f in hand, all three are immediate:

QuestionAnswer
steepest increasealong ∇ f , at rate | ∇ f |
steepest decreasealong − ∇ f , at rate − | ∇ f |
no changeperpendicular to ∇ f — the level curve

To find and classify critical points.

  1. Solve f x = 0 and f y = 0 simultaneously. Both must vanish at once; a point where only one does is not critical.
  2. Compute f x x , f y y , f x y at each solution.
  3. Form D = f x x f y y − f x y 2 .
  4. Read the verdict: D > 0 with f x x > 0 is a minimum, D > 0 with f x x < 0 a maximum, D < 0 a saddle, and D = 0 undecided, use another argument.

Check: evaluate f at a few nearby points. A claimed minimum with a smaller neighbour is wrong, and the two axis directions alone will expose most saddles.

Confirming any partial numerically. A symmetric difference quotient in one variable, holding the other fixed, at h = 10 − 6 , settles a suspected algebra error in seconds.

Example

Gradients and critical points worth recognising

A plane, f ( x , y ) = 3 x + 2 y . The gradient is ( 3 , 2 ) everywhere, constant, because a plane has the same slope at every point. The steepest increase is always along ( 3 , 2 ) at rate 13 ≈ 3.606 , and the level curves are the parallel lines 3 x + 2 y = c , perpendicular to that direction. No critical points exist: the gradient never vanishes.

A paraboloid, f ( x , y ) = x 2 + y 2 . The gradient ( 2 x , 2 y ) points radially outward, vanishing only at the origin. There f x x = f y y = 2 and f x y = 0 , so D = 4 > 0 with f x x > 0 : a minimum, value 0. The level curves are circles, and the gradient is perpendicular to each, radial lines cross circles at right angles.

A shifted bowl, g ( x , y ) = x 2 + 3 y 2 − 4 x + 6 y . Setting g x = 2 x − 4 and g y = 6 y + 6 to zero gives the single critical point ( 2 , − 1 ) , with g = − 7 . Here D = ( 2 ) ( 6 ) − 0 = 12 > 0 and g x x = 2 > 0 , so it is a minimum. Twenty thousand random points within 0.5 of it produced nothing below − 7 .

An inverted bowl, h ( x , y ) = − x 2 − y 2 + 2 x + 4 y . The critical point is ( 1 , 2 ) with h = 5 , and D = ( − 2 ) ( − 2 ) − 0 = 4 > 0 with h x x = − 2 < 0 : a maximum. Twenty thousand nearby probes found nothing above 5.

The saddle, s ( x , y ) = x 2 − y 2 . The gradient ( 2 x , − 2 y ) vanishes at the origin, where D = ( 2 ) ( − 2 ) − 0 = − 4 < 0 . Along the x -axis the surface rises, s ( ± 0.1 , 0 ) = + 0.01 , and along the y -axis it falls, s ( 0 , ± 0.1 ) = − 0.01 . Both signs occur in every neighbourhood, so the point is neither a maximum nor a minimum. This is the case with no single-variable counterpart.

A product, f ( x , y ) = x 3 y 2 . The partials are f x = 3 x 2 y 2 and f y = 2 x 3 y , giving ( 12 , 4 ) at ( 1 , 2 ) , confirmed numerically. Note that f x still contains y : holding a variable fixed does not remove it from the answer, only from the differentiation.

A double integral, ∫ 0 2 ∫ 0 3 ( 2 x + y ) d y d x = 21 . Inner over y gives 6 x + 4.5 ; outer over x gives 12 + 9 . Reversing the order gives 4 + 2 y then 12 + 9 . The same 21, as Fubini promises.

The gradient vanishes at three of these and never at the other two, and where it vanishes the determinant D separates the three outcomes. The sign of D is doing work no single number could: it reports whether the surface curves the same way in every direction or opposite ways in two.

Non-example

Errors the extra dimension invites

Dotting with an unnormalised direction. For f = x 2 y + 3 y 2 at ( 2 , 1 ) with ∇ f = ( 4 , 10 ) , the rate along ( 3 , 4 ) is not 4 ( 3 ) + 10 ( 4 ) = 52 . That vector has length 5, so the dot product reports five times the rate per unit distance. Normalising to ( 0.6 , 0.8 ) gives the correct 10.4 , and 52 / 10.4 = 5 exactly.

The check that catches it: a directional derivative can never exceed | ∇ f | , here 116 ≈ 10.77 . A reported 52 is impossible on its face.

Treating one vanishing partial as a critical point. For g = x 2 + 3 y 2 − 4 x + 6 y , the equation g x = 2 x − 4 = 0 gives x = 2 . A line of points, not a critical point. Both partials must vanish simultaneously, and g y = 6 y + 6 = 0 pins y = − 1 . The critical point is the single intersection ( 2 , − 1 ) .

Reading a vanishing gradient as an extremum. At the origin, s = x 2 − y 2 has ∇ s = 0 , and the point is neither a maximum nor a minimum: s ( ± 0.1 , 0 ) = + 0.01 while s ( 0 , ± 0.1 ) = − 0.01 . Every neighbourhood contains larger and smaller values. As in one variable, ∇ f = 0 is necessary and not sufficient, but the extra failure mode here, the saddle, does not arise with one variable.

Classifying with a single second derivative. For that same saddle, s x x = 2 > 0 might suggest a minimum. It does not: s y y = − 2 , and the test uses D = s x x s y y − s x y 2 = − 4 < 0 . A condition over all directions cannot be read from one of them.

Forgetting that D = 0 decides nothing. As with f ″ ( c ) = 0 in one variable, a vanishing determinant leaves the second-order terms silent in some direction, and the behaviour turns on higher-order terms the test never inspects. Reporting "the test gives D = 0 , so it is a saddle" asserts what the test did not say.

Assuming partials existing makes f differentiable. In one variable, a derivative existing at a point forces continuity there. In two it does not: both partials probe only the two coordinate lines, and a function can behave badly along every other direction while those two are fine, so f x and f y can exist at a point where f is not even continuous. Differentiability demands a linear approximation valid in all directions.

Expecting a partial derivative to lose the other variable. For f = x 3 y 2 , f x = 3 x 2 y 2 still contains y . Holding y fixed means treating it as a constant during differentiation, not deleting it from the result.

Application

Following the gradient

Gradient descent. Every model fitted by minimising a loss follows − ∇ f , because that is the direction of steepest decrease. The theorem in this unit is the algorithm's justification. A step is x k + 1 = x k − η ∇ f ( x k ) , with the step size η scaling a direction whose length already carries the local steepness.

The algorithm halts where ∇ f = 0 , which is exactly a critical point, and it therefore inherits the limitation: a vanishing gradient is necessary for a minimum and not sufficient. Saddle points stall it, and in high dimensions saddles vastly outnumber true minima, which is why practical optimisers add momentum or random perturbation to escape them.

Least squares. Fitting y = m x + b minimises S ( m , b ) = ∑ ( y i − m x i − b ) 2 , a function of two variables. Setting ∂ S / ∂ m = 0 and ∂ S / ∂ b = 0 gives the normal equations, and their solution is the regression line. The Hessian is positive definite whenever the x i are not all equal, which is what guarantees the critical point is a minimum rather than a saddle. The condition that the orthogonality unit expresses as the columns being independent.

Maximum likelihood. The log-likelihood of a model with several parameters is a surface, and the estimator is where its gradient vanishes. The Hessian at that point becomes the observed Fisher information, whose inverse estimates the estimator's covariance, so the second derivatives that classify the critical point also quantify the uncertainty.

Thermodynamics and physics. Heat flows along − ∇ T , from hot to cold in the direction of steepest temperature decrease. A conservative force is − ∇ V for a potential V , and equilibrium is where the gradient vanishes, stable when the Hessian is positive definite, unstable at a saddle. The classification is the same three cases with physical names.

Contour maps in the field. A hiker reading a map uses the perpendicularity directly: the steepest path crosses contours at right angles, and a path along a contour neither climbs nor descends. Tightly packed contours warn of steep ground, which is | ∇ f | read off the spacing.

The through-line. In one variable a derivative answers "how fast". In several it must first answer "which way", and the gradient packages both: a direction to move and a rate for moving that way. Every application above is one of those two readings put to work.

Next step

Practice Partial Derivatives, the Gradient and Critical Points

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Double Integrals and Fubini's Theorem

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.