Partial Derivatives, the Gradient and Critical Points

Partial derivatives as rates along the coordinate axes, the gradient that assembles them into a vector determining the rate in every other direction, and the Hessian that separates a minimum from a maximum from the saddle that one variable cannot produce.

Definition

Partial derivatives. For f ( x , y ) , the partial derivative with respect to x holds y fixed:

f x ( a , b ) = lim h → 0 f ( a + h , b ) − f ( a , b ) h ,

and symmetrically for f y . Computationally this is single-variable differentiation with the other variables treated as constants, so every rule from the single-variable unit applies unchanged.

Notation. f x , ∂ f ∂ x and ∂ x f all denote the same thing. The curved ∂ marks that other variables are being held, which a plain d would not.

Higher and mixed partials. f x x differentiates twice in x ; f x y differentiates first in x then in y . Clairaut's theorem: if f x y and f y x are continuous near a point, they are equal there, so the order of mixed differentiation does not matter for the functions ordinarily met.

The gradient. ∇ f = ( f x , f y ) , a vector field assembling the partials.

The directional derivative. For a unit vector u ,

D u f = ∇ f ⋅ u ,

the rate of change of f moving in the direction u . The requirement that | u | = 1 is load-bearing: without normalising, the answer scales with the vector's length and is not a rate per unit distance.

What the gradient points at. By ∇ f ⋅ u = | ∇ f | cos ⁡ θ , the directional derivative is largest when θ = 0 . So ∇ f points in the direction of steepest increase, its magnitude | ∇ f | is that maximum rate, and at a point where ∇ f ≠ 0 the directions perpendicular to it give zero. Those are the directions tangent to the level curve through the point: the gradient is normal to the level set, and moving along the curve keeps f constant.

Critical points. An interior extremum requires ∇ f = 0 , both partials vanishing at once. As in one variable this is necessary and not sufficient, and the second-derivative test uses the Hessian ( f x x f x y f x y f y y ) : positive definite gives a minimum, negative definite a maximum, indefinite a saddle.

Assumptions and scope

  • A directional derivative requires a unit vector. Using an unnormalised direction scales the answer by that vector's length and no longer reports a rate per unit distance.

  • Clairaut's theorem needs the mixed partials to be continuous near the point. Functions exist for which f x y ≠ f y x at a point where that fails.

  • ∇ f = 0 is necessary but not sufficient for an interior extremum, and in two variables the extra failure mode is a saddle, which has no single-variable counterpart.

  • Partial derivatives existing at a point does not make f differentiable there, nor even continuous. Differentiability is a stronger condition than in one variable, where differentiability at a point follows from the single derivative existing.

Forms this is expressed in

The same content in several forms. Each makes something visible that the others leave implicit, so moving between them is part of understanding the topic rather than a presentation choice.

geometric

The surface z = x²y + 3y² over the square −2 ≤ x, y ≤ 2

A function f ( x , y ) drawn as the surface z = f ( x , y ) standing above the plane.

This form makes the central multivariable difficulty visible: a surface has no single slope at a point. Walking east along the surface climbs at one rate, walking north at another, walking north-east at a third. That is why one derivative is replaced by a gradient, and why a directional derivative has to name its direction.

It is also the form in which a critical point looks like something. A maximum is a summit and a minimum a basin, while a saddle rises along one axis and falls along another, so it is level without being extreme. The saddle has no single-variable analogue, and the surface is where its shape is obvious.

What the surface is poor at is reading rates. Judging steepness from a drawn perspective is unreliable, and any exact value requires the symbolic form. For comparing rates across a region the contour map is the better view of the same function.

Translates into: geometric, symbolic

geometric

The same function as level curves, where spacing measures steepness

The same function drawn flat, as the level curves f ( x , y ) = c for a ladder of values c .

This form carries the surface's information in the form most useful for reading rates. Closely spaced contours mean the surface is steep; widely spaced ones mean it is flat, and the spacing is measurable in a way a perspective drawing's steepness is not.

A path along a contour has zero rate of change by construction, which is the geometric content of ∇ f ⋅ u = 0 for u tangent to the level curve. The gradient at a point is therefore perpendicular to the contour through it, and the contour map is where that relation is seen.

Critical points separate at a glance: closed contours shrinking to a point mark a maximum or a minimum, while contours crossing in an X mark a saddle. That distinction is the one the algebra does not make obvious and the contour map does.

What the map loses is the sense of height. Which of two nested families is the peak and which the basin is a matter of reading the labels, not the shape.

Translates into: geometric, geometric

symbolic

The gradient as an expression: ∇ f = ( f x , f y ) , two partial derivatives collected into one vector-valued formula.

This form makes every directional rate computable from two numbers. For f ( x , y ) = x 2 y + 3 y 2 the gradient is ∇ f = ( 2 x y ,   x 2 + 6 y ) , which at ( 2 , 1 ) evaluates to ( 4 , 10 ) ; the rate in the direction ( 3 , 4 ) / 5 is the dot product 4 ( 0.6 ) + 10 ( 0.8 ) = 10.4 . No new limit is taken. Infinitely many directional derivatives are encoded in one expression, which is what makes the gradient the right object rather than a convenient bundle.

It also answers the geometric questions by arithmetic. The steepest increase is along ∇ f itself, at rate | ∇ f | = 4 2 + 10 2 ≈ 10.770 ; the steepest decrease is along − ∇ f ; and any direction perpendicular to ∇ f gives zero. Each follows from ∇ f ⋅ u = | ∇ f | cos ⁡ θ without drawing anything.

What the expression does not show is the shape of the field it defines: which way the arrows point across a whole region, where they are long and where they vanish. That is the field's own form.

Translates into: geometric, geometric

geometric

The gradient as an arrow at every point

The same gradient drawn as a field: at each point of the plane, an arrow pointing in the direction of steepest increase, with length equal to that maximum rate.

This form makes the behaviour of ∇ f over a whole region visible at once, which the expression gives only one point at a time. Where the arrows are long the surface is steep; where they shorten it is flattening; where one vanishes the surface is level and a critical point sits there. The origin is marked for exactly that reason: ∇ f = 0 there, so no arrow is drawn and no direction is steepest. The pattern is the object optimisation acts on: gradient descent steps along − ∇ f , so the algorithm follows the arrows backwards.

It also shows the relation to the level curves directly. Every arrow crosses the contour through its point at a right angle, because a direction perpendicular to ∇ f gives rate zero, which is the level curve's defining property. Seeing the field and the contours together is seeing ∇ f ⋅ u = 0 rather than reading it.

What the field cannot supply is a number. Reading a rate off arrow lengths gives an impression; the exact value requires the symbolic gradient.

Translates into: symbolic, geometric

Worked material

Example

Gradients and critical points worth recognising

A plane, f ( x , y ) = 3 x + 2 y . The gradient is ( 3 , 2 ) everywhere, constant, because a plane has the same slope at every point. The steepest increase is always along ( 3 , 2 ) at rate 13 ≈ 3.606 , and the level curves are the parallel lines 3 x + 2 y = c , perpendicular to that direction. No critical points exist: the gradient never vanishes.

A paraboloid, f ( x , y ) = x 2 + y 2 . The gradient ( 2 x , 2 y ) points radially outward, vanishing only at the origin. There f x x = f y y = 2 and f x y = 0 , so D = 4 > 0 with f x x > 0 : a minimum, value 0. The level curves are circles, and the gradient is perpendicular to each, radial lines cross circles at right angles.

A shifted bowl, g ( x , y ) = x 2 + 3 y 2 − 4 x + 6 y . Setting g x = 2 x − 4 and g y = 6 y + 6 to zero gives the single critical point ( 2 , − 1 ) , with g = − 7 . Here D = ( 2 ) ( 6 ) − 0 = 12 > 0 and g x x = 2 > 0 , so it is a minimum. Twenty thousand random points within 0.5 of it produced nothing below − 7 .

An inverted bowl, h ( x , y ) = − x 2 − y 2 + 2 x + 4 y . The critical point is ( 1 , 2 ) with h = 5 , and D = ( − 2 ) ( − 2 ) − 0 = 4 > 0 with h x x = − 2 < 0 : a maximum. Twenty thousand nearby probes found nothing above 5.

The saddle, s ( x , y ) = x 2 − y 2 . The gradient ( 2 x , − 2 y ) vanishes at the origin, where D = ( 2 ) ( − 2 ) − 0 = − 4 < 0 . Along the x -axis the surface rises, s ( ± 0.1 , 0 ) = + 0.01 , and along the y -axis it falls, s ( 0 , ± 0.1 ) = − 0.01 . Both signs occur in every neighbourhood, so the point is neither a maximum nor a minimum. This is the case with no single-variable counterpart.

A product, f ( x , y ) = x 3 y 2 . The partials are f x = 3 x 2 y 2 and f y = 2 x 3 y , giving ( 12 , 4 ) at ( 1 , 2 ) , confirmed numerically. Note that f x still contains y : holding a variable fixed does not remove it from the answer, only from the differentiation.

A double integral, ∫ 0 2 ∫ 0 3 ( 2 x + y ) d y d x = 21 . Inner over y gives 6 x + 4.5 ; outer over x gives 12 + 9 . Reversing the order gives 4 + 2 y then 12 + 9 . The same 21, as Fubini promises.

The gradient vanishes at three of these and never at the other two, and where it vanishes the determinant D separates the three outcomes. The sign of D is doing work no single number could: it reports whether the surface curves the same way in every direction or opposite ways in two.

Non-example

Errors the extra dimension invites

Dotting with an unnormalised direction. For f = x 2 y + 3 y 2 at ( 2 , 1 ) with ∇ f = ( 4 , 10 ) , the rate along ( 3 , 4 ) is not 4 ( 3 ) + 10 ( 4 ) = 52 . That vector has length 5, so the dot product reports five times the rate per unit distance. Normalising to ( 0.6 , 0.8 ) gives the correct 10.4 , and 52 / 10.4 = 5 exactly.

The check that catches it: a directional derivative can never exceed | ∇ f | , here 116 ≈ 10.77 . A reported 52 is impossible on its face.

Treating one vanishing partial as a critical point. For g = x 2 + 3 y 2 − 4 x + 6 y , the equation g x = 2 x − 4 = 0 gives x = 2 . A line of points, not a critical point. Both partials must vanish simultaneously, and g y = 6 y + 6 = 0 pins y = − 1 . The critical point is the single intersection ( 2 , − 1 ) .

Reading a vanishing gradient as an extremum. At the origin, s = x 2 − y 2 has ∇ s = 0 , and the point is neither a maximum nor a minimum: s ( ± 0.1 , 0 ) = + 0.01 while s ( 0 , ± 0.1 ) = − 0.01 . Every neighbourhood contains larger and smaller values. As in one variable, ∇ f = 0 is necessary and not sufficient, but the extra failure mode here, the saddle, does not arise with one variable.

Classifying with a single second derivative. For that same saddle, s x x = 2 > 0 might suggest a minimum. It does not: s y y = − 2 , and the test uses D = s x x s y y − s x y 2 = − 4 < 0 . A condition over all directions cannot be read from one of them.

Forgetting that D = 0 decides nothing. As with f ″ ( c ) = 0 in one variable, a vanishing determinant leaves the second-order terms silent in some direction, and the behaviour turns on higher-order terms the test never inspects. Reporting "the test gives D = 0 , so it is a saddle" asserts what the test did not say.

Assuming partials existing makes f differentiable. In one variable, a derivative existing at a point forces continuity there. In two it does not: both partials probe only the two coordinate lines, and a function can behave badly along every other direction while those two are fine, so f x and f y can exist at a point where f is not even continuous. Differentiability demands a linear approximation valid in all directions.

Expecting a partial derivative to lose the other variable. For f = x 3 y 2 , f x = 3 x 2 y 2 still contains y . Holding y fixed means treating it as a constant during differentiation, not deleting it from the result.

Common errors

Common misconception

The directional derivative is ∇ f ⋅ v for any vector v giving the direction, so normalising is an optional tidying step.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.