Module 1 of 1 · Lesson 9 of 14

Quadratic Forms and Definiteness

Degree-two expressions as symmetric matrices, and what their eigenvalues decide.

What you will be able to do

Given a quadratic form written in variables, the learner can produce its symmetric matrix, compute the eigenvalues, classify the definiteness, give the principal axes and the diagonalised form, and confirm the verdict by evaluating the form on specific vectors.

Orientation

5 x 2 + 4 x y + 5 y 2 is always positive for nonzero ( x , y ) . x 2 − y 2 takes both signs. Neither fact is visible from the coefficients: both have positive squared terms, and the second has no cross term at all.

Writing the expression as x T A x with A symmetric turns the question into one about eigenvalues, and the spectral theorem answers it. A symmetric matrix can be diagonalised by a rotation, so there is always a set of perpendicular directions in which the cross terms disappear and the form becomes a weighted sum of squares. The weights are the eigenvalues, and their signs settle everything.

That is why this classification appears wherever second-order behaviour matters: a critical point is a minimum, a maximum or a saddle according to whether the Hessian's form is positive definite, negative definite or indefinite.

Figure

Level sets of positive definite, indefinite and semidefinite forms

the sign of the eigenvalues decides the shape

The three classifications as three families of level sets, each in its own coordinate frame centred on its own origin.

Closed nested ellipses mean every nonzero direction gives a positive value, so the origin is a strict minimum: positive definite. Hyperbolas mean the form takes both signs, with the zero set a pair of crossing lines and the origin a saddle: indefinite. Parallel lines mean the form is flat along one eigenvector, whose eigenvalue is zero: semidefinite.

The axes of each family are the eigenvectors, which is what the spectral theorem supplies: a rotation that removes the cross term. On the level set x T A x = c with c > 0 , the semiaxis along an eigenvector with eigenvalue λ > 0 has length c / λ , so at a fixed level the smallest eigenvalue gives the longest axis. A zero eigenvalue has no finite axis at all: the form does not grow in that eigendirection, so the level set runs off to infinity along it, which is the degenerate case on the right.

Definition

Splitting the cross terms, and why symmetry is forced

Reading a form into a matrix. The coefficient of x i 2 goes on the diagonal at a i i . The coefficient of x i x j is halved and placed at both a i j and a j i , because the expansion of x T A x collects those two entries into one term:

Q ( x , y ) = 5 x 2 + 4 x y + 5 y 2 ⟷ A = ( 5 2 2 5 ) .

Forgetting to halve is the most common error, and it changes the matrix without changing the intent: ( 5 4 4 5 ) represents 5 x 2 + 8 x y + 5 y 2 , a different form with different eigenvalues.

Why the symmetric representative is canonical. For any square M , the number x T M x is a 1 × 1 matrix and therefore equal to its own transpose, x T M T x . So M and M T define the same function, and so does their average 1 2 ( M + M T ) , which is symmetric.

Concretely, ( 5 4 0 5 ) gives the same form as ( 5 2 2 5 ) : both evaluate to 5, 5, 14 and 89 at ( 1 , 0 ) , ( 0 , 1 ) , ( 1 , 1 ) and ( 2 , 3 ) . The symmetric one is preferred because only it has the properties the classification uses, real eigenvalues and an orthonormal eigenbasis. The eigenvalues of the non-symmetric version are both 5, which says nothing about this form.

What the matrix is a property of. Given the form, the symmetric matrix is unique. Given a non-symmetric matrix, the form is still determined, but the matrix is not recoverable from it, information about the antisymmetric part is discarded, and it never affected the values.

Evaluation is a double sum. Q ( x ) = ∑ i , j a i j x i x j , so a 3 × 3 form has three squared terms and three cross terms. When checking a classification by evaluating at specific vectors, it is usually quicker to substitute into the original expression than into the matrix product.

Theorem

The spectral theorem, and what it gives the form

Spectral theorem. Every real symmetric matrix A has real eigenvalues and an orthonormal basis of eigenvectors, so A = P D P T with P orthogonal and D diagonal.

Real eigenvalues. Suppose A v = λ v with v ≠ 0 , allowing complex entries. Then

v ¯ T A v = λ v ¯ T v = λ ‖ v ‖ 2 .

Taking the conjugate transpose of the left side and using A T = A with A real gives v ¯ T A v ― = v ¯ T A v , so that number is real. Since ‖ v ‖ 2 > 0 is real and positive, λ is real.

Orthogonal eigenvectors for distinct eigenvalues. If A v 1 = λ 1 v 1 and A v 2 = λ 2 v 2 then

λ 1 v 1 T v 2 = ( A v 1 ) T v 2 = v 1 T A v 2 = λ 2 v 1 T v 2 ,

using symmetry in the middle step. So ( λ 1 − λ 2 ) v 1 T v 2 = 0 , and distinct eigenvalues force v 1 ⟂ v 2 . Within a repeated eigenvalue's eigenspace an orthonormal basis is chosen by Gram–Schmidt, and the full argument that the eigenvectors span proceeds by induction on dimension. ◼

This is the theorem the diagonalization and SVD units invoked without proof; here it is discharged.

Corollary (principal axes). Substituting y = P T x into Q ( x ) = x T A x gives

Q = λ 1 y 1 2 + ⋯ + λ n y n 2 .

Every real quadratic form is a weighted sum of squares in suitable perpendicular coordinates, with the eigenvalues as weights.

Corollary (classification). Q ( x ) > 0 for all x ≠ 0 exactly when every λ i > 0 . The forward direction: take x = P e i , giving Q = λ i , so each eigenvalue is a value of the form and must be positive. The converse: in principal coordinates Q = ∑ i λ i y i 2 , which is positive whenever some y i ≠ 0 , and y = 0 only when x = 0 since P is invertible. The other classes follow identically.

Why orthogonality matters here. Completing the square also removes cross terms, and Sylvester's law of inertia guarantees it produces the same counts of positive, negative and zero coefficients. But its change of variables is not orthogonal, so it does not preserve lengths: the axes it produces are not perpendicular and the level set it describes is a sheared picture of the true one. The eigenvalues are recoverable only from the orthogonal route.

Procedure

From expression to verdict

  1. Build the symmetric matrix. Squared coefficients on the diagonal; each cross-term coefficient halved and placed symmetrically.
  2. If given a matrix instead, check it is symmetric. If not, replace it by 1 2 ( M + M T ) . The form is unchanged.
  3. Find the eigenvalues.
  4. Classify by their signs: all positive, all negative, mixed, or some zero with the rest of one sign.
  5. For the axes, compute an orthonormal eigenbasis; these are the principal directions, and Q = ∑ i λ i y i 2 in those coordinates.
  6. Confirm by evaluating Q on a few vectors, including an eigenvector, Q ( v i ) = λ i for a unit eigenvector.

Shortcuts for 2 × 2 . With t = tr ⁡ A and d = det A :

ConditionVerdict
d > 0 , a 11 > 0 positive definite
d > 0 , a 11 < 0 negative definite
d < 0 indefinite
d = 0 semidefinite, sign from t

These follow from d = λ 1 λ 2 and t = λ 1 + λ 2 : a negative product forces opposite signs, and a positive product with the trace's sign settles which.

The analogue for larger matrices is Sylvester's criterion: positive definite exactly when every leading principal minor is positive. It does not extend naively to semidefiniteness, where all principal minors, not only the leading ones, must be checked.

Checks worth running.

  • ∑ i λ i = tr ⁡ A and ∏ i λ i = det A , both free.
  • A positive definite matrix has every diagonal entry positive, since Q ( e i ) = a i i . The converse fails, which is what the non-example block is about.
  • For a semidefinite verdict, exhibit the nonzero vector on which Q vanishes; it lies in the kernel.

A note on what not to do. Do not classify by inspecting the coefficients of the squared terms. 2 x 2 − 4 x y + 5 y 2 has a negative cross term and is positive definite; x 2 + 4 x y + y 2 has all-positive squared coefficients and is indefinite. The cross terms are not a perturbation. They are half the matrix.

Worked example

Classifying a form and finding its axes

Q ( x , y ) = 5 x 2 + 4 x y + 5 y 2 .

Step 1: the symmetric matrix. Squared coefficients 5 and 5 on the diagonal; the cross coefficient 4 halved to 2 off it:

A = ( 5 2 2 5 ) .

Check. Q ( 1 , 0 ) = 5 , Q ( 0 , 1 ) = 5 , Q ( 1 , 1 ) = 5 + 4 + 5 = 14 , and x T A x at ( 1 , 1 ) gives 5 + 2 + 2 + 5 = 14 .

Step 2–3: eigenvalues. tr ⁡ A = 10 , det A = 25 − 4 = 21 , so p ( λ ) = λ 2 − 10 λ + 21 = ( λ − 7 ) ( λ − 3 ) :

λ 1 = 7 , λ 2 = 3 .

Check. 7 + 3 = 10 = tr ⁡ A and 7 × 3 = 21 = det A .

Step 4: classify. Both eigenvalues positive, so Q is positive definite: Q ( x , y ) > 0 for every ( x , y ) ≠ ( 0 , 0 ) . The 2 × 2 shortcut agrees, det A = 21 > 0 and a 11 = 5 > 0 .

Step 5: principal axes. For λ = 7 : A − 7 I = ( − 2 2 2 − 2 ) , giving v 1 ∝ ( 1 , 1 ) . For λ = 3 : ( 2 2 2 2 ) , giving v 2 ∝ ( 1 , − 1 ) . Normalising,

v 1 = 1 2 ( 1 , 1 ) , v 2 = 1 2 ( 1 , − 1 ) ,

perpendicular as the spectral theorem promised. In these coordinates

Q = 7 y 1 2 + 3 y 2 2 .

Step 6: confirm. Evaluating the original form on the unit eigenvectors:

Q ( v 1 ) = 5 ⋅ 1 2 + 4 ⋅ 1 2 + 5 ⋅ 1 2 = 7 = λ 1 Q ( v 2 ) = 5 2 − 2 + 5 2 = 3 = λ 2

Each eigenvalue is literally the value of the form along its own axis.

The geometry. The level set Q = 1 is an ellipse with axes along ( 1 , 1 ) and ( 1 , − 1 ) , of semi-axis lengths 1 / 7 ≈ 0.378 and 1 / 3 ≈ 0.577 . The larger eigenvalue gives the shorter axis: the form reaches 1 sooner in the direction where it grows faster.

Completing the square, for contrast.

5 x 2 + 4 x y + 5 y 2 = 5 ( x + 2 5 y ) 2 + 21 5 y 2 .

Checking at ( 2 , − 1 ) : directly 20 − 8 + 5 = 17 , and 5 ( 2 − 2 5 ) 2 + 21 5 = 5 ⋅ 64 25 + 21 5 = 64 5 + 21 5 = 17 .

Both coefficients, 5 and 21 / 5 , are positive, confirming positive definiteness. But they are not the eigenvalues 7 and 3, and the variables x + 2 5 y and y are not perpendicular directions. Sylvester's law guarantees the signs agree; nothing guarantees the values do.

Example

One of each class

Positive definite, despite a negative cross term. Q = 2 x 2 − 4 x y + 5 y 2 has matrix ( 2 − 2 − 2 5 ) , trace 7 and determinant 10 − 4 = 6 . The characteristic polynomial λ 2 − 7 λ + 6 factors as ( λ − 6 ) ( λ − 1 ) , so the eigenvalues are 6 and 1: positive definite.

The sign of the cross term was irrelevant. Checking a few values: Q ( 1 , 0 ) = 2 , Q ( 0 , 1 ) = 5 , Q ( 1 , 1 ) = 2 − 4 + 5 = 3 , all positive, and the minimum over the unit circle is 1, the smaller eigenvalue.

Its principal axes are 1 5 ( 1 , − 2 ) for λ = 6 and 1 5 ( 2 , 1 ) for λ = 1 , perpendicular as the theorem requires, and Q evaluates to exactly 6 and 1 on them.

Indefinite. Q = x 2 − y 2 has matrix diag ⁡ ( 1 , − 1 ) , eigenvalues 1 and − 1 . Already diagonal, so the standard axes are the principal ones. Q ( 1 , 0 ) = 1 > 0 , Q ( 0 , 1 ) = − 1 < 0 , and Q ( 1 , 1 ) = 0 . The form vanishes on the lines y = ± x without being semidefinite, because it takes both signs elsewhere. Its level set Q = 1 is a hyperbola.

Negative definite. Q = − 3 x 2 − 2 x y − 3 y 2 has matrix ( − 3 − 1 − 1 − 3 ) , trace − 6 , determinant 9 − 1 = 8 , eigenvalues − 2 and − 4 . Every value is negative: Q ( 1 , 0 ) = − 3 , Q ( 1 , 1 ) = − 8 , Q ( 1 , − 1 ) = − 4 . Note the determinant is positive here, for a 2 × 2 , a positive determinant means same-signed eigenvalues, and the diagonal entry settles which sign.

Positive semidefinite in three variables. Q = x 2 + y 2 + z 2 + 2 x y has matrix

( 1 1 0 1 1 0 0 0 1 ) ,

which is block diagonal: the 2 × 2 block has eigenvalues 2 and 0, and the isolated entry gives 1. So the eigenvalues are 2, 1, 0, semidefinite, not definite.

The zero eigenvalue's eigenvector is ( 1 , − 1 , 0 ) , and indeed Q ( 1 , − 1 , 0 ) = 1 + 1 + 0 − 2 = 0 with the vector nonzero. Along the whole line through it the form vanishes; everywhere off that line it is strictly positive, as Q ( 1 , 1 , 0 ) = 4 , Q ( 0 , 0 , 3 ) = 9 and Q ( 1 , − 1 , 5 ) = 25 show.

Cross terms decide as much as squared ones; a positive determinant does not mean positive definite; and the gap between definite and semidefinite is a single nonzero vector on which the form vanishes. None of this is visible without the eigenvalues.

Non-example

Five ways a classification goes wrong

Reading definiteness off the squared coefficients. " x 2 + 4 x y + y 2 has positive coefficients on x 2 and y 2 , so it is positive definite." Its matrix is ( 1 2 2 1 ) with det = 1 − 4 = − 3 < 0 , so the eigenvalues are 3 and − 1 : indefinite. Evaluating confirms it, Q ( 1 , 1 ) = 6 > 0 and Q ( 1 , − 1 ) = 1 − 4 + 1 = − 2 < 0 .

The converse error is just as common: 2 x 2 − 4 x y + 5 y 2 has a negative cross term and is positive definite, with eigenvalues 6 and 1 .

Forgetting to halve the cross coefficient. Writing ( 5 4 4 5 ) for 5 x 2 + 4 x y + 5 y 2 . That matrix represents 5 x 2 + 8 x y + 5 y 2 , whose eigenvalues are 9 and 1 rather than 7 and 3 . The verdict happens to survive here, both are still positive, but the axes, the level set and any quantitative conclusion are wrong.

Using a non-symmetric matrix's eigenvalues. ( 5 4 0 5 ) defines the same form as ( 5 2 2 5 ) , as evaluating at any point confirms. Its eigenvalues are both 5, which are neither 7 nor 3 and say nothing about this form. Definiteness is a property of the symmetric representative; the eigenvalues of any other representative are an artefact of how it was written.

Calling a semidefinite form definite. ( x + y ) 2 has matrix ( 1 1 1 1 ) with eigenvalues 2 and 0 . It is never negative, but it is not positive definite: Q ( 1 , − 1 ) = 0 with ( 1 , − 1 ) ≠ 0 . The distinction is the difference between > and ≥ , and it matters. A positive semidefinite Hessian does not establish a strict minimum.

Taking completing-the-square coefficients as eigenvalues. From 5 ( x + 2 5 y ) 2 + 21 5 y 2 , reporting the eigenvalues as 5 and 21 / 5 . They are 7 and 3 . Sylvester's law guarantees only that the signs match; the values differ because the change of variables was not orthogonal.

The first three mistake a representation for the object, coefficients, an unhalved matrix, an asymmetric one. The last two mistake a weaker conclusion for a stronger one. All five are caught by the same habit: build the symmetric matrix, take its eigenvalues, and test the verdict on a concrete vector.

Optional enrichment (1)

Application

Second derivatives, covariance, and energy

Classifying a critical point. For a twice-differentiable f : R n → R with ∇ f ( a ) = 0 , the second-order Taylor expansion near a is

f ( a + h ) ≈ f ( a ) + 1 2 h T H h ,

where H is the Hessian of second partial derivatives, symmetric whenever those derivatives are continuous, which is what brings it into this unit's scope. The behaviour near a is the behaviour of the quadratic form h T H h :

Hessian formCritical point
positive definitelocal minimum
negative definitelocal maximum
indefinitesaddle
semidefiniteundetermined at this order

The last row is the one to watch. A zero eigenvalue means the form vanishes along some direction, and the second-order term says nothing about what happens there, f ( x , y ) = x 2 + y 4 and f ( x , y ) = x 2 − y 4 have the same Hessian at the origin and different behaviour.

For two variables this is the familiar test: det H > 0 with f x x > 0 gives a minimum, det H < 0 gives a saddle. Those rules are the 2 × 2 definiteness shortcuts, not separate facts.

Covariance matrices. For a random vector X with covariance Σ , the variance of the linear combination a T X is a T Σ a . Variance is never negative, so Σ is positive semidefinite, necessarily, not incidentally. A zero eigenvalue means some combination of the variables has zero variance, which is an exact linear relationship among them.

That is also why the principal components of the SVD unit are eigenvectors of Σ : the largest eigenvalue is the greatest variance achievable by any unit combination, attained along its eigenvector. The optimisation "maximise a T Σ a subject to ‖ a ‖ = 1 " has answer λ 1 , because in principal coordinates it is maximising a weighted sum of squares whose weights sum to one.

Energy and stability. In mechanics the potential energy near equilibrium is a quadratic form in the displacements, and the equilibrium is stable exactly when that form is positive definite. Every displacement raises the energy. An indefinite form gives an unstable equilibrium with a direction of descent, and the eigenvector for the most negative eigenvalue is the direction the system falls fastest.

What the classification does not do. It is local and second-order. Positive definiteness of a Hessian at a point establishes a local minimum, not a global one, and says nothing away from that point. For the semidefinite case it establishes nothing at all, and higher-order terms must be examined.

Next step

Practice Quadratic Forms and Definiteness

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Matrix Inverses and Elementary Matrices

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.