Orthogonality and Projection

What you will be able to do

Given a basis of a subspace, the learner can produce an orthogonal basis by Gram–Schmidt, project a vector onto that subspace, verify that the residual is orthogonal to the subspace, and set up the normal equations for a least-squares problem.

Orientation

Finding coordinates in an arbitrary basis means solving a linear system: every coordinate depends on every other. In an orthonormal basis it means taking one inner product per coordinate, because the directions do not interfere.

That independence follows from orthogonality, and it answers a second question at the same time. Given a point off a subspace, the closest point in it is found by keeping the part that lies along the subspace and discarding the part perpendicular to it, and the discarded part being perpendicular is exactly what makes the remainder closest.

Least squares is that construction applied to a system with no solution. When A x = b cannot be solved, replace b by the nearest point of the column space. The normal equations that result are not a trick; they say the residual is perpendicular to every column.

Definition

The inner product axioms and what depends on the choice

Positive definiteness is what makes a norm. Symmetry and linearity alone would permit ⟨ v , v ⟩ < 0 , and ⟨ v , v ⟩ would not exist. Requiring ⟨ v , v ⟩ > 0 for v ≠ 0 is what lets an inner product define a length, and with it distance, angle and the whole geometry.

Orthogonality is relative to the inner product. The vectors ( 1 , 1 ) and ( 1 , − 1 ) are orthogonal under the dot product. Under ⟨ u , v ⟩ = u 1 v 1 + 2 u 2 v 2 , also a valid inner product, their pairing is 1 − 2 = − 1 ≠ 0 , so they are not. The word means nothing until the inner product is named, which is why a space with a chosen inner product is a single object, an inner product space, rather than a space and a separate decoration.

Why an orthogonal set is independent. Suppose ∑ j c j u j = 0 with the u j pairwise orthogonal and nonzero. Take the inner product of both sides with u k : every cross term vanishes, leaving c k ⟨ u k , u k ⟩ = 0 , and positive definiteness forces c k = 0 . So orthogonality is a strictly stronger condition than independence, and an orthogonal spanning set is automatically a basis.

Why the coefficients are inner products. For an orthonormal basis, apply ⟨ ⋅ , e k ⟩ to v = ∑ j c j e j . All terms but one vanish and ⟨ e k , e k ⟩ = 1 , leaving c k = ⟨ v , e k ⟩ . The same computation with a merely orthogonal basis leaves ⟨ u k , u k ⟩ behind, which is where the denominator in the projection formula comes from. It is not a normalisation added by hand.

The denominator is the commonest omission. Writing ∑ j ⟨ b , u j ⟩ u j for an orthogonal but not orthonormal basis scales each component by ‖ u j ‖ 2 , giving a vector in the right subspace with the wrong length. The error is invisible unless the residual is checked for orthogonality.

Projection is idempotent. proj W is a linear map with proj W ∘ proj W = proj W : a point already in W is its own closest point. Its kernel is W ⟂ , the set of vectors orthogonal to all of W , and its image is W , so by rank–nullity dim ⁡ W + dim ⁡ W ⟂ = n .

Figure

Projection, residual, and the right angle between them

b, its projection onto a line, and the perpendicular residual

b is projected onto the line W = span ⁡ { ( 2 , 1 ) } . The green arrow is p = proj W ⁡ b , the closest point of W to b ; the red segment is the residual b − p .

The right angle is the whole theorem. Because b − p ⟂ W , Pythagoras gives ‖ b − w ‖ 2 = ‖ b − p ‖ 2 + ‖ p − w ‖ 2 for every w ∈ W , so any other point of the line is further from b by exactly the amount it differs from p . That is why the perpendicular foot is the minimiser rather than merely a natural choice, and why least squares solves a minimisation by solving an orthogonality condition.

Example

Orthogonal sets in four different spaces

The same definition, applied wherever an inner product exists.

The standard basis of R 3 . e 1 , e 2 , e 3 pair to zero with each other and to one with themselves, so the set is orthonormal. This is why coordinates in the standard basis are so easy to read: each is already an inner product, x j = ⟨ x , e j ⟩ .

A rotated pair in R 2 . ( 1 , 1 ) and ( 1 , − 1 ) have ⟨ u , v ⟩ = 1 − 1 = 0 , so they are orthogonal, but ‖ u ‖ 2 = ‖ v ‖ 2 = 2 , so the set is orthogonal, not orthonormal. Dividing each by 2 gives

1 2 ( 1 , 1 ) , 1 2 ( 1 , − 1 ) ,

which pair to zero and have norm 1. This is the distinction that decides whether the denominators in the projection formula may be dropped.

Polynomials on [ 0 , 1 ] , with ⟨ f , g ⟩ = ∫ 0 1 f g d x . The basis { 1 , x } is not orthogonal:

⟨ 1 , x ⟩ = ∫ 0 1 x d x = 1 2 ≠ 0 .

One Gram–Schmidt step replaces x by x − 1 2 , and ⟨ 1 , x − 1 2 ⟩ = 1 2 − 1 2 = 0 . Since ‖ x − 1 2 ‖ 2 = ∫ 0 1 ( x − 1 2 ) 2 d x = 1 12 , normalising gives 12 ( x − 1 2 ) . These are the first two Legendre polynomials for this interval.

Two functions can look entirely unalike and still fail to be orthogonal; the integral decides, not the appearance.

Trigonometric functions on [ 0 , 2 π ] , with ⟨ f , g ⟩ = ∫ 0 2 π f g d x .

⟨ sin , cos ⟩ = 0 , ⟨ 1 , sin ⟩ = ⟨ 1 , cos ⟩ = 0 ,
⟨ sin ⁡ x , sin ⁡ 2 x ⟩ = ⟨ cos ⁡ x , cos ⁡ 2 x ⟩ = ⟨ sin ⁡ x , cos ⁡ 2 x ⟩ = 0 ,

while ⟨ sin , sin ⟩ = ⟨ cos , cos ⟩ = π . So { 1 , sin ⁡ x , cos ⁡ x , sin ⁡ 2 x , cos ⁡ 2 x , … } is an orthogonal set. The basis Fourier series are written in. Projecting a function onto its span is exactly how Fourier coefficients are computed, and the π in the denominators of the usual formulas is the ⟨ u j , u j ⟩ of the projection formula, not a convention.

Symmetric matrices, with ⟨ A , B ⟩ = tr ⁡ ( A T B ) . This inner product is the sum of entrywise products. The basis from the basis-and-dimension unit,

( 1 0 0 0 ) , ( 0 0 0 1 ) , ( 0 1 1 0 ) ,

is pairwise orthogonal, with squared norms 1 , 1 and 2 . Again orthogonal but not orthonormal: the third must be divided by 2 , because its single off-diagonal value appears twice.

Nothing about the objects, lists of numbers, polynomials, functions, matrices. What they share is an inner product satisfying the three axioms, and every construction in this unit is written in those terms alone. What differs is which sets count as orthogonal, and that depends entirely on the inner product chosen.

Procedure

Orthogonalise, project, verify

Gram–Schmidt. Given a basis v 1 , … , v k of a subspace:

  1. u 1 = v 1 .
  2. For j = 2 , … , k : subtract from v j its projection onto each earlier u i ,
u j = v j − ∑ i < j ⟨ v j , u i ⟩ ⟨ u i , u i ⟩ u i .
  1. Optionally divide each u j by ‖ u j ‖ to normalise.

At each step, check ⟨ u j , u i ⟩ = 0 for every i < j before continuing. An error at step j contaminates every later step, and the check costs one inner product per pair.

Projecting onto a subspace W . With an orthogonal basis u 1 , … , u k of W :

proj W ⁡ ( b ) = ∑ j = 1 k c j u j , c j = ⟨ b , u j ⟩ ⟨ u j , u j ⟩ .

If the basis is orthonormal the denominators are 1 and may be dropped, but only then.

Verification, in two independent ways.

CheckWhat it catches
⟨ b − proj W ⁡ ( b ) , u j ⟩ = 0 for every j a wrong coefficient, a missing denominator
‖ b ‖ 2 = ‖ proj W ⁡ ( b ) ‖ 2 + ‖ e ‖ 2 an arithmetic slip anywhere

The first is the definitive one: the projection is characterised by having an orthogonal residual, so a residual that fails the test means the answer is not the projection.

Least squares for A x ≈ b .

  1. Form A T A and A T b .
  2. Solve A T A x ^ = A T b .
  3. The fitted vector is A x ^ = proj col ⁡ A ⁡ ( b ) , and the residual is e = b − A x ^ .
  4. Check A T e = 0 . The residual is orthogonal to every column.

For fitting y = a + b x to points ( x i , y i ) , the columns of A are a column of ones and the column of x i , so the normal equations read

( n ∑ x i ∑ x i ∑ x i 2 ) ( a b ) = ( ∑ y i ∑ x i y i ) .

The two rows say ∑ i e i = 0 and ∑ i x i e i = 0 , which are the orthogonality conditions written out.

Worked example

Gram–Schmidt, then a projection, then the check

Let W = span ⁡ { ( 1 , 1 , 0 ) , ( 1 , 0 , 1 ) } ⊆ R 3 , and project b = ( 1 , 2 , 3 ) onto it.

---

Step 1: orthogonalise. Take u 1 = v 1 = ( 1 , 1 , 0 ) , so ⟨ u 1 , u 1 ⟩ = 1 + 1 + 0 = 2 .

For v 2 = ( 1 , 0 , 1 ) : ⟨ v 2 , u 1 ⟩ = 1 + 0 + 0 = 1 , so the projection of v 2 onto u 1 is 1 2 ( 1 , 1 , 0 ) = ( 1 2 , 1 2 , 0 ) , and

u 2 = ( 1 , 0 , 1 ) − ( 1 2 , 1 2 , 0 ) = ( 1 2 , − 1 2 , 1 ) .

Check. ⟨ u 1 , u 2 ⟩ = 1 2 − 1 2 + 0 = 0 . Also ⟨ u 2 , u 2 ⟩ = 1 4 + 1 4 + 1 = 3 2 .

The pair { u 1 , u 2 } is an orthogonal basis of the same plane W , Gram–Schmidt never leaves the span.

Step 2: the coefficients.

⟨ b , u 1 ⟩ = 1 + 2 + 0 = 3 , c 1 = 3 2 ,
⟨ b , u 2 ⟩ = 1 2 − 1 + 3 = 5 2 , c 2 = 5 / 2 3 / 2 = 5 3 .

The denominators matter: dropping them would give c 2 = 5 2 instead of 5 3 , and a projection pointing the right way with the wrong length.

Step 3: the projection.

proj W ⁡ ( b ) = 3 2 ( 1 , 1 , 0 ) + 5 3 ( 1 2 , − 1 2 , 1 ) = ( 3 2 + 5 6 , 3 2 − 5 6 , 5 3 ) = ( 7 3 , 2 3 , 5 3 ) .

Step 4: the residual and the checks.

e = b − proj W ⁡ ( b ) = ( 1 − 7 3 , 2 − 2 3 , 3 − 5 3 ) = ( − 4 3 , 4 3 , 4 3 ) .

Orthogonality. ⟨ e , u 1 ⟩ = − 4 3 + 4 3 + 0 = 0 and ⟨ e , u 2 ⟩ = − 2 3 − 2 3 + 4 3 = 0 . Since e is orthogonal to a basis of W , it is orthogonal to all of W .

Pythagoras. ‖ b ‖ 2 = 1 + 4 + 9 = 14 . And

‖ proj W ⁡ ( b ) ‖ 2 = 49 9 + 4 9 + 25 9 = 78 9 = 26 3 , ‖ e ‖ 2 = 16 9 ⋅ 3 = 48 9 = 16 3 ,

and 26 3 + 16 3 = 42 3 = 14 .

( 7 3 , 2 3 , 5 3 ) is the point of the plane W closest to ( 1 , 2 , 3 ) , and the distance between them is ‖ e ‖ = 16 / 3 ≈ 2.31 . Any other point of W is further, because moving within W adds a vector orthogonal to e , and by Pythagoras that only increases the total.

On the ordering. Starting Gram–Schmidt from v 2 instead would give a different orthogonal basis of the same plane, and the same projection, since the projection depends on W and not on which basis describes it. That invariance is worth checking once.

Theorem

The projection is the closest point

Theorem. Let W be a subspace and b any vector. The point p = proj W ⁡ ( b ) is the unique element of W minimising ‖ b − w ‖ over w ∈ W .

Proof. Write e = b − p , which is orthogonal to every element of W by construction. For any w ∈ W ,

b − w = ( b − p ) + ( p − w ) = e + ( p − w ) ,

and p − w ∈ W because W is a subspace, so e ⟂ ( p − w ) . By Pythagoras,

‖ b − w ‖ 2 = ‖ e ‖ 2 + ‖ p − w ‖ 2   ≥   ‖ e ‖ 2 ,

with equality exactly when ‖ p − w ‖ = 0 , that is w = p . ◼

The proof uses only that e is orthogonal to W and that W is closed under subtraction. Nothing depends on coordinates or on the space being R n , so the result holds in any inner product space, including spaces of functions, where it underlies Fourier approximation.

Corollary (Pythagoras for the decomposition). Taking w = 0 gives ‖ b ‖ 2 = ‖ p ‖ 2 + ‖ e ‖ 2 , which is the second verification check.

Corollary (the normal equations). Let W = col ⁡ A . Then x ^ minimises ‖ b − A x ‖ exactly when A x ^ = proj W ⁡ ( b ) , which holds exactly when the residual is orthogonal to every column:

A T ( b − A x ^ ) = 0 ⟺ A T A x ^ = A T b .

So the normal equations are the orthogonality condition rearranged, not a separate construction. They are solvable for any A and b , and the solution is unique exactly when the columns of A are independent, in which case A T A is invertible.

Why uniqueness can fail. If the columns are dependent, many x ^ produce the same A x ^ . The projection is still unique, it is a point of W , and the theorem gives uniqueness there, but its representation in terms of the columns is not. Distinguishing the two is what separates "the fit is unique" from "the coefficients are unique", and only the first is guaranteed.

Non-example

Projecting without an orthogonal basis, and three other errors

Using the formula on a non-orthogonal basis. Take W = span ⁡ { ( 1 , 1 , 0 ) , ( 1 , 0 , 1 ) } and b = ( 1 , 2 , 3 ) , and apply the projection formula directly to the given basis:

⟨ b , v 1 ⟩ ⟨ v 1 , v 1 ⟩ v 1 + ⟨ b , v 2 ⟩ ⟨ v 2 , v 2 ⟩ v 2 = 3 2 ( 1 , 1 , 0 ) + 4 2 ( 1 , 0 , 1 ) = ( 7 2 , 3 2 , 2 ) .

The correct projection is ( 7 3 , 2 3 , 5 3 ) . The two differ, and the test that exposes it is orthogonality of the residual: with the wrong answer, b − ( 7 2 , 3 2 , 2 ) = ( − 5 2 , 1 2 , 1 ) , whose inner product with v 1 is − 5 2 + 1 2 = − 2 ≠ 0 .

The formula presumes the cross terms vanish. When they do not, each term double-counts the overlap between the basis vectors, and the result is a vector in W that is simply not the nearest one.

Dropping the denominator. For an orthogonal but not orthonormal basis, writing ∑ j ⟨ b , u j ⟩ u j scales each component by ‖ u j ‖ 2 . In the worked example that turns c 2 = 5 3 into 5 2 . The direction survives and the length does not, so the residual is not orthogonal and the point is not closest.

Calling vectors orthogonal without naming the inner product. ( 1 , 1 ) and ( 1 , − 1 ) are orthogonal under the dot product and not under ⟨ u , v ⟩ = u 1 v 1 + 2 u 2 v 2 , where their pairing is − 1 . Orthogonality is a relation between two vectors and an inner product, and a claim omitting the third is incomplete.

Reading uniqueness of the fit as uniqueness of the coefficients. When the columns of A are dependent, the projection A x ^ is still the unique closest point, but many x ^ produce it. Reporting "the least-squares solution" then names something that is not determined. The theorem gives uniqueness in W ; it says nothing about the representation.

Each drops a condition that the construction depends on, orthogonality of the basis, the normalisation, the choice of inner product, independence of the columns, and each produces an answer that looks well-formed. Only the residual check catches the first two, and only attention to the hypotheses catches the last two.

Optional enrichment (1)

Application

Fitting a line as a projection

Fit y = a + b x to the points ( 1 , 1 ) , ( 2 , 2 ) , ( 3 , 2 ) .

As a system with no solution. Demanding exact fit gives three equations in two unknowns:

A = ( 1 1 1 2 1 3 ) , x = ( a b ) , b obs = ( 1 2 2 ) .

The three points are not collinear, so b obs is not in the column space of A , a plane inside R 3 , and A x = b obs has no solution. Least squares replaces b obs with the closest point of that plane.

The normal equations. With n = 3 , ∑ x i = 6 , ∑ x i 2 = 14 , ∑ y i = 5 , ∑ x i y i = 11 :

( 3 6 6 14 ) ( a b ) = ( 5 11 ) .

The determinant is 3 ( 14 ) − 36 = 6 , so

a = 5 ( 14 ) − 6 ( 11 ) 6 = 4 6 = 2 3 , b = 3 ( 11 ) − 6 ( 5 ) 6 = 3 6 = 1 2 .

The fitted line is y = 2 3 + 1 2 x .

The residuals, and what the normal equations asserted. Fitted values at x = 1 , 2 , 3 are 7 6 , 5 3 , 13 6 , so

e = ( 1 − 7 6 , 2 − 5 3 , 2 − 13 6 ) = ( − 1 6 , 1 3 , − 1 6 ) .

Row 1 of the normal equations says ∑ i e i = 0 : − 1 6 + 1 3 − 1 6 = 0 .

Row 2 says ∑ i x i e i = 0 : − 1 6 + 2 3 − 1 2 = 0 .

Those two conditions are A T e = 0 written out. The residual orthogonal to each column of A , the column of ones and the column of x values. The sum of residuals vanishing is not a coincidence of this data; it is what having an intercept column means.

The sum of squared residuals is 1 36 + 1 9 + 1 36 = 1 6 , and no other line achieves less, by the theorem, since any other choice adds a vector in the column space, orthogonal to e .

What this account does not establish. That the line is a good model, that the errors behave in any particular way, or that b = 1 2 supports a claim about how y responds to x . Those are inferential questions, and they need assumptions the geometry never makes. Projection finds the closest point in a subspace; whether that point means anything is decided elsewhere.

What the geometry explains. It accounts for why the normal equations look as they do, why an intercept forces the residuals to sum to zero, why dependent columns leave the fit unique but the coefficients not, and why the computation is stable when the columns are far from parallel and delicate when they are nearly aligned.

Next step

Practice Orthogonality and Projection

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.