Orthogonality and Projection

An inner product gives a vector space lengths and angles. Orthonormal bases make coordinates into inner products, Gram–Schmidt produces one from any basis, and projecting a vector onto a subspace finds the closest point in it, which is what least squares computes.

Definition

An inner product on a real vector space V assigns to each pair u , v a number ⟨ u , v ⟩ satisfying, for all u , v , w ∈ V and a ∈ R :

Axiom
symmetry ⟨ u , v ⟩ = ⟨ v , u ⟩
linearity ⟨ a u + w , v ⟩ = a ⟨ u , v ⟩ + ⟨ w , v ⟩
positive definiteness ⟨ v , v ⟩ ≥ 0 , with equality only for v = 0

The norm is ‖ v ‖ = ⟨ v , v ⟩ , and u and v are orthogonal when ⟨ u , v ⟩ = 0 . On R n the standard inner product is the dot product, but others exist: ⟨ f , g ⟩ = ∫ 0 1 f ( x ) g ( x ) d x makes a space of functions an inner product space.

Orthogonal and orthonormal sets. A set is orthogonal when its vectors are pairwise orthogonal, orthonormal when in addition each has norm 1. An orthogonal set of nonzero vectors is automatically independent, so an orthogonal spanning set is a basis.

What an orthonormal basis supplies. For an orthonormal basis { e 1 , … , e n } , the coordinates of v are inner products:

v = ∑ j ⟨ v , e j ⟩ e j ,

so finding coordinates needs no linear system. For a merely orthogonal basis the same holds with ⟨ v , u j ⟩ / ⟨ u j , u j ⟩ as the coefficients.

Projection onto a subspace. For a subspace W with orthogonal basis { u 1 , … , u k } ,

proj W ⁡ ( b ) = ∑ j = 1 k ⟨ b , u j ⟩ ⟨ u j , u j ⟩ u j .

The residual e = b − proj W ⁡ ( b ) is orthogonal to every element of W , and proj W ⁡ ( b ) is the unique closest point of W to b . Since b = proj W ⁡ ( b ) + e with the two parts orthogonal, ‖ b ‖ 2 = ‖ proj W ⁡ ( b ) ‖ 2 + ‖ e ‖ 2 .

Gram–Schmidt. Given a basis v 1 , … , v k , set u 1 = v 1 and

u j = v j − ∑ i < j ⟨ v j , u i ⟩ ⟨ u i , u i ⟩ u i ,

subtracting from each vector its projection onto the span of those already produced. The result is an orthogonal basis of the same subspace; dividing each by its norm makes it orthonormal.

Least squares. When A x = b has no solution, the closest achievable A x is proj col ⁡ A ⁡ ( b ) , and the minimising x ^ satisfies the normal equations

A T A x ^ = A T b ,

which say exactly that the residual b − A x ^ is orthogonal to every column of A .

Assumptions and scope

  • Orthogonality depends on the inner product. Two functions orthogonal under ∫ 0 1 f g need not be under ∫ − 1 1 f g , and the word carries no meaning until the inner product is named.

  • An orthogonal set of nonzero vectors is independent, but an independent set is not generally orthogonal. Gram–Schmidt is what converts one to the other.

  • The projection formula requires an orthogonal basis of the subspace. Applying it with a non-orthogonal spanning set gives a vector that is not the projection, and the error is silent.

  • The normal equations have a unique solution exactly when the columns of A are independent. With dependent columns the least-squares problem still has a minimiser, but not a unique one.

  • Gram–Schmidt is numerically unstable in its classical form; implementations use the modified variant or a Householder QR factorisation. The mathematics is unaffected.

  • Least squares here is a geometric construction. Whether the fitted coefficients support an inferential claim is a separate question, governed by the modelling assumptions rather than by the projection.

Worked material

Example

Orthogonal sets in four different spaces

The same definition, applied wherever an inner product exists.

The standard basis of R 3 . e 1 , e 2 , e 3 pair to zero with each other and to one with themselves, so the set is orthonormal. This is why coordinates in the standard basis are so easy to read: each is already an inner product, x j = ⟨ x , e j ⟩ .

A rotated pair in R 2 . ( 1 , 1 ) and ( 1 , − 1 ) have ⟨ u , v ⟩ = 1 − 1 = 0 , so they are orthogonal, but ‖ u ‖ 2 = ‖ v ‖ 2 = 2 , so the set is orthogonal, not orthonormal. Dividing each by 2 gives

1 2 ( 1 , 1 ) , 1 2 ( 1 , − 1 ) ,

which pair to zero and have norm 1. This is the distinction that decides whether the denominators in the projection formula may be dropped.

Polynomials on [ 0 , 1 ] , with ⟨ f , g ⟩ = ∫ 0 1 f g d x . The basis { 1 , x } is not orthogonal:

⟨ 1 , x ⟩ = ∫ 0 1 x d x = 1 2 ≠ 0 .

One Gram–Schmidt step replaces x by x − 1 2 , and ⟨ 1 , x − 1 2 ⟩ = 1 2 − 1 2 = 0 . Since ‖ x − 1 2 ‖ 2 = ∫ 0 1 ( x − 1 2 ) 2 d x = 1 12 , normalising gives 12 ( x − 1 2 ) . These are the first two Legendre polynomials for this interval.

Two functions can look entirely unalike and still fail to be orthogonal; the integral decides, not the appearance.

Trigonometric functions on [ 0 , 2 π ] , with ⟨ f , g ⟩ = ∫ 0 2 π f g d x .

⟨ sin , cos ⟩ = 0 , ⟨ 1 , sin ⟩ = ⟨ 1 , cos ⟩ = 0 ,
⟨ sin ⁡ x , sin ⁡ 2 x ⟩ = ⟨ cos ⁡ x , cos ⁡ 2 x ⟩ = ⟨ sin ⁡ x , cos ⁡ 2 x ⟩ = 0 ,

while ⟨ sin , sin ⟩ = ⟨ cos , cos ⟩ = π . So { 1 , sin ⁡ x , cos ⁡ x , sin ⁡ 2 x , cos ⁡ 2 x , … } is an orthogonal set. The basis Fourier series are written in. Projecting a function onto its span is exactly how Fourier coefficients are computed, and the π in the denominators of the usual formulas is the ⟨ u j , u j ⟩ of the projection formula, not a convention.

Symmetric matrices, with ⟨ A , B ⟩ = tr ⁡ ( A T B ) . This inner product is the sum of entrywise products. The basis from the basis-and-dimension unit,

( 1 0 0 0 ) , ( 0 0 0 1 ) , ( 0 1 1 0 ) ,

is pairwise orthogonal, with squared norms 1 , 1 and 2 . Again orthogonal but not orthonormal: the third must be divided by 2 , because its single off-diagonal value appears twice.

Nothing about the objects, lists of numbers, polynomials, functions, matrices. What they share is an inner product satisfying the three axioms, and every construction in this unit is written in those terms alone. What differs is which sets count as orthogonal, and that depends entirely on the inner product chosen.

Non-example

Projecting without an orthogonal basis, and three other errors

Using the formula on a non-orthogonal basis. Take W = span ⁡ { ( 1 , 1 , 0 ) , ( 1 , 0 , 1 ) } and b = ( 1 , 2 , 3 ) , and apply the projection formula directly to the given basis:

⟨ b , v 1 ⟩ ⟨ v 1 , v 1 ⟩ v 1 + ⟨ b , v 2 ⟩ ⟨ v 2 , v 2 ⟩ v 2 = 3 2 ( 1 , 1 , 0 ) + 4 2 ( 1 , 0 , 1 ) = ( 7 2 , 3 2 , 2 ) .

The correct projection is ( 7 3 , 2 3 , 5 3 ) . The two differ, and the test that exposes it is orthogonality of the residual: with the wrong answer, b − ( 7 2 , 3 2 , 2 ) = ( − 5 2 , 1 2 , 1 ) , whose inner product with v 1 is − 5 2 + 1 2 = − 2 ≠ 0 .

The formula presumes the cross terms vanish. When they do not, each term double-counts the overlap between the basis vectors, and the result is a vector in W that is simply not the nearest one.

Dropping the denominator. For an orthogonal but not orthonormal basis, writing ∑ j ⟨ b , u j ⟩ u j scales each component by ‖ u j ‖ 2 . In the worked example that turns c 2 = 5 3 into 5 2 . The direction survives and the length does not, so the residual is not orthogonal and the point is not closest.

Calling vectors orthogonal without naming the inner product. ( 1 , 1 ) and ( 1 , − 1 ) are orthogonal under the dot product and not under ⟨ u , v ⟩ = u 1 v 1 + 2 u 2 v 2 , where their pairing is − 1 . Orthogonality is a relation between two vectors and an inner product, and a claim omitting the third is incomplete.

Reading uniqueness of the fit as uniqueness of the coefficients. When the columns of A are dependent, the projection A x ^ is still the unique closest point, but many x ^ produce it. Reporting "the least-squares solution" then names something that is not determined. The theorem gives uniqueness in W ; it says nothing about the representation.

Each drops a condition that the construction depends on, orthogonality of the basis, the normalisation, the choice of inner product, independence of the columns, and each produces an answer that looks well-formed. Only the residual check catches the first two, and only attention to the hypotheses catches the last two.

Common errors

Common misconception

The projection formula ∑ j ⟨ b , v j ⟩ ⟨ v j , v j ⟩ v j works for any spanning set of the subspace, so a basis need not be orthogonalised before projecting.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.