Linear Transformations

What you will be able to do

Given a map between vector spaces, the learner can decide whether it is linear and justify the verdict, construct its matrix from the images of a basis when it is, and determine the kernel and image with their dimensions, checking the result against rank–nullity.

Orientation

A function on R 3 generally has to be specified everywhere: its value at one point tells you nothing about its value at another.

A linear one is fixed by three numbers' worth of choices, its values on a basis. Everything else follows, because every vector is a combination of those three and a linear map sends combinations to the corresponding combinations of images.

That collapse is what a matrix is. Its columns are the images of the basis vectors, and multiplying by it reconstructs the map's value anywhere. Reading a matrix as the record of a transformation, rather than as a grid of numbers with a multiplication rule, is what this unit adds.

Two subspaces come with any such map: what it sends to zero, and what it reaches. Their dimensions are not independent. They sum to the dimension of the space you started from.

Definition

Dependence of the matrix on the chosen bases

The definition leaves three points implicit, and each is a place the notation misleads.

The matrix is not the transformation. T is a map between spaces; A is a table of numbers that reproduces it once bases are chosen for both sides. Change the basis and A changes while T does not. Writing "the matrix of T " without naming the bases is an abbreviation, acceptable when the standard basis is understood and misleading otherwise. This is the whole content of similarity: A and P − 1 A P record the same map read in two bases.

Column j is the image of basis vector j . Not a row. The convention follows from T ( v ) = ∑ j c j T ( v j ) with c the coordinate vector: writing that as A c requires the T ( v j ) to be columns. A matrix built with the images as rows represents a different map, the transpose's, and will produce consistent, wrong answers.

Order in a composition runs right to left. T ∘ S means apply S first, and its matrix is A B with A for T and B for S . The notation reverses the reading order, and since matrix multiplication is not commutative, getting it backwards gives a genuinely different map rather than a harmless variant.

Why kernel and image are subspaces. Both follow from the three conditions of the previous unit, and each takes one line.

For the kernel: T ( 0 ) = 0 , so 0 ∈ ker ⁡ T . If T ( u ) = 0 and T ( v ) = 0 then T ( u + v ) = T ( u ) + T ( v ) = 0 , and T ( a v ) = a T ( v ) = a ⋅ 0 = 0 . Closed under both.

For the image: 0 = T ( 0 ) is reached. If w 1 = T ( u 1 ) and w 2 = T ( u 2 ) then w 1 + w 2 = T ( u 1 + u 2 ) , and a w 1 = T ( a u 1 ) . Closed under both.

Neither argument mentions coordinates, so both hold for maps between spaces of matrices, polynomials or functions.

Figure

A coordinate grid before and after a linear map

The grid before and after, with the basis vectors and their images

The grid before and after. This nonsingular example sends lines to lines and keeps parallel families parallel, which is linearity seen rather than asserted, and the origin does not move because T ( 0 ) = 0 .

Stated for an arbitrary linear map the fact is weaker: a line maps to a line or to a point, collapsing when its direction lies in the kernel, and distinct parallel lines can land on the same image. Invertibility is what rules those cases out here.

Read the matrix off the picture: e 1 lands on the first column and e 2 on the second. Everything else follows, because a linear map is fixed by what it does to a basis.

Keep this image. The area factor of the transformed grid is the determinant, the directions it does not turn are the eigenvectors, and the ellipse a circle becomes under it carries the singular values. Three later units read this one picture.

Example

A catalogue of maps on the plane

Each map below is linear, each is given by its matrix, and each is worth recognising on sight.

Scaling by k . ( k 0 0 k ) . Stretches everything away from the origin by a factor k . For k = 0 it is the zero map, whose kernel is all of R 2 and whose image is { 0 } : nullity 2, rank 0. For k = 1 it is the identity: nullity 0, rank 2.

Reflection in the x -axis. ( 1 0 0 − 1 ) , sending ( x , y ) to ( x , − y ) . Determinant − 1 , kernel trivial, image everything. Applying it twice gives the identity.

Rotation by θ . ( cos ⁡ θ − sin ⁡ θ sin ⁡ θ cos ⁡ θ ) . The columns are the images of ( 1 , 0 ) and ( 0 , 1 ) , which land at angle θ from where they started. Determinant cos 2 ⁡ θ + sin 2 ⁡ θ = 1 , so nothing is collapsed and nothing is scaled.

Projection onto the x -axis. ( 1 0 0 0 ) , sending ( x , y ) to ( x , 0 ) . The kernel is the y -axis and the image is the x -axis: nullity 1, rank 1, summing to 2. Applying it twice changes nothing after the first time, which is what makes it a projection.

Shear. ( 1 s 0 1 ) , sending ( x , y ) to ( x + s y , y ) . Horizontal lines slide by an amount proportional to their height. Determinant 1, so areas are preserved even though the picture is distorted. A reminder that the determinant measures area scaling, not distortion.

Map det nullityrank
zero020
projection011
identity102
rotation102
reflection − 1 02
shear102

The determinant is zero exactly when the kernel is nontrivial, and in the plane the three rank values 0, 1 and 2 are the only ones available, collapse everything, collapse a line, collapse nothing. Rank–nullity accounts for each row: the entries in the last two columns sum to 2 throughout.

Two that are not on the list. Translation by a fixed vector is not linear, since it moves the origin. "Reflection in the line y = 1 " is not linear for the same reason; only reflections in lines through the origin qualify.

Worked example

From a formula to a matrix, kernel and image

The map. T : R 2 → R 2 given by T ( x , y ) = ( x + 2 y , 3 x − y ) .

---

Step 1: test linearity. Take u = ( x 1 , y 1 ) and v = ( x 2 , y 2 ) .

T ( u + v ) = ( ( x 1 + x 2 ) + 2 ( y 1 + y 2 ) , 3 ( x 1 + x 2 ) − ( y 1 + y 2 ) ) ,
T ( u ) + T ( v ) = ( ( x 1 + 2 y 1 ) + ( x 2 + 2 y 2 ) , ( 3 x 1 − y 1 ) + ( 3 x 2 − y 2 ) ) .

Regrouping shows these agree. For scaling, T ( a u ) = ( a x 1 + 2 a y 1 , 3 a x 1 − a y 1 ) = a ( x 1 + 2 y 1 , 3 x 1 − y 1 ) = a T ( u ) .

A numerical check on u = ( 1 , 2 ) and v = ( − 3 , 4 ) : T ( u ) = ( 5 , 1 ) , T ( v ) = ( 5 , − 13 ) , T ( u + v ) = T ( − 2 , 6 ) = ( 10 , − 12 ) , and T ( u ) + T ( v ) = ( 10 , − 12 ) . Scaling by 1 3 : T ( 1 3 , 2 3 ) = ( 5 3 , 1 3 ) = 1 3 T ( u ) . Also T ( 0 , 0 ) = ( 0 , 0 ) .

Step 2: build the matrix. Apply T to the standard basis:

T ( 1 , 0 ) = ( 1 , 3 ) , T ( 0 , 1 ) = ( 2 , − 1 ) .

These are the columns:

A = ( 1 2 3 − 1 ) .

Checking on ( 1 , 2 ) : A ( 1 2 ) = ( 1 + 4 3 − 2 ) = ( 5 1 ) , matching T ( 1 , 2 ) = ( 5 , 1 ) .

Step 3: kernel and image. det A = ( 1 ) ( − 1 ) − ( 2 ) ( 3 ) = − 7 ≠ 0 , so A is invertible: the only solution of A v = 0 is v = 0 .

ker ⁡ T = { 0 } , nullity  0 ; im ⁡ T = R 2 , rank  2 .

Rank–nullity: 0 + 2 = 2 = dim ⁡ R 2 .

---

A singular map, for contrast. K ( x , y ) = ( x + 2 y , 2 x + 4 y ) , with matrix ( 1 2 2 4 ) and det = 4 − 4 = 0 .

Kernel. x + 2 y = 0 means x = − 2 y , so every kernel element is y ( − 2 , 1 ) . Checking ( 2 , − 1 ) : K ( 2 , − 1 ) = ( 2 − 2 , 4 − 4 ) = ( 0 , 0 ) . The kernel is the line spanned by ( 2 , − 1 ) , nullity 1.

Image. K ( 1 , 0 ) = ( 1 , 2 ) and K ( 0 , 1 ) = ( 2 , 4 ) = 2 ( 1 , 2 ) . Both columns are multiples of ( 1 , 2 ) , so the image is the line spanned by ( 1 , 2 ) , rank 1.

Rank–nullity: 1 + 1 = 2 .

Both maps are linear and both are given by equally ordinary formulas. One loses nothing and reaches everything; the other collapses a line to the origin and reaches only a line. The determinant distinguishes them before either kernel is computed, and rank–nullity then fixes the second dimension once the first is known.

---

Composition, briefly. Let U ( x , y ) = ( y , x ) , with matrix B = ( 0 1 1 0 ) . Then

A B = ( 2 1 − 1 3 ) , B A = ( 3 − 1 1 2 ) ,

which differ. A B represents T ∘ U , swap first, then apply T , and checking on ( 1 , 2 ) : U ( 1 , 2 ) = ( 2 , 1 ) , T ( 2 , 1 ) = ( 4 , 5 ) , while A B ( 1 2 ) = ( 2 + 2 − 1 + 6 ) = ( 4 5 ) .

Determinants multiply: det ( A B ) = 7 = ( − 7 ) ( − 1 ) = det A ⋅ det B .

Derivation

The rank-nullity theorem

Claim. For a linear T : V → W with dim ⁡ V = n finite,

dim ⁡ ker ⁡ T + dim ⁡ im ⁡ T = n .

The construction. Let dim ⁡ ker ⁡ T = k and take a basis { u 1 , … , u k } of ker ⁡ T . A basis of a subspace extends to a basis of the whole space, so choose w 1 , … , w n − k making

{ u 1 , … , u k , w 1 , … , w n − k }

a basis of V . The claim reduces to showing { T ( w 1 ) , … , T ( w n − k ) } is a basis of im ⁡ T , which gives dim ⁡ im ⁡ T = n − k directly.

They span the image. Any element of im ⁡ T is T ( v ) for some v ∈ V . Write v in the basis:

v = ∑ i = 1 k a i u i + ∑ j = 1 n − k b j w j .

Applying T and using linearity, the first sum vanishes because each u i is in the kernel:

T ( v ) = ∑ i a i T ( u i ) + ∑ j b j T ( w j ) = 0 + ∑ j b j T ( w j ) .

So every image element is a combination of the T ( w j ) .

They are independent. Suppose ∑ j b j T ( w j ) = 0 . By linearity T ( ∑ j b j w j ) = 0 , so ∑ j b j w j ∈ ker ⁡ T and is therefore a combination of the u i :

∑ j b j w j = ∑ i a i u i ⟹ ∑ j b j w j − ∑ i a i u i = 0 .

That is a vanishing combination of the full basis of V , whose coefficients must all be zero. In particular every b j = 0 , which is independence.

Being a spanning independent set, { T ( w j ) } is a basis of im ⁡ T , so dim ⁡ im ⁡ T = n − k and the identity follows. ◼

What the proof used. Linearity, the existence of a basis, and the extension of a basis of a subspace to one of the whole space. No coordinates, no matrix, and no assumption about W beyond its being a vector space, so the theorem applies to maps between polynomial or function spaces exactly as stated.

Where it fails. Finite-dimensionality of V is essential. On the space of infinite sequences, the shift ( a 1 , a 2 , a 3 , … ) ↦ ( a 2 , a 3 , … ) is surjective with a one-dimensional kernel, and the shift the other way is injective without being surjective. Neither is possible when dim ⁡ V is finite and W = V , where the theorem makes injectivity and surjectivity equivalent.

The immediate consequence. For T : V → V with dim ⁡ V = n , injective ⟺ nullity 0 ⟺ rank n ⟺ surjective. One-sided invertibility is two-sided, which is why a square matrix with a left inverse has it as a right inverse too.

Non-example

Maps that fail linearity

A shift. S ( x , y ) = ( x + 1 , y ) . Immediately S ( 0 , 0 ) = ( 1 , 0 ) ≠ 0 , so it is not linear. Additivity fails too: S ( 1 , 2 ) + S ( − 3 , 4 ) = ( 2 , 2 ) + ( − 2 , 4 ) = ( 0 , 6 ) , while S ( − 2 , 6 ) = ( − 1 , 6 ) . Any map with a constant term fails this way, including x ↦ m x + c for c ≠ 0 . The schoolroom "linear function" is affine, not linear.

A squaring. Q ( x , y ) = ( x 2 , y ) . Here Q ( 0 , 0 ) = ( 0 , 0 ) , so the origin test passes and gives no information. Homogeneity is what fails: Q ( 2 ⋅ ( 1 , 1 ) ) = Q ( 2 , 2 ) = ( 4 , 2 ) , while 2 ⋅ Q ( 1 , 1 ) = 2 ( 1 , 1 ) = ( 2 , 2 ) . The map is nonlinear in a way the origin never reveals, which is why the origin test rules out rather than rules in.

A norm. N ( x , y ) = x 2 + y 2 , as a map R 2 → R . It fixes the origin and it satisfies N ( a v ) = | a | N ( v ) , but | a | , not a . Taking a = − 1 : N ( − v ) = N ( v ) while − N ( v ) is negative for v ≠ 0 . Additivity fails too, since N ( 1 , 0 ) + N ( 0 , 1 ) = 2 while N ( 1 , 1 ) = 2 .

Transposition, which is linear. T ( M ) = M T on 2 × 2 matrices. It looks like a rearrangement rather than an algebraic operation, and it is linear: ( M + N ) T = M T + N T and ( a M ) T = a M T . Its kernel is { 0 } , since only the zero matrix transposes to zero, so by rank–nullity its image is all of M 2 × 2 , nullity 0 , rank 4 , summing to dim ⁡ M 2 × 2 = 4 .

Linearity is not about looking like a formula with no exponents. It is two equations that either hold for all inputs or do not, and the cases divide by how they fail: S at the origin, Q under scaling only, N under negative scaling and addition both. A map is linear when it respects the operations, whatever it looks like, which is why transposition qualifies and squaring does not.

Optional enrichment (1)

Application

Differentiation as a linear map

Let P 3 be the polynomials of degree at most 3, and D : P 3 → P 3 differentiation. Writing p = a 0 + a 1 x + a 2 x 2 + a 3 x 3 by its coefficient list ( a 0 , a 1 , a 2 , a 3 ) :

D ( p ) = a 1 + 2 a 2 x + 3 a 3 x 2 ⟷ ( a 1 , 2 a 2 , 3 a 3 , 0 ) .

Linear. ( p + q ) ′ = p ′ + q ′ and ( a p ) ′ = a p ′ are the sum and constant-multiple rules, which is exactly the definition. The rules learned as calculus facts are the statement that D is a linear map.

Its matrix. Apply D to the basis 1 , x , x 2 , x 3 : the images are 0 , 1 , 2 x , 3 x 2 , so in coordinates the columns are ( 0 , 0 , 0 , 0 ) , ( 1 , 0 , 0 , 0 ) , ( 0 , 2 , 0 , 0 ) , ( 0 , 0 , 3 , 0 ) :

A = ( 0 1 0 0 0 0 2 0 0 0 0 3 0 0 0 0 ) .

Checking on x 3 , coordinates ( 0 , 0 , 0 , 1 ) : A ( 0 , 0 , 0 , 1 ) T = ( 0 , 0 , 3 , 0 ) T , which is 3 x 2 .

Kernel and image. D ( p ) = 0 exactly when p is constant, so ker ⁡ D is the constants, nullity 1. The image is the polynomials of degree at most 2, every such polynomial has an antiderivative in P 3 , so rank 3. Rank–nullity: 1 + 3 = 4 = dim ⁡ P 3 .

What this makes possible. Two facts usually met separately become one statement. That an antiderivative is determined only up to a constant is the kernel being one-dimensional. That every polynomial of degree at most 2 has an antiderivative is the image being everything of that degree. The + C in an indefinite integral is a coset of ker ⁡ D .

Beyond polynomials. The same reading applies to linear differential operators. The solutions of y ″ + y = 0 are the kernel of L ( y ) = y ″ + y , which is why they form a space spanned by sin and cos , and why a general solution is a combination of two independent ones. The solutions of y ″ + y = f are a coset of that kernel: one particular solution plus anything in it. The structure of ODE solution sets is a statement about kernels, and nothing about it is specific to differentiation.

The limit worth naming. D on all polynomials, or on smooth functions, is a map on an infinite-dimensional space. It still has a one-dimensional kernel, but the rank–nullity bookkeeping no longer applies, and D is surjective there while having a nontrivial kernel, impossible for a map from a finite-dimensional space to itself.

Next step

Practice Linear Transformations

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.