Module 6 of 8 · Lesson 1 of 1

Shrinkage, and the Trade It Makes

Ridge and lasso penalties, and selecting the penalty strength by held-out error.

What you will be able to do

The learner can explain why least squares becomes unstable when predictors are nearly collinear, compute ridge and lasso solutions on a small design, state what each penalty does differently to a coefficient that duplicates another, and justify a penalty strength by held-out error rather than by the fit it produces on the data it was estimated from.

Orientation

Collinearity and unstable least-squares coefficients

Five observations, two predictors. The second is the first with a few values nudged by a tenth. They correlate at 0.999032 . The response increases steadily with both.

Least squares returns

β 1 = + 2.605263 , β 2 = − 0.631579 .

One predictor is credited with more than twice the effect that is there, and the other, nearly indistinguishable from it, rising just as the response rises, is assigned a negative coefficient.

Nothing is broken. This is the exact minimiser of squared error, and recomputing it by a different route gives the same answer to ten decimal places. The fit is optimal and the coefficients are not usable.

The cause is visible in the matrix being inverted: X T X has determinant 0.19 , close enough to zero that inverting it magnifies small features of the data into large swings in the answer. Change one response value by 0.2 and the coefficients move by 0.269605 . A displacement comparable to the quantities being estimated.

The repair is not a better algorithm. It is to change what is being minimised: add a term penalising the size of the coefficients, and the cancelling pair stops being attractive. At λ = 1 the same data gives ( + 0.965969 , + 0.926702 ) . Two nearly equal positive coefficients, which is the accurate reading when two predictors are two measurements of one thing. Under the same perturbation these move by 0.027302 , less by a factor of 9.875079 .

What that costs is the part most easily left out. The penalised fit is worse on the data it was estimated from, necessarily, since the unpenalised one was optimal there, and its coefficients are biased toward zero. A penalty trades accuracy for stability, and whether the exchange was worth making cannot be read off the fit that made it.

This unit covers where the instability comes from, how the two standard penalties differ, one shrinks, the other also selects, and why the penalty strength has to be argued for on data that was held back.

Definition

What each piece of the definition commits you to

The canonical statement above gives the two estimators. What follows is what each choice inside them decides, since those are the decisions a reader of a fitted model needs back.

Why λ I and not something else. Adding λ to every diagonal entry shifts each eigenvalue of X T X up by exactly λ . Near-collinearity means a small eigenvalue; a small eigenvalue means a large reciprocal in the inverse; a large reciprocal means small changes in y arriving in β magnified. Raising the floor under the eigenvalues is precisely what caps that magnification, which is why the fix takes this form rather than being an arbitrary nudge toward smaller numbers.

QuantityUnpenalisedWith penalty λ
matrix inverted X T X X T X + λ I
smallest eigenvaluemay be near 0 at least λ
training errorminimalstrictly larger
biasnone, under the modelgrows with λ
variancelarge when collinearreduced
a coefficient can be exactly 0 nolasso yes, ridge no

The two penalties are not two strengths of one idea. ‖ β ‖ 2 2 and ‖ β ‖ 1 differ in the shape of the region they confine β to, and that shape decides whether solutions land on an axis. This is a qualitative difference no choice of λ erases, so the two are alternatives rather than points on a scale.

What is excluded, and why it matters. The intercept is not penalised. Shrinking it would drag fitted values toward zero, which regularisation is not for. The location of the response is not what is unstable. This is why the data is centred first: centring makes the intercept separable so that leaving it out of the penalty is well defined.

Standardisation is part of the estimator, not preparation for it. A predictor recorded in smaller units carries a proportionally larger coefficient, and a larger coefficient is penalised more. So the same data in different units gives different answers, and the choice of scale is a statement about how much each predictor should be allowed to matter.

What λ cannot be chosen by. Training error is minimised at λ = 0 by construction and rises monotonically thereafter. Any procedure choosing λ to improve the fit on the estimation data will choose zero, which is the estimator the penalty exists to avoid. Selection therefore requires error measured on data held back.

What the estimates are, afterwards. They are biased for every λ > 0 , and knowingly so. A penalised coefficient is a component of a prediction rule; reading it as an estimate of a structural parameter reports a number the procedure never claimed to provide.

Intuition

An impossible question, answered by changing the question

When two predictors say almost the same thing, least squares is being asked something it cannot answer. If x 1 and x 2 agree to within a rounding error, then 3 x 1 − x 2 and x 1 + x 2 are nearly the same vector, so the two coefficient pairs fit the data nearly equally well. The procedure has no basis for preferring either, and it settles the tie on whichever matches the noise marginally better.

That is how a predictor rising with the response acquires a coefficient of − 0.631579 . The number is not describing the predictor. It is the other half of a cancelling pair, and the cancellation is what the data supports.

The instability is a direct consequence. Nearly identical columns make X T X nearly singular, determinant 0.19 here, and inverting a nearly singular matrix multiplies small perturbations by something large. Moving one response value by 0.2 moves the coefficients by 0.269605 , which is not a rounding artefact but a swing comparable to the effects being estimated. Another sample from the same process would give a visibly different answer.

Adding a penalty changes what counts as a good fit. Score a solution on squared error plus the size of its coefficients, and the cancelling pair suddenly looks expensive: ( + 2.605263 , − 0.631579 ) has a much larger squared length than ( + 0.965969 , + 0.926702 ) , while fitting only slightly better. The second wins, and it happens to be the description a person would give. Two measurements of one quantity, each carrying about half the effect.

The tie was never broken by evidence; it was broken by a preference the analyst supplied. That is the accurate account of what regularisation does.

Why the lasso eliminates and ridge does not. Picture the set of coefficient vectors a penalty budget allows. For ridge it is a disc; for lasso, a diamond with corners on the axes. The fit's error contours grow outward until they touch that set, and where they touch is the solution. A disc can be touched almost anywhere, so a coefficient shrinks toward zero without arriving. A diamond is most easily touched at a corner, and a corner is exactly a point where one coefficient is zero.

On this data the lasso drops the second predictor from λ ≈ 0.006030 onward. Ridge never does: even at λ = 10 it returns ( + 0.665220 , + 0.656121 ) , both alive, converging toward each other rather than toward elimination.

Shrinkage is a trade, not a free improvement. Reading it as repairing a broken estimate drops the cost from view. The penalised fit is worse on the data it was estimated from, always, necessarily, and the coefficients are biased toward zero: the ridge pair sums to 1.892670 where the true signal sums to 2 . What is gained is stability, that displacement falling from 0.269605 to 0.027302 . A trade, both halves real, and the only evidence about whether it was worth making lies in data the fit has not seen.

Example

The ridge path, the lasso path, and two endpoints

The ridge path on the unit's design. Solving ( X T X + λ I ) β = X T y at increasing λ :

λ β 1 β 2 sum ‖ β ‖ 2
0 + 2.605263 − 0.631579 1.973684 2.680725
0.01 + 1.784703 + 0.195467 1.980170 1.795375
0.1 + 1.134280 + 0.842805 1.977085 1.413121
0.5 + 1.004348 + 0.934783 1.939130 1.372054
1 + 0.965969 + 0.926702 1.892670 1.338608
2 + 0.914784 + 0.891170 1.805955 1.277112
5 + 0.800507 + 0.787111 1.587618 1.122655
10 + 0.665220 + 0.656121 1.321341 0.934351

Three things are visible. The sign flip is repaired almost immediately, by λ = 0.01 the second coefficient is already positive. The two coefficients converge toward each other, which is what splitting the effect between duplicate measurements looks like. And the sum decays from 1.973684 toward zero, so the bias grows steadily: at λ = 10 the fitted total effect is 1.321341 where the truth is about 2 .

Notice also that most of the stabilising happens between λ = 0 and λ = 0.1 , while most of the bias accumulates after it. That is the shape of the trade on this data, and it is why λ is worth choosing rather than setting large for safety.

The lasso path on the same design.

λ β 1 β 2 nonzero
0 + 2.605263 − 0.631579 2
0.01 + 1.979000 0 1
0.5 + 1.930000 0 1
1 + 1.880000 0 1
5 + 1.480000 0 1
10 + 0.980000 0 1
19.6 + 0.020000 0 1
20 0 0 0

The second predictor is dropped at λ ≈ 0.006030 and never returns. After that the surviving coefficient declines linearly, absorbing the whole effect, and staying nearer the true total of 2 than the ridge sum does at every comparable λ . At λ = 20 both coefficients are zero and the model predicts the mean.

The two endpoints. At λ = 0 both estimators are ordinary least squares, with all of its instability. As λ → ∞ , ridge shrinks the coefficients toward zero without reaching it while the lasso arrives at exactly zero, and in both limits the model predicts the response mean for every observation, having learned nothing.

Every useful model is somewhere between, and nothing in these tables says where.

---

Both fix the sign problem. They disagree about what to do afterwards: ridge keeps two measurements and reports each as carrying about half the effect, while the lasso keeps one and attributes the whole effect to it. On this data the lasso's total is closer to the truth, which is a fact about a design where the second predictor really is redundant, not a general ranking. When two correlated predictors measure genuinely different things, discarding one is discarding information.

Procedure

Fitting a penalised model, and defending the penalty

To diagnose whether a penalty is called for.

  1. Form X T X and look at its determinant relative to the size of its diagonal entries. A determinant of 0.19 against diagonals near 10 is a warning; a determinant near the product of the diagonals is not.
  2. Compute the correlations between predictor columns. Anything above about 0.95 deserves attention before the coefficients are read.
  3. Perturb and refit. Change one response value slightly, solve again, and measure how far the coefficients moved. This is the direct evidence, and it costs one extra solve.
  4. Read the signs against what is known. A coefficient whose sign contradicts the relationship visible in a scatter of that predictor against the response is a symptom, not a finding.

To fit a ridge model by hand.

  1. Centre the response and the predictors, and standardise the predictors so each has comparable spread. Record that this was done; the coefficients are not comparable across scalings.
  2. Form X T X + λ I , adding λ to each diagonal entry only, leaving the off-diagonals alone.
  3. Solve ( X T X + λ I ) β = X T y .
  4. Do not penalise the intercept. Centring in step 1 is what makes this well defined.
  5. Repeat at several λ and tabulate the path. One λ shows a point; the path shows the behaviour.

To choose λ .

  1. Split the data, or set up k -fold resampling, before looking at any penalised fit.
  2. For each candidate λ , fit on the training part and measure error on the held-out part.
  3. Choose the λ minimising held-out error, or the largest λ within one standard error of that minimum if a simpler model is wanted, and say which rule was used, since they give different answers.
  4. Never choose λ by training error. It is minimised at λ = 0 by construction, so the procedure would always return the unpenalised estimator.
  5. Do not reuse the held-out data to report performance. The λ was chosen using it, so error measured there is optimistic for the same reason training error is.

To choose between ridge and lasso.

  1. Ask what the answer should look like. If a subset of predictors is wanted, the lasso produces one; ridge will not, at any λ .
  2. Ask whether the correlated predictors are duplicates or distinct measurements. Dropping one duplicate is reasonable; dropping one of two genuinely different quantities that happen to correlate in this sample discards information.
  3. If the lasso is used, check the stability of what it selected by refitting on resamples. Which of several near-identical predictors survives can change, and an unstable selection should not be reported as a finding about which variable matters.

Checks. Confirm the penalised determinant is substantially larger than the unpenalised one, if not, λ is too small to have done anything. Confirm the training error rose: if it fell, the penalty was applied incorrectly, since the unpenalised fit is optimal there by construction. And before reporting coefficients, state the bias: compare the penalised coefficient sum with the unpenalised one, so that what shrinkage moved is on the record alongside what it steadied.

Worked example

One design, three estimators, and a perturbation

The data. Five observations, already centred, two predictors.

x 1 x 2 y
− 2 − 2 − 3.9
− 1 − 0.9 − 2.1
0 − 0.1 0
1 1 2.1
2 2 3.9

x 2 is x 1 with three values nudged by a tenth. The response is close to 2 x 1 , so the effect being estimated totals about 2 spread across two nearly identical predictors.

Step 1: the normal equations.

X T X = ( 10 9.9 9.9 9.82 ) , X T y = ( 19.8 19.59 ) .

The determinant is 10 × 9.82 − 9.9 2 = 98.2 − 98.01 = 0.19 , against diagonal entries near 10 . The correlation between the columns is 0.999032 .

That determinant is the warning. Everything below follows from it.

Step 2: least squares. Solving by Cramer's rule,

β 1 = 9.82 × 19.8 − 9.9 × 19.59 0.19 = + 2.605263 , β 2 = 10 × 19.59 − 9.9 × 19.8 0.19 = − 0.631579 .

A predictor that rises with the response has been assigned a negative coefficient. Note the sum: + 1.973684 , close to the true total of about 2 . The combination is well determined even though neither coefficient is.

Step 3: ridge at λ = 1 . Add 1 to each diagonal entry and solve again:

( 11 9.9 9.9 10.82 ) β = ( 19.8 19.59 ) ⟹ β = ( + 0.965969 , + 0.926702 ) .

The determinant is now 11 × 10.82 − 9.9 2 = 119.02 − 98.01 = 21.01 , over a hundred times larger. Both coefficients are positive and nearly equal, splitting the effect between the two measurements.

Their sum is 1.892670 , further from 2 than the unpenalised 1.973684 . That is the bias, and it is not a defect to be explained away, shrinkage pulls the total toward zero.

Verification. Minimising 1 2 ‖ y − X β ‖ 2 + 1 2 λ ‖ β ‖ 2 over a grid of 121 × 121 candidates around this point returns the same pair, with objective 0.9798691099 at both. The analytic solution is the minimiser.

Step 4: the lasso. At λ = 0.506030 , coordinate descent converges to

β = ( + 1.929397 , 0 ) .

The second coefficient is exactly zero. The predictor has been dropped. On this data that happens from λ ≈ 0.006030 upward, so the lasso discards the duplicate almost as soon as any penalty is applied.

At that same λ = 0.506030 , ridge returns ( + 1.003705 , + 0.934853 ) : both nonzero. Identical data, identical penalty strength, and a categorical difference in the answer.

Step 5: the perturbation. Increase the first response from − 3.9 to − 3.7 , one value, by 0.2 , and refit.

beforeafterdisplacement
least squares ( + 2.605263 , − 0.631579 ) ( + 2.773684 , − 0.842105 ) 0.269605
ridge, λ = 1 ( + 0.965969 , + 0.926702 ) ( + 0.948453 , + 0.905759 ) 0.027302
0.269605 0.027302 = 9.875079 .

The unpenalised coefficients move almost ten times as far in response to the same nudge. This is the variance the penalty removed, stated as a number rather than as a property.

The unpenalised fit is optimal on this data and unusable beyond it. The penalty makes the fit worse here, necessarily, and biases the total from 1.973684 to 1.892670 . In exchange, the answer stops depending so heavily on the particular sample. Which of those matters more is not decidable from anything on this page, because every figure above was computed from the same five observations. That is what makes held-out error necessary rather than merely advisable.

Contrast

Pairs that differ in one respect

Ridge against lasso, same data, same λ = 0.506030 .

ridgelasso
β 1 + 1.003705 + 1.929397
β 2 + 0.934853 0
predictors retained 2 1
sum 1.938558 1.929397

Identical inputs and identical penalty strength, and the answers differ categorically rather than by degree. The cause is the shape of the constraint region, a disc against a diamond with corners on the axes, so no adjustment of λ turns one into the other.

A fit that is optimal against a fit that is usable.

Least squares gives ( + 2.605263 , − 0.631579 ) and attains the lowest possible squared error on this data. Ridge at λ = 1 gives ( + 0.965969 , + 0.926702 ) and is strictly worse there. The second is the one to report, and the reason has nothing to do with the fit: it is that the first moves by 0.269605 when a single response value shifts by 0.2 , against 0.027302 for the second.

Optimality on the estimation data and usefulness beyond it are different properties, and this is the clearest case in the curriculum where they point in opposite directions.

A determined combination against determined coefficients.

The unpenalised coefficients are wild, yet their sum is 1.973684 , close to the true total. The data determines what x 1 and x 2 jointly contribute and says almost nothing about how to divide it. So the instability is not uniform across everything one might want to know: a prediction from this model is far more trustworthy than either coefficient in it.

This is why the diagnosis matters. If only predictions are needed, the unpenalised model may be adequate; if the coefficients are to be interpreted, it is not.

Bias that is a defect against bias that was purchased.

An omitted confounder biases a coefficient and gives nothing back; the estimate is simply wrong about the quantity it names. A penalty biases the coefficients toward zero, the sum falling from 1.973684 to 1.892670 at λ = 1 , and returns a tenfold reduction in variance. Both are bias; only one is a trade, and only one was chosen deliberately.

A penalty strength chosen by training error against one chosen by held-out error.

Training error is minimised at λ = 0 and rises monotonically, so any selection rule based on it returns the unpenalised estimator. The very thing being avoided. Held-out error falls and then rises, so it has an interior minimum. One procedure cannot produce an answer; the other can.

Two correlated predictors that are duplicates against two that are not.

Here x 2 is x 1 with three values nudged, so dropping it loses essentially nothing and the lasso's single-predictor answer is honest. Had they been distinct quantities that merely correlate in this sample, height and weight, say, the same procedure would discard a real predictor because of an accident of the sample. The arithmetic is identical; what differs is whether the redundancy is structural or incidental, and no penalty can tell the difference.

Warning

Reports that leave out the half that cost something

Presenting a penalty as a repair. The unpenalised fit is optimal on the estimation data by construction, so a penalised fit is worse there, always. It is also biased: the coefficient sum falls from 1.973684 to 1.892670 at λ = 1 , moving away from the true total of about 2 . A report describing shrinkage as correcting a bad fit has described a different procedure from the one that was run.

Choosing λ by the fit it produces. Training error is minimised at λ = 0 and rises monotonically with the penalty, so selecting λ to improve the fit returns the unpenalised estimator every time. The choice requires data the fit has not seen; there is no way around this, and no diagnostic computed from the training data substitutes for it.

Reporting performance on the data that chose λ . Holding out a set, using it to pick λ , and then quoting the error on that same set is the reuse problem in a new place. The held-out error is optimistic because λ was selected to minimise it, which is why a separate test set, or nested resampling, is needed when both a tuned parameter and an error estimate are wanted.

Interpreting penalised coefficients as effects. A penalised coefficient is a component of a prediction rule that was deliberately biased. Reading + 0.965969 as an estimate of x 1 's effect reports a number the procedure never claimed to produce, and the bias is toward zero, so effects are systematically understated.

Treating lasso selection as a finding about which variable matters. When predictors are nearly identical, which one survives is close to arbitrary and can change under a small perturbation of the data. The lasso's output names a variable; it does not establish that this variable rather than its near-twin is the operative one. Checking selection stability across resamples is what separates the two claims.

Forgetting that the penalty is scale-dependent. A predictor in smaller units carries a larger coefficient and is penalised more for it, so unstandardised fitting silently penalises variables according to the units they arrived in. Standardisation is part of the estimator, not tidying done beforehand.

Penalising the intercept. Shrinking the intercept pulls fitted values toward zero rather than toward a simpler relationship, which is not what regularisation is for. Centring the data first is what makes excluding it well defined.

---

Two that follow from the diagnosis rather than the fix.

Concluding from a wrong-signed coefficient that the predictor acts in that direction. On this data β 2 = − 0.631579 for a predictor that plainly rises with the response. The number is the other half of a cancelling pair, and reading it as a reversed effect is reading noise as a finding.

Assuming instability contaminates everything the model produces. The unpenalised coefficients here are unusable while their sum, 1.973684 , is close to correct: the data determines the joint contribution and not its division. A model may therefore predict well while its coefficients cannot be interpreted, and whether the instability matters depends on which of the two is wanted.

---

And one about how the choice gets made in practice. Fitting at several λ , looking at the coefficient paths, and picking the value where they look reasonable is selection by appearance. The analyst's prior expectations deciding the answer while a resampling procedure is named in the report. If a prior belief about the coefficients is doing the work, it belongs in the write-up as such.

Application

When to penalise, and how the deliverable decides

More predictors than observations. With p > n the matrix X T X is singular outright, so least squares has no unique solution, infinitely many coefficient vectors fit perfectly. A ridge penalty makes the problem well posed for any λ > 0 , because adding λ to the diagonal lifts every eigenvalue above zero. Here regularisation is not an improvement on an existing estimator; it is what makes estimation possible at all. Genomic and text data routinely arrive in this shape.

Signals measured twice. Two sensors recording the same physical quantity, or two survey items asking the same question, produce exactly the structure in this unit's design. Ridge is usually the right choice: it reports each measurement as carrying part of the effect, which is what is true. Dropping one by lasso is defensible when only a prediction is wanted and cheaper data collection matters, and it discards the information that the two agreed.

Screening in a wide, sparse problem. When most predictors are believed to have no effect at all, the lasso's ability to return exact zeros is the point. It produces a shortlist, which is a different deliverable from a shrunken coefficient for every variable. The stability caveat applies with force: among correlated candidates, the chosen one is partly arbitrary, so a shortlist is a hypothesis-generating output rather than a set of findings.

Prediction where interpretation is not wanted. For a forecasting rule that will be judged only on held-out accuracy, the bias in penalised coefficients costs nothing, because no one will read them. This is the easy case, and it is worth naming because most of the cautions in this unit are about interpretation and simply do not apply.

Econometric estimation, where it usually is not wanted. When the object is a structural parameter, the effect of schooling on wages, a deliberately biased estimator is the wrong tool, since the quantity of interest is the coefficient itself and the procedure is known to pull it toward zero. The instability from collinearity is real and the response is usually different: better identification, more data, or an honest statement that the data cannot separate the two predictors.

Ridge as a stated prior. The ridge estimate is the posterior mode under a normal prior on the coefficients centred at zero, and the lasso under a Laplace prior. This makes explicit what the penalty always was: a preference about plausible coefficients, supplied by the analyst rather than found in the data. Seen that way, λ measures how strongly that preference is held, which is why arguing for a λ means arguing for a belief.

---

What decides the direction. Every case above turns on the same question: is the coefficient the deliverable, or is the prediction? Where predictions are wanted, a penalty is nearly free and sometimes necessary. Where the coefficient is the answer, the bias is a cost paid in the currency of the result, and the trade needs a much stronger argument.

Next step

Practice Shrinkage, and the Trade It Makes

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Association Rules, and What Confidence Leaves Out

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.