Shrinkage, and the Trade It Makes

Why least squares returns wild coefficients when two predictors carry nearly the same information, how adding λ to the diagonal of X T X steadies them, what the lasso does differently in setting coefficients to exactly zero, and why a penalty is a trade rather than a repair. Least squares already minimises the unpenalised training sum of squares, so no penalised fit can have a lower training sum of squares, and in practice it is higher. The case for a penalty is therefore made on held-out error.

Definition

Ordinary least squares solves the normal equations X T X β = X T y . When the columns of X are nearly linearly dependent, X T X is nearly singular: it has at least one very small singular value, equivalently a large condition number, and inverting it amplifies small changes in y into large changes in β . Near-singularity is diagnosed by small singular values or a large condition number rather than by a small determinant, which is scale-dependent and shrinks with dimension for reasons unrelated to collinearity.

Ridge regression adds a multiple of the identity before inverting:

β ^ ridge = ( X T X + λ I ) − 1 X T y , λ > 0 ,

which is the minimiser of ‖ y − X β ‖ 2 + λ ‖ β ‖ 2 2 . Adding λ to each diagonal entry moves the matrix away from singularity, so the solution exists and is stable even when the unpenalised one is not.

The lasso replaces the squared penalty with an absolute one, minimising ‖ y − X β ‖ 2 + λ ‖ β ‖ 1 . This has no general closed-form matrix solution and is computed numerically, by coordinate descent or a path algorithm. It differs from ridge in one consequential way: it can set coefficients to exactly zero, so it selects predictors as well as shrinking them.

In the formulation used here the predictors are centred and standardised and the intercept is not penalised, which is the usual convention rather than a mathematical requirement of regularisation; penalising the intercept would shrink the fitted response toward zero for no substantive reason. The location of the response is not what is being regularised. Both are also scale-dependent: a predictor measured in smaller units has a larger coefficient and is penalised more, so predictors are standardised first.

What a penalty costs. Ordinary least squares minimises the unpenalised training sum of squares by definition, so no penalised fit can achieve a lower one; the training fit is non-decreasing relative to the OLS minimum and in practice strictly worse. The penalty also biases the estimates toward zero in general. What it returns is variance: the coefficients move less when the data move.

Assumptions and scope

  • In the formulation used here the predictors are centred and standardised and the intercept is excluded from the penalty. This is the usual convention rather than a requirement of regularisation itself: penalising the intercept would pull the fitted response toward zero rather than toward a simpler relationship, which is not what regularisation is for.

  • Both penalties are scale-dependent, so predictors are standardised before fitting. A predictor recorded in smaller units carries a larger coefficient and is therefore penalised more heavily for no substantive reason.

  • λ cannot be chosen by error on the estimation data: that error is minimised at λ = 0 by construction and rises monotonically as the penalty grows. Selection requires held-out or resampled error, and the same reuse caution applies as for any tuned parameter.

  • A positive penalty shrinks the coefficient vector toward zero, introducing bias in general and increasing it with λ ; the exception is the degenerate case where the unpenalised solution already minimises the penalty. Reporting a penalised coefficient as an estimate of a structural parameter, rather than as a component of a prediction rule, misrepresents what was computed.

  • The lasso's selection among correlated predictors is unstable: which of several near-identical predictors survives can change with a small perturbation of the data, so the identity of the selected variable carries less information than its inclusion suggests.

  • The figures in this unit come from exact rational arithmetic on the stated five-observation design, with the ridge solutions verified against a direct grid minimisation of the penalised objective and the lasso solutions from coordinate descent run to convergence.

Worked material

Example

The ridge path, the lasso path, and two endpoints

The ridge path on the unit's design. Solving ( X T X + λ I ) β = X T y at increasing λ :

λ β 1 β 2 sum ‖ β ‖ 2
0 + 2.605263 − 0.631579 1.973684 2.680725
0.01 + 1.784703 + 0.195467 1.980170 1.795375
0.1 + 1.134280 + 0.842805 1.977085 1.413121
0.5 + 1.004348 + 0.934783 1.939130 1.372054
1 + 0.965969 + 0.926702 1.892670 1.338608
2 + 0.914784 + 0.891170 1.805955 1.277112
5 + 0.800507 + 0.787111 1.587618 1.122655
10 + 0.665220 + 0.656121 1.321341 0.934351

Three things are visible. The sign flip is repaired almost immediately, by λ = 0.01 the second coefficient is already positive. The two coefficients converge toward each other, which is what splitting the effect between duplicate measurements looks like. And the sum decays from 1.973684 toward zero, so the bias grows steadily: at λ = 10 the fitted total effect is 1.321341 where the truth is about 2 .

Notice also that most of the stabilising happens between λ = 0 and λ = 0.1 , while most of the bias accumulates after it. That is the shape of the trade on this data, and it is why λ is worth choosing rather than setting large for safety.

The lasso path on the same design.

λ β 1 β 2 nonzero
0 + 2.605263 − 0.631579 2
0.01 + 1.979000 0 1
0.5 + 1.930000 0 1
1 + 1.880000 0 1
5 + 1.480000 0 1
10 + 0.980000 0 1
19.6 + 0.020000 0 1
20 0 0 0

The second predictor is dropped at λ ≈ 0.006030 and never returns. After that the surviving coefficient declines linearly, absorbing the whole effect, and staying nearer the true total of 2 than the ridge sum does at every comparable λ . At λ = 20 both coefficients are zero and the model predicts the mean.

The two endpoints. At λ = 0 both estimators are ordinary least squares, with all of its instability. As λ → ∞ , ridge shrinks the coefficients toward zero without reaching it while the lasso arrives at exactly zero, and in both limits the model predicts the response mean for every observation, having learned nothing.

Every useful model is somewhere between, and nothing in these tables says where.

---

Both fix the sign problem. They disagree about what to do afterwards: ridge keeps two measurements and reports each as carrying about half the effect, while the lasso keeps one and attributes the whole effect to it. On this data the lasso's total is closer to the truth, which is a fact about a design where the second predictor really is redundant, not a general ranking. When two correlated predictors measure genuinely different things, discarding one is discarding information.

Contrast

Pairs that differ in one respect

Ridge against lasso, same data, same λ = 0.506030 .

ridgelasso
β 1 + 1.003705 + 1.929397
β 2 + 0.934853 0
predictors retained 2 1
sum 1.938558 1.929397

Identical inputs and identical penalty strength, and the answers differ categorically rather than by degree. The cause is the shape of the constraint region, a disc against a diamond with corners on the axes, so no adjustment of λ turns one into the other.

A fit that is optimal against a fit that is usable.

Least squares gives ( + 2.605263 , − 0.631579 ) and attains the lowest possible squared error on this data. Ridge at λ = 1 gives ( + 0.965969 , + 0.926702 ) and is strictly worse there. The second is the one to report, and the reason has nothing to do with the fit: it is that the first moves by 0.269605 when a single response value shifts by 0.2 , against 0.027302 for the second.

Optimality on the estimation data and usefulness beyond it are different properties, and this is the clearest case in the curriculum where they point in opposite directions.

A determined combination against determined coefficients.

The unpenalised coefficients are wild, yet their sum is 1.973684 , close to the true total. The data determines what x 1 and x 2 jointly contribute and says almost nothing about how to divide it. So the instability is not uniform across everything one might want to know: a prediction from this model is far more trustworthy than either coefficient in it.

This is why the diagnosis matters. If only predictions are needed, the unpenalised model may be adequate; if the coefficients are to be interpreted, it is not.

Bias that is a defect against bias that was purchased.

An omitted confounder biases a coefficient and gives nothing back; the estimate is simply wrong about the quantity it names. A penalty biases the coefficients toward zero, the sum falling from 1.973684 to 1.892670 at λ = 1 , and returns a tenfold reduction in variance. Both are bias; only one is a trade, and only one was chosen deliberately.

A penalty strength chosen by training error against one chosen by held-out error.

Training error is minimised at λ = 0 and rises monotonically, so any selection rule based on it returns the unpenalised estimator. The very thing being avoided. Held-out error falls and then rises, so it has an interior minimum. One procedure cannot produce an answer; the other can.

Two correlated predictors that are duplicates against two that are not.

Here x 2 is x 1 with three values nudged, so dropping it loses essentially nothing and the lasso's single-predictor answer is honest. Had they been distinct quantities that merely correlate in this sample, height and weight, say, the same procedure would discard a real predictor because of an accident of the sample. The arithmetic is identical; what differs is whether the redundancy is structural or incidental, and no penalty can tell the difference.

Common errors

Common misconception

That a penalty corrects a defective fit, so a penalised model is simply better than the unpenalised one. A penalty makes the fit on the estimation data strictly worse, that is what it is for. On the two-predictor design of this unit, ordinary least squares gives the lowest possible residual sum of squares by construction, and ridge at λ = 1 gives coefficients ( + 0.965969 , + 0.926702 ) whose residuals are larger. What is gained is stability: perturbing one response value by 0.2 moves the least-squares coefficients by 0.269605 and the ridge coefficients by 0.027302 , a factor of 9.875079 . The penalty also introduces bias, the ridge coefficients sum to 1.892670 where the unpenalised pair sums to 1.973684 and the signal being estimated sums to 2 , so shrinkage pulls the total toward zero and further from the truth even as it stops the individual coefficients from swinging. The trade is variance down, bias up, and it is a trade rather than a repair: reporting one side of it describes a different procedure from the one that was run.

Common misconception

That ridge and lasso are two ways of doing the same thing, so either will drop a redundant predictor if the penalty is strong enough. The ridge penalty shrinks coefficients toward zero without reaching it: on the design in this unit, ridge at λ = 0.506030 returns ( + 1.003705 , + 0.934853 ) , and raising λ to 10 gives ( + 0.665220 , + 0.656121 ) , smaller, still both nonzero, and converging toward each other rather than toward elimination. The lasso at that same λ = 0.506030 returns ( + 1.929397 , 0 ) , having set the second coefficient to exactly zero; on this data it does so from λ ≈ 0.006030 onward. The difference is the shape of the constraint region, not the strength of the penalty, so no amount of ridge penalty performs selection. Which behaviour is wanted depends on the question: lasso yields one predictor to report and discards the information that two measurements agreed, while ridge keeps both and splits the signal between them, which is the accurate description when the predictors are genuinely two measurements of one thing.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.