Shrinkage, and the Trade It Makes
Why least squares returns wild coefficients when two predictors carry nearly the same information, how adding
Definition
Ordinary least squares solves the normal equations
Ridge regression adds a multiple of the identity before inverting:
which is the minimiser of
The lasso replaces the squared penalty with an absolute one, minimising
In the formulation used here the predictors are centred and standardised and the intercept is not penalised, which is the usual convention rather than a mathematical requirement of regularisation; penalising the intercept would shrink the fitted response toward zero for no substantive reason. The location of the response is not what is being regularised. Both are also scale-dependent: a predictor measured in smaller units has a larger coefficient and is penalised more, so predictors are standardised first.
What a penalty costs. Ordinary least squares minimises the unpenalised training sum of squares by definition, so no penalised fit can achieve a lower one; the training fit is non-decreasing relative to the OLS minimum and in practice strictly worse. The penalty also biases the estimates toward zero in general. What it returns is variance: the coefficients move less when the data move.
Assumptions and scope
In the formulation used here the predictors are centred and standardised and the intercept is excluded from the penalty. This is the usual convention rather than a requirement of regularisation itself: penalising the intercept would pull the fitted response toward zero rather than toward a simpler relationship, which is not what regularisation is for.
Both penalties are scale-dependent, so predictors are standardised before fitting. A predictor recorded in smaller units carries a larger coefficient and is therefore penalised more heavily for no substantive reason.
cannot be chosen by error on the estimation data: that error is minimised at by construction and rises monotonically as the penalty grows. Selection requires held-out or resampled error, and the same reuse caution applies as for any tuned parameter.A positive penalty shrinks the coefficient vector toward zero, introducing bias in general and increasing it with
; the exception is the degenerate case where the unpenalised solution already minimises the penalty. Reporting a penalised coefficient as an estimate of a structural parameter, rather than as a component of a prediction rule, misrepresents what was computed. The lasso's selection among correlated predictors is unstable: which of several near-identical predictors survives can change with a small perturbation of the data, so the identity of the selected variable carries less information than its inclusion suggests.
The figures in this unit come from exact rational arithmetic on the stated five-observation design, with the ridge solutions verified against a direct grid minimisation of the penalised objective and the lasso solutions from coordinate descent run to convergence.
Worked material
Example
The ridge path, the lasso path, and two endpoints
The ridge path on the unit's design. Solving
| sum | ||||
|---|---|---|---|---|
Three things are visible. The sign flip is repaired almost immediately, by
Notice also that most of the stabilising happens between
The lasso path on the same design.
| nonzero | |||
|---|---|---|---|
The second predictor is dropped at
The two endpoints. At
Every useful model is somewhere between, and nothing in these tables says where.
---
Both fix the sign problem. They disagree about what to do afterwards: ridge keeps two measurements and reports each as carrying about half the effect, while the lasso keeps one and attributes the whole effect to it. On this data the lasso's total is closer to the truth, which is a fact about a design where the second predictor really is redundant, not a general ranking. When two correlated predictors measure genuinely different things, discarding one is discarding information.
Contrast
Pairs that differ in one respect
Ridge against lasso, same data, same
| ridge | lasso | |
|---|---|---|
| predictors retained | ||
| sum |
Identical inputs and identical penalty strength, and the answers differ categorically rather than by degree. The cause is the shape of the constraint region, a disc against a diamond with corners on the axes, so no adjustment of
A fit that is optimal against a fit that is usable.
Least squares gives
Optimality on the estimation data and usefulness beyond it are different properties, and this is the clearest case in the curriculum where they point in opposite directions.
A determined combination against determined coefficients.
The unpenalised coefficients are wild, yet their sum is
This is why the diagnosis matters. If only predictions are needed, the unpenalised model may be adequate; if the coefficients are to be interpreted, it is not.
Bias that is a defect against bias that was purchased.
An omitted confounder biases a coefficient and gives nothing back; the estimate is simply wrong about the quantity it names. A penalty biases the coefficients toward zero, the sum falling from
A penalty strength chosen by training error against one chosen by held-out error.
Training error is minimised at
Two correlated predictors that are duplicates against two that are not.
Here
Common errors
Common misconception
That a penalty corrects a defective fit, so a penalised model is simply better than the unpenalised one. A penalty makes the fit on the estimation data strictly worse, that is what it is for. On the two-predictor design of this unit, ordinary least squares gives the lowest possible residual sum of squares by construction, and ridge at
Common misconception
That ridge and lasso are two ways of doing the same thing, so either will drop a redundant predictor if the penalty is strong enough. The ridge penalty shrinks coefficients toward zero without reaching it: on the design in this unit, ridge at
Related units
Requires
Connected
- Choosing a Regression Form (related)
- The Singular Value Decomposition (related)