Module 6 of 8 · Lesson 1 of 1
Shrinkage, and the Trade It Makes
Ridge and lasso penalties, and selecting the penalty strength by held-out error.
What you will be able to do
The learner can explain why least squares becomes unstable when predictors are nearly collinear, compute ridge and lasso solutions on a small design, state what each penalty does differently to a coefficient that duplicates another, and justify a penalty strength by held-out error rather than by the fit it produces on the data it was estimated from.
Orientation
Collinearity and unstable least-squares coefficients
Five observations, two predictors. The second is the first with a few values nudged by a tenth. They correlate at
Least squares returns
One predictor is credited with more than twice the effect that is there, and the other, nearly indistinguishable from it, rising just as the response rises, is assigned a negative coefficient.
Nothing is broken. This is the exact minimiser of squared error, and recomputing it by a different route gives the same answer to ten decimal places. The fit is optimal and the coefficients are not usable.
The cause is visible in the matrix being inverted:
The repair is not a better algorithm. It is to change what is being minimised: add a term penalising the size of the coefficients, and the cancelling pair stops being attractive. At
What that costs is the part most easily left out. The penalised fit is worse on the data it was estimated from, necessarily, since the unpenalised one was optimal there, and its coefficients are biased toward zero. A penalty trades accuracy for stability, and whether the exchange was worth making cannot be read off the fit that made it.
This unit covers where the instability comes from, how the two standard penalties differ, one shrinks, the other also selects, and why the penalty strength has to be argued for on data that was held back.
Definition
What each piece of the definition commits you to
The canonical statement above gives the two estimators. What follows is what each choice inside them decides, since those are the decisions a reader of a fitted model needs back.
Why
| Quantity | Unpenalised | With penalty |
|---|---|---|
| matrix inverted | ||
| smallest eigenvalue | may be near | at least |
| training error | minimal | strictly larger |
| bias | none, under the model | grows with |
| variance | large when collinear | reduced |
| a coefficient can be exactly | no | lasso yes, ridge no |
The two penalties are not two strengths of one idea.
What is excluded, and why it matters. The intercept is not penalised. Shrinking it would drag fitted values toward zero, which regularisation is not for. The location of the response is not what is unstable. This is why the data is centred first: centring makes the intercept separable so that leaving it out of the penalty is well defined.
Standardisation is part of the estimator, not preparation for it. A predictor recorded in smaller units carries a proportionally larger coefficient, and a larger coefficient is penalised more. So the same data in different units gives different answers, and the choice of scale is a statement about how much each predictor should be allowed to matter.
What
What the estimates are, afterwards. They are biased for every
Intuition
An impossible question, answered by changing the question
When two predictors say almost the same thing, least squares is being asked something it cannot answer. If
That is how a predictor rising with the response acquires a coefficient of
The instability is a direct consequence. Nearly identical columns make
Adding a penalty changes what counts as a good fit. Score a solution on squared error plus the size of its coefficients, and the cancelling pair suddenly looks expensive:
The tie was never broken by evidence; it was broken by a preference the analyst supplied. That is the accurate account of what regularisation does.
Why the lasso eliminates and ridge does not. Picture the set of coefficient vectors a penalty budget allows. For ridge it is a disc; for lasso, a diamond with corners on the axes. The fit's error contours grow outward until they touch that set, and where they touch is the solution. A disc can be touched almost anywhere, so a coefficient shrinks toward zero without arriving. A diamond is most easily touched at a corner, and a corner is exactly a point where one coefficient is zero.
On this data the lasso drops the second predictor from
Shrinkage is a trade, not a free improvement. Reading it as repairing a broken estimate drops the cost from view. The penalised fit is worse on the data it was estimated from, always, necessarily, and the coefficients are biased toward zero: the ridge pair sums to
Example
The ridge path, the lasso path, and two endpoints
The ridge path on the unit's design. Solving
| sum | ||||
|---|---|---|---|---|
Three things are visible. The sign flip is repaired almost immediately, by
Notice also that most of the stabilising happens between
The lasso path on the same design.
| nonzero | |||
|---|---|---|---|
The second predictor is dropped at
The two endpoints. At
Every useful model is somewhere between, and nothing in these tables says where.
---
Both fix the sign problem. They disagree about what to do afterwards: ridge keeps two measurements and reports each as carrying about half the effect, while the lasso keeps one and attributes the whole effect to it. On this data the lasso's total is closer to the truth, which is a fact about a design where the second predictor really is redundant, not a general ranking. When two correlated predictors measure genuinely different things, discarding one is discarding information.
Procedure
Fitting a penalised model, and defending the penalty
To diagnose whether a penalty is called for.
- Form
and look at its determinant relative to the size of its diagonal entries. A determinant of against diagonals near is a warning; a determinant near the product of the diagonals is not. - Compute the correlations between predictor columns. Anything above about
deserves attention before the coefficients are read. - Perturb and refit. Change one response value slightly, solve again, and measure how far the coefficients moved. This is the direct evidence, and it costs one extra solve.
- Read the signs against what is known. A coefficient whose sign contradicts the relationship visible in a scatter of that predictor against the response is a symptom, not a finding.
To fit a ridge model by hand.
- Centre the response and the predictors, and standardise the predictors so each has comparable spread. Record that this was done; the coefficients are not comparable across scalings.
- Form
, addingto each diagonal entry only, leaving the off-diagonals alone. - Solve
. - Do not penalise the intercept. Centring in step 1 is what makes this well defined.
- Repeat at several
and tabulate the path. One shows a point; the path shows the behaviour.
To choose
- Split the data, or set up
-fold resampling, before looking at any penalised fit. - For each candidate
, fit on the training part and measure error on the held-out part. - Choose the
minimising held-out error, or the largest within one standard error of that minimum if a simpler model is wanted, and say which rule was used, since they give different answers. - Never choose
by training error. It is minimised at by construction, so the procedure would always return the unpenalised estimator. - Do not reuse the held-out data to report performance. The
was chosen using it, so error measured there is optimistic for the same reason training error is.
To choose between ridge and lasso.
- Ask what the answer should look like. If a subset of predictors is wanted, the lasso produces one; ridge will not, at any
. - Ask whether the correlated predictors are duplicates or distinct measurements. Dropping one duplicate is reasonable; dropping one of two genuinely different quantities that happen to correlate in this sample discards information.
- If the lasso is used, check the stability of what it selected by refitting on resamples. Which of several near-identical predictors survives can change, and an unstable selection should not be reported as a finding about which variable matters.
Checks. Confirm the penalised determinant is substantially larger than the unpenalised one, if not,
Worked example
One design, three estimators, and a perturbation
The data. Five observations, already centred, two predictors.
Step 1: the normal equations.
The determinant is
That determinant is the warning. Everything below follows from it.
Step 2: least squares. Solving by Cramer's rule,
A predictor that rises with the response has been assigned a negative coefficient. Note the sum:
Step 3: ridge at
The determinant is now
Their sum is
Verification. Minimising
Step 4: the lasso. At
The second coefficient is exactly zero. The predictor has been dropped. On this data that happens from
At that same
Step 5: the perturbation. Increase the first response from
| before | after | displacement | |
|---|---|---|---|
| least squares | |||
| ridge, |
The unpenalised coefficients move almost ten times as far in response to the same nudge. This is the variance the penalty removed, stated as a number rather than as a property.
The unpenalised fit is optimal on this data and unusable beyond it. The penalty makes the fit worse here, necessarily, and biases the total from
Contrast
Pairs that differ in one respect
Ridge against lasso, same data, same
| ridge | lasso | |
|---|---|---|
| predictors retained | ||
| sum |
Identical inputs and identical penalty strength, and the answers differ categorically rather than by degree. The cause is the shape of the constraint region, a disc against a diamond with corners on the axes, so no adjustment of
A fit that is optimal against a fit that is usable.
Least squares gives
Optimality on the estimation data and usefulness beyond it are different properties, and this is the clearest case in the curriculum where they point in opposite directions.
A determined combination against determined coefficients.
The unpenalised coefficients are wild, yet their sum is
This is why the diagnosis matters. If only predictions are needed, the unpenalised model may be adequate; if the coefficients are to be interpreted, it is not.
Bias that is a defect against bias that was purchased.
An omitted confounder biases a coefficient and gives nothing back; the estimate is simply wrong about the quantity it names. A penalty biases the coefficients toward zero, the sum falling from
A penalty strength chosen by training error against one chosen by held-out error.
Training error is minimised at
Two correlated predictors that are duplicates against two that are not.
Here
Warning
Reports that leave out the half that cost something
Presenting a penalty as a repair. The unpenalised fit is optimal on the estimation data by construction, so a penalised fit is worse there, always. It is also biased: the coefficient sum falls from
Choosing
Reporting performance on the data that chose
Interpreting penalised coefficients as effects. A penalised coefficient is a component of a prediction rule that was deliberately biased. Reading
Treating lasso selection as a finding about which variable matters. When predictors are nearly identical, which one survives is close to arbitrary and can change under a small perturbation of the data. The lasso's output names a variable; it does not establish that this variable rather than its near-twin is the operative one. Checking selection stability across resamples is what separates the two claims.
Forgetting that the penalty is scale-dependent. A predictor in smaller units carries a larger coefficient and is penalised more for it, so unstandardised fitting silently penalises variables according to the units they arrived in. Standardisation is part of the estimator, not tidying done beforehand.
Penalising the intercept. Shrinking the intercept pulls fitted values toward zero rather than toward a simpler relationship, which is not what regularisation is for. Centring the data first is what makes excluding it well defined.
---
Two that follow from the diagnosis rather than the fix.
Concluding from a wrong-signed coefficient that the predictor acts in that direction. On this data
Assuming instability contaminates everything the model produces. The unpenalised coefficients here are unusable while their sum,
---
And one about how the choice gets made in practice. Fitting at several
Application
When to penalise, and how the deliverable decides
More predictors than observations. With
Signals measured twice. Two sensors recording the same physical quantity, or two survey items asking the same question, produce exactly the structure in this unit's design. Ridge is usually the right choice: it reports each measurement as carrying part of the effect, which is what is true. Dropping one by lasso is defensible when only a prediction is wanted and cheaper data collection matters, and it discards the information that the two agreed.
Screening in a wide, sparse problem. When most predictors are believed to have no effect at all, the lasso's ability to return exact zeros is the point. It produces a shortlist, which is a different deliverable from a shrunken coefficient for every variable. The stability caveat applies with force: among correlated candidates, the chosen one is partly arbitrary, so a shortlist is a hypothesis-generating output rather than a set of findings.
Prediction where interpretation is not wanted. For a forecasting rule that will be judged only on held-out accuracy, the bias in penalised coefficients costs nothing, because no one will read them. This is the easy case, and it is worth naming because most of the cautions in this unit are about interpretation and simply do not apply.
Econometric estimation, where it usually is not wanted. When the object is a structural parameter, the effect of schooling on wages, a deliberately biased estimator is the wrong tool, since the quantity of interest is the coefficient itself and the procedure is known to pull it toward zero. The instability from collinearity is real and the response is usually different: better identification, more data, or an honest statement that the data cannot separate the two predictors.
Ridge as a stated prior. The ridge estimate is the posterior mode under a normal prior on the coefficients centred at zero, and the lasso under a Laplace prior. This makes explicit what the penalty always was: a preference about plausible coefficients, supplied by the analyst rather than found in the data. Seen that way,
---
What decides the direction. Every case above turns on the same question: is the coefficient the deliverable, or is the prediction? Where predictions are wanted, a penalty is nearly free and sometimes necessary. Where the coefficient is the answer, the bias is a cost paid in the currency of the result, and the trade needs a much stronger argument.