Practice: Shrinkage, and the Trade It Makes

Direct application

For a centred two-predictor design,

X T X = ( 10 9.9 9.9 9.82 ) , X T y = ( 19.8 19.59 ) .

Solve the ridge normal equations ( X T X + λ I ) β = X T y at λ = 1 .

Report β 1 to six decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

λ I adds λ to each diagonal entry and nothing to the off-diagonals.

Hint 2: Next step

By Cramer's rule, β 1 = ( 10.82 × 19.8 − 9.9 × 19.59 ) / 21.01 .

Direct application

On the design of this unit, ordinary least squares gives coefficients summing to 1.973684 , and ridge at λ = 10 gives β 1 = + 0.665220 and β 2 = + 0.656121 .

By how much has the penalty reduced the fitted total effect? Report the difference between the two sums, to six decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

Add the two ridge coefficients first.

Hint 2: Concept cue

Shrinkage pulls coefficients toward zero, so the penalised total should be the smaller of the two.

Direct application

On the design of this unit at penalty strength λ = 0.506030 :

  • ridge returns β = ( + 1.003705 , + 0.934853 ) ;
  • the lasso returns β = ( + 1.929397 , 0 ) .

How many predictors does the lasso solution retain, that is, how many of its coefficients are nonzero?

Enter the value. It is checked against the answer and the precision this task asks for.

1 hint available, least help first.

Hint 1: Retrieval cue

A coefficient of exactly zero removes that predictor from the model.

Direct application

A two-predictor design gives

X T X = ( 10 9.9 9.9 9.82 ) .

Compute its determinant, exactly.

(Compare your answer with the diagonal entries: that comparison, rather than the number alone, is what indicates near-collinearity.)

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

For ( a b c d ) the determinant is a d − b c .

Hint 2: Next step

Both products are close to 98 , so keep the full decimals rather than rounding before subtracting.

Error diagnosis · Explanation

An analyst fits a two-predictor least-squares model and obtains β 1 = + 2.605263 and β 2 = − 0.631579 . Plotting each predictor against the response shows both rising steadily. The design matrix gives

X T X = ( 10 9.9 9.9 9.82 ) .

The analyst concludes there is an arithmetic error somewhere and recomputes, getting the same answer.

(a) Compute the determinant of X T X and the correlation between the two predictor columns, and say what they indicate.

(b) Explain why the negative coefficient is not an arithmetic error and not evidence that x 2 acts negatively.

(c) The coefficients sum to + 1.973684 , and the response is close to 2 x 1 . Say what this tells you about which quantities the data does and does not determine.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Solving the normal equations means inverting X T X . What does its determinant do to that inverse?

Hint 2: Concept cue

If x 1 and x 2 are nearly the same vector, how many coefficient pairs fit the data about equally well?

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The determinant and the correlation.

det ( X T X ) = 10 × 9.82 − 9.9 2 = 98.2 − 98.01 = 0.19 .

The correlation between the columns is

9.9 10 × 9.82 = 9.9 98.2 = 0.999032 .

A determinant of 0.19 against diagonal entries near 10 means X T X is nearly singular, and a correlation of 0.999032 says the two predictors carry almost the same information. These are the same fact stated two ways. (b) Why it is neither an error nor an effect. It is not an arithmetic error because ( + 2.605263 , − 0.631579 ) genuinely minimises squared error on this data, recomputation confirms it, and it would confirm it however many times the analyst repeated the calculation. The procedure did what it is defined to do. It is not evidence that x 2 acts negatively because the coefficient is not describing x 2 on its own. When two predictors are nearly identical, many coefficient pairs fit almost equally well: 3 x 1 − x 2 and x 1 + x 2 are nearly the same vector, so least squares has no substantive basis for choosing between them and settles the tie on whichever matches the noise marginally better. The negative value is the second half of a cancelling pair, and it is a property of the pair rather than of the predictor. Mechanically: solving requires inverting a matrix with determinant 0.19 , and dividing by a small number magnifies small features of the data into large coefficients. The instability is real, the arithmetic is right, and the interpretation is what fails. (c) What the data determines. The sum + 1.973684 is close to the true total effect of about 2 , while neither individual coefficient is close to anything sensible. So the data determines the joint contribution of x 1 and x 2 well, and says almost nothing about how to divide that contribution between them. This matters for what the model can be used for. A prediction uses the coefficients only through the combination β 1 x 1 + β 2 x 2 , which is well determined, so predictions from this fit may be perfectly serviceable. Reading either coefficient as the effect of its predictor uses exactly the part the data cannot resolve. The instability is therefore not uniform across everything one might want to know, and whether it is a problem depends on which of the two is wanted.

A complete answer does each of these:

  • diagnoses collinearity

Direct application · Interpretation

On this unit's design, one response value is increased by 0.2 and both models are refitted:

beforeafter
least squares ( + 2.605263 , − 0.631579 ) ( + 2.773684 , − 0.842105 )
ridge, λ = 1 ( + 0.965969 , + 0.926702 ) ( + 0.948453 , + 0.905759 )

(a) Compute the Euclidean displacement of the coefficient vector in each case, and their ratio.

(b) State what quantity this comparison estimates, and why a single fit cannot reveal it.

(c) A colleague says the comparison shows ridge is more accurate. Correct that, saying what the comparison does and does not show.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

The displacement is the Euclidean distance between the before and after coefficient pairs.

Hint 2: Concept cue

Ask what was held fixed in this experiment and what was varied, that determines which property is being measured.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The displacements.

Least squares:

( 2.773684 − 2.605263 ) 2 + ( − 0.842105 + 0.631579 ) 2 = 0.168421 2 + 0.210526 2 = 0.269605 .

Ridge:

( 0.948453 − 0.965969 ) 2 + ( 0.905759 − 0.926702 ) 2 = 0.017516 2 + 0.020943 2 = 0.027302 .

Ratio:

0.269605 0.027302 = 9.875079 .

The unpenalised coefficients move almost ten times as far in response to the same change in one observation.

(b) What it estimates.

It estimates variance, how much the fitted coefficients depend on the particular sample rather than on the underlying relationship. A different draw from the same process would give somewhat different responses, and this comparison shows how much the answer would move in consequence.

A single fit cannot reveal it because a single fit produces one number per coefficient, with nothing to compare against. + 2.605263 does not announce that it would have been + 2.773684 had one observation differed slightly. Variance is a property of the procedure across datasets, so it becomes visible only when the data is varied and the procedure rerun, by perturbation, as here, or by resampling.

(c) Correcting the accuracy claim.

The comparison says nothing about accuracy. Both fits were computed from the same data and compared with each other, not with any true value.

On accuracy the evidence points the other way. The unpenalised fit minimises squared error on this data by construction, so ridge is strictly worse there. And the ridge coefficients sum to 1.892670 where the unpenalised pair sums to 1.973684 and the underlying signal totals about 2 , so shrinkage has pulled the estimated total further from the truth. That is bias, and it is the price rather than a side effect.

What the comparison shows is stability: the penalised estimate depends far less on which particular observations were collected. The summary is a trade: variance down by a factor of about ten, bias up, training fit worse, and whether it was worth making is not decidable from these figures, since every one of them comes from the same five observations. That question needs error measured on data the fit has not seen.

A complete answer does each of these:

  • quantifies stability gain

Error diagnosis · Method selection

An analyst writes: "We fitted ridge at λ ∈ { 0 , 0.01 , 0.1 , 0.5 , 1 , 2 , 5 , 10 } and selected the value giving the lowest residual sum of squares on our dataset. We then reported that model's residual sum of squares as its performance."

(a) Say which λ this procedure will select, and why the answer does not depend on the data.

(b) Name the two distinct errors in the quoted procedure.

(c) Describe a procedure that would answer both questions the analyst was trying to answer, which λ , and how well the resulting model performs.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

What is ordinary least squares defined to minimise?

Hint 2: Concept cue

Two things are being asked of one dataset here. A choice and a measurement. Can the same observations supply both?

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) It will select λ = 0 , whatever the data. Ordinary least squares is defined as the minimiser of the residual sum of squares on the fitting data. Every λ > 0 minimises a different objective, squared error plus a penalty, so its solution cannot beat the least-squares solution on squared error alone. Residual sum of squares therefore rises monotonically with λ , and the minimum is at λ = 0 by construction. The answer is fixed before the data is seen. A selection procedure whose outcome is determined in advance is not selecting anything, and it returns precisely the unpenalised estimator the penalty existed to avoid. (b) The two errors. First: selecting λ by training error. As above, this cannot work even in principle. The criterion has no interior minimum, so it cannot express a preference for any positive penalty. Second: reporting performance on the data used to fit and select. Even with a valid selection rule, quoting the error on data that participated in choosing λ gives an optimistic figure. The parameter was chosen to make that number small, so the number no longer estimates performance on new data. These are separate mistakes: the first makes the choice impossible, and the second would remain wrong even after the first was fixed. (c) A procedure that answers both. Split the data three ways, or use nested resampling: 1. Hold out a test set and set it aside untouched.
2. On the remainder, use k -fold cross-validation: for each candidate λ , fit on k − 1 folds and measure error on the held-out fold, averaging over folds. Held-out error falls and then rises as λ grows, so it has an interior minimum and can express a genuine preference.
3. Choose the λ minimising cross-validated error, or the largest λ within one standard error of that minimum if a simpler model is wanted, and say which rule was used, since they give different answers.
4. Refit at the chosen λ on all the non-test data.
5. Report performance on the test set, which took no part in fitting or selection. The structure matters: step 2 chooses, step 5 measures, and they use different data because a dataset used to choose a parameter can no longer give an unbiased estimate of the result. One caveat worth adding. Cross-validation estimates the performance of the procedure, and the selected λ can vary across splits. Reporting the chosen value as though it were a stable property of the problem overstates what the resampling established.

A complete answer does each of these:

  • selects penalty honestly

Interpretation · Comparison

Two sensors measure the same physical quantity on the same units, correlating at 0.999032 . Fitted against a response, three models give:

β 1 β 2 sum
least squares + 2.605263 − 0.631579 1.973684
ridge, λ = 0.506030 + 1.003705 + 0.934853 1.938558
lasso, λ = 0.506030 + 1.929397 0 1.929397

(a) For each model, say what it asserts about the two sensors, and whether that assertion is credible given what the sensors measure.

(b) The lasso's total is closest to the true effect of about 2 . Say whether that makes it the best choice here, and why the answer is not simply yes.

(c) The team wants to report "the effect of the measured quantity" as a single number with an interpretation. Say what you would report and what you would say about it.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Compare the three sums with each other, then compare the three values of β 1 with each other.

Hint 2: Concept cue

Which quantity is stable across the three fits? That is the one the data determines.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) What each model asserts. Least squares asserts that sensor 1 has a strong positive effect of + 2.605263 and sensor 2 a negative effect of − 0.631579 . Given that both measure the same physical quantity, this is not credible: two readings of one thing cannot act in opposite directions. The pair is a cancelling artefact of near-singularity, not a description of the sensors. Ridge asserts that each sensor carries roughly half the effect, + 1.003705 and + 0.934853 . This is credible and is arguably the most honest of the three: the underlying quantity has one effect, both sensors measure it, and neither reading is privileged. The split reflects a genuine symmetry in what is known. Lasso asserts that sensor 1 carries the entire effect of + 1.929397 and sensor 2 contributes nothing. As a prediction rule this is fine; as a statement about the sensors it is misleading, since sensor 2 measures the quantity just as well and its exclusion reflects the optimisation breaking a near-tie rather than any difference between the instruments. Had the sample differed slightly, the lasso might well have kept sensor 2 and dropped sensor 1. (b) Why closest-to-the-truth does not settle it. The lasso total 1.929397 is nearer 2 than ridge's 1.938558 , but the gap is 0.009161 , far smaller than the uncertainty in any of these estimates from five observations. Treating that as a ranking reads more into the figures than they carry. More importantly, it is a fact about this sample. The comparison was computed on the same data that produced the fits, so it measures which estimator happened to land closer here, not which is generally better. A different draw could reverse it. And the criterion is the wrong one for the stated purpose. If the deliverable is a prediction rule, both are fine and the choice can rest on held-out error. If it is a statement about the sensors, the question is which description is true, and "both sensors measure the quantity, each accounting for part of the fitted effect" is true, while "sensor 2 has no effect" is false. (c) What to report. Report the sum, about 1.93 to 1.97 depending on the estimator, as the effect of the underlying quantity, with the explicit statement that it is the combined contribution of the two sensors and that the data does not determine how to divide it between them. The justification is in the fits themselves: the individual coefficients swing wildly across the three models, from − 0.631579 to + 1.929397 for the same sensor, while the sums agree to within about 0.05 . What is stable is what the data determines, and what is unstable is what it does not. Also report that the two sensors correlate at 0.999032 , so a reader can see why the split is unavailable; state which penalty was used and at what strength; and if the penalised coefficients are shown, note that they are biased toward zero by construction and so understate the effect slightly. Presenting any single coefficient as "the effect" would report a number that the next sample would contradict.

A complete answer does each of these:

  • contrasts penalty geometry
  • diagnoses collinearity

Transfer · Evaluation · Explanation

A lender builds a default-risk model on 400 historical applicants with 10,000 candidate predictors, many of them near-duplicates, "months at address" and "months since last move", several overlapping credit-bureau aggregates.

Their report states: "Ordinary least squares fitted the training data perfectly, with zero residual error. We then applied a lasso penalty, selecting λ as the value minimising residual error on the same data. The lasso identified 37 predictors as the true drivers of default. We recommend collecting only these 37 going forward, and we interpret their coefficients as the effect of each factor on default risk."

Write a review covering:

(a) What a zero training error on 400 observations with 10,000 predictors indicates, and what least squares can and cannot deliver in this setting.

(b) What their λ selection procedure will return, and why.

(c) Whether the 37 selected predictors can be described as the true drivers, referring to the near-duplicates.

(d) Whether the coefficients can be interpreted as effects on default risk.

(e) What you would do instead, including how you would choose between the two penalties and what you would report.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

With 10,000 predictors and 400 observations, what is the rank of X T X ?

Hint 2: Concept cue

Separate three claims in their report: the fit, the selection, and the interpretation. Each fails for a different reason.

Hint 3: Strategy cue

For (e), decide what the deliverable is, a prediction rule or a statement about factors, before choosing a penalty.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) Zero training error is the warning, not the result. With p = 10,000 predictors and n = 400 observations, X T X is a 10,000 × 10,000 matrix of rank at most 400 , singular, not merely near-singular. Least squares has no unique solution: infinitely many coefficient vectors reproduce the training responses exactly, and the fitting software returned one of them arbitrarily. A perfect fit here carries no information about the data. It is guaranteed by the shape of the problem whenever p ≥ n , and would have appeared just as readily had the predictors been random numbers. Least squares can deliver a fit; it cannot deliver an estimate. This is the extreme of the situation the unit's small design illustrates. There, two predictors correlating at 0.999032 gave det ( X T X ) = 0.19 and coefficients of + 2.605263 and − 0.631579 . A cancelling pair. Here the determinant is exactly zero and the arbitrariness is total. (b) Their selection procedure will return λ = 0 . Residual error on the fitting data is minimised by ordinary least squares by definition, and rises monotonically as λ grows, because every λ > 0 minimises a different objective. So the rule as stated selects no penalty at all. The answer is fixed before any data is examined. That this did not visibly happen suggests something else occurred: perhaps the search grid excluded zero, or the software's default cross-validation was used and the report describes it incorrectly. Either way the procedure as written cannot have produced a meaningful 37 -predictor model, and the discrepancy needs resolving before anything else in the report can be trusted. (c) The 37 are not the true drivers. The lasso returns a set of predictors sufficient to predict well, not the set that matters. Where predictors are near-duplicates, it keeps roughly one from each redundant group and zeroes the rest, and which one survives is close to arbitrary, decided by whichever matches the sample noise marginally better. On the unit's two-predictor design this is visible in miniature: the lasso zeroes the second predictor from λ ≈ 0.006030 onward, though both sensors measure the same quantity equally well. Nothing distinguishes them except the sample. So "months at address" surviving while "months since last move" is dropped is not evidence that the first drives default and the second does not. A different 400 applicants could easily reverse it. The right check is refitting on resamples and recording how often each predictor is selected; a variable chosen in 30 % of resamples is not a finding. The recommendation to collect only these 37 inherits the problem. If a dropped near-duplicate is the one that remains measurable in future, discarding it on this evidence is a costly mistake. (d) The coefficients cannot be interpreted as effects. Three reasons, each sufficient. First, lasso coefficients are deliberately biased toward zero. On the unit's design, ridge at λ = 1 shrinks the coefficient sum from 1.973684 to 1.892670 when the underlying signal is about 2 . Penalised coefficients systematically understate effects, by design. Second, a coefficient in a model with near-duplicate predictors describes the predictor's contribution given the others in the model, and with duplicates that contribution is not identifiable. The unit's design shows the point sharply: the same sensor takes coefficients from − 0.631579 to + 1.929397 across three fits, while the sum stays near 1.93 – 1.97 . The data determines combinations, not individual shares. Third, these are observational credit records. Even a perfectly estimated coefficient would be an association, and calling it "the effect on default risk" asserts a causal claim no regression on this data supports. A claim with consequences, given that the model will be used to decline applicants. (e) What to do instead. Structure the evidence properly. Hold out a test set and touch it once. On the remainder, use k -fold cross-validation to choose λ : held-out error falls and then rises, so it has an interior minimum and can genuinely express a preference, which training error cannot. Refit at the chosen λ on all non-test data and report performance on the test set alone. If the selected λ varies substantially across folds, say so. Choose the penalty by the deliverable. If the goal is a cheaper data-collection pipeline, the lasso's sparsity is the right tool, accompanied by a selection-stability analysis over resamples, reporting selection frequencies rather than a list of 37 names. If the goal is the best risk prediction from data already collected, ridge is usually preferable with grouped near-duplicates, since it keeps correlated predictors and splits the signal rather than picking arbitrarily among them. The elastic net, combining both penalties, is the standard compromise and selects groups together instead of one member. Whichever is chosen, the decision should be stated with its reason. Quantify the stability gained. Refit on perturbed or resampled data and report how far the coefficients move, with and without the penalty. On the unit's design the displacement falls from 0.269605 to 0.027302 , a factor of 9.875079 . A magnitude a reader can weigh, unlike an assertion that the penalty stabilises the fit. Report the trade, both halves. The penalised model fits the training data worse, necessarily, and its coefficients are biased toward zero. What it gains is that the answer depends far less on which 400 applicants happened to be in the sample. State both, and state that no figure computed from the fitting data can settle whether the exchange was worthwhile. Separate the claims. "These predictors give the best held-out prediction we could obtain" is supportable. "These are the drivers of default, and their coefficients are the effects" is not, and in a lending context the difference is not academic: the second invites decisions about individual applicants that the evidence does not license.

A complete answer does each of these:

  • diagnoses collinearity
  • computes penalized solution
  • contrasts penalty geometry
  • quantifies stability gain
  • selects penalty honestly
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.