Practice: Estimating Out-of-Sample Error
Question
Direct application
A 4-fold cross-validation of a model produces held-out mean squared errors of
Enter the value. It is checked against the answer and the precision this task asks for.
2 hints available, least help first.
Hint 1: Retrieval cue
Each fold was held out exactly once, so each contributes equally.
Hint 2: Next step
Add the four losses and divide by 4.
Classification · Prediction
A straight line is replaced by a degree-8 polynomial, fitted to the same 15 observations. Which description of the change in expected test error is correct?
Method selection · Interpretation
A model forecasts next month's demand from the preceding months. Three years of monthly observations are available. Which evaluation scheme estimates the quantity the model will actually be asked for, and why?
Error diagnosis
An analyst has 60 observations and 5000 candidate predictors. They keep the 20 predictors most correlated with the response, then run 10-fold cross-validation using only those 20, and report a cross-validation error far below the variance of the response. What is wrong?
Interpretation
Two models are compared by 5-fold cross-validation on the same folds. Model A has mean held-out error
Transfer · Evaluation
A competition scores submissions on a public leaderboard computed from a fixed held-out set. Over three months a team submits 200 times, each time adjusting the model according to its leaderboard score, and finishes first with a public score of
Construction · Evaluation · Explanation
A subscription business wants a model predicting which customers will cancel in the coming month. You are given 24 months of records; each row is one customer in one month, customers appear in many months, and a customer who cancels stops appearing. A colleague has fitted a gradient-boosted model, tuned its depth and learning rate by 10-fold cross-validation over all rows, and reports the best cross-validation error obtained during that search as the expected error in production.
Write an evaluation plan and an assessment of the colleague's figure. Address all of the following.
- The resampling scheme. Say how you would split these records and carry out the estimate, and state what each fit is trained on.
- The structure of the data. Identify every feature of these records that forbids random folds, and say what a random fold would let the model know that it will not know in production.
- The colleague's figure. State whether it is an honest estimate of production error, name the mechanism if it is not, and describe the procedure that would produce an honest one.
- The error components. The colleague proposes a deeper model because training error is still falling. Say what this would do to bias, to variance, and to the irreducible term, and what evidence would tell you whether it helps.
- Reporting. State what you would report alongside the error estimate, and how you would decide whether a rival model is genuinely better.
Write your answer, then compare it with the worked solution.
3 hints available, least help first.
Hint 1: Retrieval cue
Ask what will be true at the moment a prediction is made in production, and arrange the held-out data to match it.
Hint 2: Concept cue
Two separate properties of these rows each break random folds. One is about when the observation occurred; the other is about who it belongs to.
Hint 3: Strategy cue
For the colleague's figure, ask what the search actually optimised, and what that implies about the value it returned.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
1. The resampling scheme. Split by time and by customer together. Choose cut points in the 24 months, for example after months 12, 15, 18 and 21, and for each, fit on every row from a customer whose records all fall before the cut, then evaluate on rows after it. Each fit is trained only on months earlier than the ones it is scored on, and no customer contributes rows to both sides of a split. Average the held-out losses across the cut points. 2. The structure of the data. Two features forbid random folds, and both must be addressed. Time. The rows are ordered, and the model will be used to predict forward. A random fold trains on months after the row being scored, so the model uses information that will not exist when the prediction is actually made. The estimate that results describes filling a gap in a known period, not forecasting. Repeated customers. One customer appears in many rows. Random assignment puts the same customer in both the training and held-out sets, so the model can learn that individual's pattern and recognise them rather than learning behaviour that generalises to a customer it has never seen. a customer who cancels stops appearing, so the number of rows a customer contributes is itself informative about the outcome. Any split must not let that leak either. 3. The colleague's figure. It is not an honest estimate of production error, for two independent reasons. The folds were random, so it suffers both leaks in point 2. Separately, the depth and learning rate were chosen by minimising that same cross-validation error. The reported value is therefore the minimum over many noisy estimates, which sits below what the chosen configuration would score on fresh data. The bias grows with the number of configurations tried. An honest figure needs a nested scheme: within each outer training period, run the full tuning search; fit with the winning values on that whole training period; and score once on the outer held-out months. Average those outer scores. That estimates the entire procedure, tuning included. 4. The error components. Greater depth lowers squared bias, raises variance, and leaves the irreducible term unchanged. The last is a property of how predictable cancellation is from the recorded variables, and no model alters it. Training error falling is not evidence for the change: it falls with depth whether or not test error does. The evidence that decides it is held-out error from the scheme in point 1. If that rises while training error falls, the additional depth is fitting noise. 5. Reporting. Report the mean held-out error together with the spread across the outer folds, and state the splitting scheme and that tuning was nested inside it. Treat a rival model as better only when the difference in means exceeds the fold-to-fold variation; when it does not, the comparison has not separated them, and saying so is the accurate report.
A complete answer does each of these:
- computes resampled error
- attributes error components
- selects split for structure
- identifies selection bias
- reports uncertainty
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.