Estimating Out-of-Sample Error
What you will be able to do
The learner can estimate a predictive model's performance on unseen data by an appropriate resampling scheme, attribute the remaining error to bias, variance or irreducible noise, and state what a reported error figure does and does not license.
Orientation
The number a model reports about itself
A fitted model can always be asked how well it does on the data it was fitted to. That number is easy to obtain and answers a question nobody asked: how well the model reproduces observations whose answers were already available.
The question that matters is what it will predict for an observation it has never seen. Those two numbers differ, and the difference is not small. A polynomial of high enough degree passes exactly through every training point and reports zero error, while predicting new observations worse than a straight line would.
This unit covers how to estimate the second number: the resampling procedure that withholds data and asks, the decomposition that says where the remaining error comes from, and the two situations where the standard procedure gives an answer that is still too optimistic, data ordered in time, and data reused to choose the model.
Definition
Reading the three terms, and what each depends on
Why the expectation in the decomposition runs over training sets. Bias and variance are not properties of one fitted model. Both are statements about what happens when the same procedure is applied to a fresh sample: bias asks where those fits sit on average relative to the truth, variance asks how far they scatter around that average. Neither can be computed from the single sample in hand, which is why they are reasoned about rather than measured.
| Term | Depends on | Changes when |
|---|---|---|
| squared bias | the model's capacity to represent | the model family changes |
| variance | capacity relative to the sample size | either the model or |
| the data-generating process | never, for any model |
Flexibility moves the first two in opposite directions. A more flexible model follows the true
A floor that no model clears. Because
What
Loss functions other than squared error. The three-term decomposition is exact for squared-error loss and does not carry over unchanged to classification error or log-loss. The qualitative trade between bias and variance survives; the clean additive separation does not.
Intuition
Training error, optimism, and sampling variability
Fitting a model means choosing parameters that make it agree with a particular set of observations. Those observations carry noise, and the fitting procedure cannot separate noise from signal. It minimises disagreement with whatever it was given. Some of the agreement it achieves is therefore agreement with accidents of that sample, which will not recur. Asking the model about those same observations counts that agreement as success.
The more flexible the model, the more of the sample's accidents it can accommodate, so training error keeps falling while the agreement it achieves becomes less transferable.
Bias and variance are answers to two different questions about a procedure. Imagine refitting the same model on many independent samples from the same population.
- Where do the fits sit on average, compared with the truth? A systematic offset is bias. A straight line fitted to genuinely curved data misses the same way in every sample, however many observations each contains.
- How much do the fits differ from one another? That spread is variance. A model with many parameters relative to the data produces a visibly different fit each time, because each sample's noise pulls it somewhere else.
A rigid model has high bias and low variance; a flexible one has the reverse. Neither quantity is observable from a single sample, which is why they are reasoned about rather than measured directly.
The irreducible term explains why error has a floor. If two observations share the same
Cross-validation withholds repeatedly rather than once. A single split gives one estimate, and which observations landed in the held-out portion affects it; an unlucky split can make a good model look poor. Rotating the held-out fold means every observation is tested on exactly once, and the spread across folds shows how much that choice mattered.
Example
Error estimates that answer different questions
Each case below reports an error figure. What separates them is not the arithmetic but which quantity the figure estimates.
A figure that estimates nothing about future performance. A shop fits a model to last year's transactions, predicts those same transactions, and reports 94% accuracy. The number is correct and describes how well the model reproduces answers it was given. It places no bound on next month's accuracy, and would be near 100% for a model that had memorised the file.
A figure that estimates performance on a fresh draw from the same population. The same shop holds out a random 20% of last year's transactions, fits on the rest, and reports 87% on the held-out portion. This estimates accuracy on another transaction from last year, useful if the population is stable, and silent about drift.
A figure that estimates next month's performance. The shop fits on January–October and evaluates on November–December, reporting 81%. The drop from 87% is the part of the earlier figure that came from knowing the future. This is the number to quote when the model will be used to predict forward.
A figure made optimistic by selection. The shop tries 40 feature sets, cross-validating each, and reports the best score of 89%. That value is the minimum of 40 noisy estimates and sits below what the chosen feature set would score on fresh data. An outer split not used in the search would give the honest figure.
A figure that reveals leakage. A model predicting whether a customer will return within 30 days reports 99.2% accuracy. Inspection shows one predictor is the date of the customer's next visit, recorded after the outcome it predicts. An error far below what the problem plausibly allows is evidence about the data pipeline rather than about the model.
---
The first and last are not estimates of test error at all. The middle three are, and they estimate different populations: another transaction from the same period, a transaction from a later period, and the performance of a whole selection procedure. Quoting one where another is meant is the most common way an honest calculation supports a claim it does not license.
Procedure
Running a cross-validation, and reporting what it gives
To estimate test error by
- Check what the observations are before splitting. Ordered in time, repeated on the same subject, or clustered within a group each forbid random folds. Step 5 covers those cases.
- Choose
. Five or ten are the usual choices. Larger fits on more data, so each fit is closer to the one the full sample would give, at proportionally more computation; is leave-one-out. - Partition the observations into
folds of roughly equal size, assigned at random. - For each fold
: fit the model on the other folds, predict the observations in fold, and record the loss . Every step of fitting, including any scaling, imputation or variable selection, must happen inside this loop, using only the training folds. - Report the mean
together with the spread of the, usually their standard deviation divided by .
When the observations are ordered in time. Use a temporal holdout: fit on observations up to some time
When observations are grouped. Assign whole groups to folds rather than individual observations, so no subject or cluster appears on both sides of a split.
To choose a tuning parameter and still report an honest error. Cross-validation used to select cannot also estimate: the winning value was chosen because it scored well on those folds. Use a nested scheme.
- Split into outer folds.
- Within each outer training set, run a complete inner cross-validation to choose the tuning parameter.
- Fit with the chosen value on the whole outer training set, and evaluate once on the outer held-out fold.
- Average the outer fold losses. That average is the estimate of the procedure's test error, tuning included.
Checks. Compare the cross-validation figure with training error: the first should be larger, and a training error near zero alongside a much larger held-out error indicates a model flexible enough to interpolate. If the fold losses vary by more than the difference between two models under comparison, the comparison does not distinguish them.
Worked example
Equal training error, different held-out error
Twelve observations,
The underlying relationship is close to
Step 1: fit each on all twelve and record training error.
| Model | Training MSE |
|---|---|
| degree 1 | |
| degree 5 |
On this evidence the two are indistinguishable, and the polynomial is marginally ahead. Training error cannot separate them, and it never favours the simpler model: adding parameters can only reduce the quantity being minimised.
Step 2: partition into 4 folds. Observations are assigned by position,
Step 3: for each fold, fit on the other nine and score the held-out three.
| Fold | degree 1 | degree 5 |
|---|---|---|
| 1 | ||
| 2 | ||
| 3 | ||
| 4 |
Step 4: average, and report the spread.
| Model | fold s.d. | s.e. of the mean | |
|---|---|---|---|
| degree 1 | |||
| degree 5 |
Reading the result.
The line's held-out error is
The polynomial's held-out error is
The difference between the two,
The fourth fold is the informative one. The polynomial scores
Reporting
Contrast
Splits that differ by one decision
Random folds against a temporal holdout, on daily sales.
| random 5-fold | temporal holdout | |
|---|---|---|
| held-out day | any day | the last 20% of days |
| fitted on | days before and after it | days before it only |
| answers | how well the model fills a gap | how well it forecasts |
| typical error | lower | higher |
Both are correct procedures; they estimate different quantities. Random folds let a model learn from Tuesday and Thursday to predict Wednesday, which is unavailable when Wednesday is genuinely next. The lower error is not a better model but a different, easier question.
Cross-validation for selection against cross-validation for estimation.
| selecting | estimating | |
|---|---|---|
| question | which candidate is best? | how will the chosen one perform? |
| quantity used | the ranking of the scores | the value of one score |
| valid on the same folds | yes | no |
The ranking is comparatively robust: shared fold noise affects all candidates together. The winner's value is not, because selecting the minimum selects partly for favourable noise.
Scaling inside the fold loop against scaling before it.
Centre and scale using the whole dataset, and each training fit has used the held-out fold's mean. The leak is small for a mean over many observations and can be large for a variable-selection step on few. The distinction is not the size of the effect but whether the procedure can be described honestly: a figure obtained with outside information does not estimate performance on data the model has not seen.
| 5-fold | leave-one-out | |
|---|---|---|
| fits required | 5 | |
| training set size | ||
| bias of the estimate | slightly pessimistic | nearly unbiased |
| variance of the estimate | lower | higher |
Leave-one-out fits on almost the full sample, so it estimates the error of the model actually being fitted. Its
In each, one decision changes what quantity is being estimated rather than how precisely. That is why the question to ask about a validation scheme is which quantity it estimates, before asking whether the number it produced is good.
Warning
Reusing cross-validation for selection and reporting
Cross-validation is usually run many times: once per candidate tuning value, or per feature set, or per model family. The best score among those runs is then reported as the estimate of test error. That figure is optimistic, and the reason is the same one that makes training error optimistic, one level up.
Why selecting on a score biases it. Each candidate's cross-validation score is an estimate carrying its own random error from the particular fold assignment. Taking the minimum over many candidates selects partly for genuinely lower error and partly for a favourable draw. The winner's score is therefore below what that candidate would achieve on fresh data, and the gap grows with the number of candidates tried.
The effect is not small when the candidate set is large. Comparing forty tuning values on twelve folds means taking a minimum over forty noisy estimates, and the minimum of forty draws sits well below their common mean.
What this looks like in practice. A tuning curve is computed, its lowest point is read off, and that value is quoted as the expected error. Nothing in the procedure signals a problem: the folds were held out, no observation was evaluated by a model fitted to it, and the arithmetic is correct. The reuse is at the level of the selection, which the fold structure does not protect.
The fix is an outer split. Reserve data the selection never touches, or nest the cross-validations as the procedure block describes. The outer estimate then measures the whole procedure, tuning included, rather than the tuned model alone. It will be larger than the inner best score, and that is the honest number.
A related failure with the same shape. Any step fitted on all the data before splitting leaks in the same way: scaling by the overall mean and standard deviation, imputing with a column median, selecting variables by their correlation with the response. Each uses held-out observations to shape the fit later evaluated on them. Variable selection on the full dataset is the most damaging, because it can produce a respectable cross-validation score from predictors that are pure noise.
What survives the warning. Cross-validation remains the right tool for choosing among candidates; comparing scores computed the same way is what it is for. The caution concerns the second use, reporting the winner's score as an estimate of future performance.
Application
Where the estimate is the deliverable
Clinical prediction. A model estimating a patient's risk is evaluated on patients from hospitals that contributed no training data, because performance falls when case mix or measurement practice differs. A model validated only by cross-validation within one hospital's records has been shown to work where it was built, which is the weakest claim of interest.
Credit scoring. Applicants are scored over time and the population shifts, so evaluation uses a temporal holdout: fit on applications up to a date, evaluate on applications after it. A random split would report the error of filling gaps in a known period rather than the error of scoring next quarter's applicants.
Recommendation. Users and items both recur, so a random split over interactions can place a user's later interaction in the training set and an earlier one in the held-out fold. Evaluation is therefore usually split by user and by time together, holding out each user's most recent interactions.
Competition leaderboards. A public leaderboard scored on a fixed held-out set is a shared validation set, and participants who submit repeatedly are selecting on it. Final standings are computed on a second, private set precisely because the public score stops estimating test error once it has been optimised against. The selection bias of the warning block, at the scale of a whole community.
Scientific reporting. Where a model supports a claim rather than a product, what is reported is the estimate together with how it was obtained: the splitting scheme, whether tuning was nested, and the spread across folds. Two papers reporting the same accuracy on the same data are not making the same claim if one tuned on its evaluation folds.
---
The common shape. Each case asks the same question the unit began with, how will this perform on data it has not seen, and each answers it by making the held-out data resemble the unseen data in the way that matters: a different hospital, a later date, a different user, a second private set. The choice of scheme follows from what "unseen" will actually mean when the model is used, which is a question about the application rather than about statistics.