Estimating Out-of-Sample Error
Why the error a model reports on the data it was fitted to is not the error it will make on new data, how resampling estimates the second, what bias, variance and irreducible noise each contribute, and which splitting schemes remain valid when the observations are ordered in time or reused for tuning.
Definition
The training error of a fitted model is its average loss on the observations used to fit it. The test error is its expected loss on an independent observation from the same distribution. These differ, and the gap widens with model flexibility, because fitting chooses parameters that suit the particular sample including its noise.
where
For squared-error loss the expected test error at a point decomposes as
the expectations taken over training sets. The third term is a property of the data, not of the model, and no method reduces it.
Assumptions and scope
Cross-validation estimates the error of the procedure applied to a training set of size about
, not of one particular fitted model. For small the difference matters, since each fold fits on visibly less data than the full sample.The folds must be drawn so that held-out observations are independent of the ones used for fitting. Random folds break this whenever observations are ordered in time, repeated on the same subject, or clustered.
Any choice made by looking at held-out error is itself fitting. A tuning parameter selected by cross-validation makes that same cross-validation figure optimistic as an estimate of test error.
The decomposition into bias, variance and irreducible noise is exact for squared-error loss. It does not carry over unchanged to classification error or other losses.
Under squared-error loss and the model
with , the EXPECTED test error at a point is at least , since bias and variance are non-negative. A cross-validation or held-out figure is a finite-sample ESTIMATE of that expectation and carries its own sampling variation, so a single estimate below is not by itself evidence of a defect. Persistently optimistic estimates are worth investigating: leakage between fitting and evaluation data is one explanation, a mis-specified and an evaluation set too small or unrepresentative are others.
Worked material
Example
Error estimates that answer different questions
Each case below reports an error figure. What separates them is not the arithmetic but which quantity the figure estimates.
A figure that estimates nothing about future performance. A shop fits a model to last year's transactions, predicts those same transactions, and reports 94% accuracy. The number is correct and describes how well the model reproduces answers it was given. It places no bound on next month's accuracy, and would be near 100% for a model that had memorised the file.
A figure that estimates performance on a fresh draw from the same population. The same shop holds out a random 20% of last year's transactions, fits on the rest, and reports 87% on the held-out portion. This estimates accuracy on another transaction from last year, useful if the population is stable, and silent about drift.
A figure that estimates next month's performance. The shop fits on January–October and evaluates on November–December, reporting 81%. The drop from 87% is the part of the earlier figure that came from knowing the future. This is the number to quote when the model will be used to predict forward.
A figure made optimistic by selection. The shop tries 40 feature sets, cross-validating each, and reports the best score of 89%. That value is the minimum of 40 noisy estimates and sits below what the chosen feature set would score on fresh data. An outer split not used in the search would give the honest figure.
A figure that reveals leakage. A model predicting whether a customer will return within 30 days reports 99.2% accuracy. Inspection shows one predictor is the date of the customer's next visit, recorded after the outcome it predicts. An error far below what the problem plausibly allows is evidence about the data pipeline rather than about the model.
---
The first and last are not estimates of test error at all. The middle three are, and they estimate different populations: another transaction from the same period, a transaction from a later period, and the performance of a whole selection procedure. Quoting one where another is meant is the most common way an honest calculation supports a claim it does not license.
Contrast
Splits that differ by one decision
Random folds against a temporal holdout, on daily sales.
| random 5-fold | temporal holdout | |
|---|---|---|
| held-out day | any day | the last 20% of days |
| fitted on | days before and after it | days before it only |
| answers | how well the model fills a gap | how well it forecasts |
| typical error | lower | higher |
Both are correct procedures; they estimate different quantities. Random folds let a model learn from Tuesday and Thursday to predict Wednesday, which is unavailable when Wednesday is genuinely next. The lower error is not a better model but a different, easier question.
Cross-validation for selection against cross-validation for estimation.
| selecting | estimating | |
|---|---|---|
| question | which candidate is best? | how will the chosen one perform? |
| quantity used | the ranking of the scores | the value of one score |
| valid on the same folds | yes | no |
The ranking is comparatively robust: shared fold noise affects all candidates together. The winner's value is not, because selecting the minimum selects partly for favourable noise.
Scaling inside the fold loop against scaling before it.
Centre and scale using the whole dataset, and each training fit has used the held-out fold's mean. The leak is small for a mean over many observations and can be large for a variable-selection step on few. The distinction is not the size of the effect but whether the procedure can be described honestly: a figure obtained with outside information does not estimate performance on data the model has not seen.
| 5-fold | leave-one-out | |
|---|---|---|
| fits required | 5 | |
| training set size | ||
| bias of the estimate | slightly pessimistic | nearly unbiased |
| variance of the estimate | lower | higher |
Leave-one-out fits on almost the full sample, so it estimates the error of the model actually being fitted. Its
In each, one decision changes what quantity is being estimated rather than how precisely. That is why the question to ask about a validation scheme is which quantity it estimates, before asking whether the number it produced is good.
Common errors
Common misconception
That a model's error on the data it was fitted to estimates how it will perform on new data, so a fit with low training error is a good predictor. Fitting chooses parameters that suit the particular sample including its noise, so training error falls as flexibility rises even where test error is climbing. A model flexible enough to interpolate the training points reports zero training error and may predict new observations worse than a straight line.
Common misconception
That randomly assigning observations to folds is always the correct way to cross-validate, whatever the data represent. Random folds assume held-out observations are independent of those used for fitting. When observations are ordered in time, random folds place later observations in the training set and earlier ones in the held-out fold, so the model is fitted with information that would not have been available when the held-out observation occurred. The resulting error estimate is optimistic, and the same failure arises for repeated measurements on one subject or observations clustered within a group.
Related units
Requires
- Random Variables, Expectation and Variance
- Covariance, Independence and the Variance of a Sum
- Linear Regression for Experimental Research