Estimating Out-of-Sample Error

Why the error a model reports on the data it was fitted to is not the error it will make on new data, how resampling estimates the second, what bias, variance and irreducible noise each contribute, and which splitting schemes remain valid when the observations are ordered in time or reused for tuning.

Definition

The training error of a fitted model is its average loss on the observations used to fit it. The test error is its expected loss on an independent observation from the same distribution. These differ, and the gap widens with model flexibility, because fitting chooses parameters that suit the particular sample including its noise.

k -fold cross-validation estimates test error without a separate test set. Partition the observations into k roughly equal folds. For each fold in turn, fit on the other k − 1 and record the loss on the held-out fold. The cross-validation estimate is the average of the k held-out losses,

CV k = 1 k ∑ i = 1 k L i ,

where L i is the loss on fold i . Every observation is held out exactly once and used for fitting k − 1 times.

For squared-error loss the expected test error at a point decomposes as

E [ ( y − f ^ ( x ) ) 2 ] = ( E [ f ^ ( x ) ] − f ( x ) ) 2 ⏟ bias 2 + Var ⁡ ( f ^ ( x ) ) ⏟ variance + σ 2 ⏟ irreducible ,

the expectations taken over training sets. The third term is a property of the data, not of the model, and no method reduces it.

Assumptions and scope

  • Cross-validation estimates the error of the procedure applied to a training set of size about n ( k − 1 ) / k , not of one particular fitted model. For small n the difference matters, since each fold fits on visibly less data than the full sample.

  • The folds must be drawn so that held-out observations are independent of the ones used for fitting. Random folds break this whenever observations are ordered in time, repeated on the same subject, or clustered.

  • Any choice made by looking at held-out error is itself fitting. A tuning parameter selected by cross-validation makes that same cross-validation figure optimistic as an estimate of test error.

  • The decomposition into bias, variance and irreducible noise is exact for squared-error loss. It does not carry over unchanged to classification error or other losses.

  • Under squared-error loss and the model y = f ( x ) + ε with Var ⁡ ( ε ∣ x ) = σ 2 , the EXPECTED test error at a point is at least σ 2 , since bias 2 and variance are non-negative. A cross-validation or held-out figure is a finite-sample ESTIMATE of that expectation and carries its own sampling variation, so a single estimate below σ 2 is not by itself evidence of a defect. Persistently optimistic estimates are worth investigating: leakage between fitting and evaluation data is one explanation, a mis-specified σ 2 and an evaluation set too small or unrepresentative are others.

Worked material

Example

Error estimates that answer different questions

Each case below reports an error figure. What separates them is not the arithmetic but which quantity the figure estimates.

A figure that estimates nothing about future performance. A shop fits a model to last year's transactions, predicts those same transactions, and reports 94% accuracy. The number is correct and describes how well the model reproduces answers it was given. It places no bound on next month's accuracy, and would be near 100% for a model that had memorised the file.

A figure that estimates performance on a fresh draw from the same population. The same shop holds out a random 20% of last year's transactions, fits on the rest, and reports 87% on the held-out portion. This estimates accuracy on another transaction from last year, useful if the population is stable, and silent about drift.

A figure that estimates next month's performance. The shop fits on January–October and evaluates on November–December, reporting 81%. The drop from 87% is the part of the earlier figure that came from knowing the future. This is the number to quote when the model will be used to predict forward.

A figure made optimistic by selection. The shop tries 40 feature sets, cross-validating each, and reports the best score of 89%. That value is the minimum of 40 noisy estimates and sits below what the chosen feature set would score on fresh data. An outer split not used in the search would give the honest figure.

A figure that reveals leakage. A model predicting whether a customer will return within 30 days reports 99.2% accuracy. Inspection shows one predictor is the date of the customer's next visit, recorded after the outcome it predicts. An error far below what the problem plausibly allows is evidence about the data pipeline rather than about the model.

---

The first and last are not estimates of test error at all. The middle three are, and they estimate different populations: another transaction from the same period, a transaction from a later period, and the performance of a whole selection procedure. Quoting one where another is meant is the most common way an honest calculation supports a claim it does not license.

Contrast

Splits that differ by one decision

Random folds against a temporal holdout, on daily sales.

random 5-foldtemporal holdout
held-out dayany daythe last 20% of days
fitted ondays before and after itdays before it only
answershow well the model fills a gaphow well it forecasts
typical errorlowerhigher

Both are correct procedures; they estimate different quantities. Random folds let a model learn from Tuesday and Thursday to predict Wednesday, which is unavailable when Wednesday is genuinely next. The lower error is not a better model but a different, easier question.

Cross-validation for selection against cross-validation for estimation.

selectingestimating
questionwhich candidate is best?how will the chosen one perform?
quantity usedthe ranking of the scoresthe value of one score
valid on the same foldsyesno

The ranking is comparatively robust: shared fold noise affects all candidates together. The winner's value is not, because selecting the minimum selects partly for favourable noise.

Scaling inside the fold loop against scaling before it.

Centre and scale using the whole dataset, and each training fit has used the held-out fold's mean. The leak is small for a mean over many observations and can be large for a variable-selection step on few. The distinction is not the size of the effect but whether the procedure can be described honestly: a figure obtained with outside information does not estimate performance on data the model has not seen.

k = 5 against k = n .

5-foldleave-one-out
fits required5 n
training set size 0.8 n n − 1
bias of the estimateslightly pessimisticnearly unbiased
variance of the estimatelowerhigher

Leave-one-out fits on almost the full sample, so it estimates the error of the model actually being fitted. Its n training sets overlap in all but one observation, making the held-out losses strongly dependent and their average more variable. Five or ten folds trade a little bias for a steadier estimate.

In each, one decision changes what quantity is being estimated rather than how precisely. That is why the question to ask about a validation scheme is which quantity it estimates, before asking whether the number it produced is good.

Common errors

Common misconception

That a model's error on the data it was fitted to estimates how it will perform on new data, so a fit with low training error is a good predictor. Fitting chooses parameters that suit the particular sample including its noise, so training error falls as flexibility rises even where test error is climbing. A model flexible enough to interpolate the training points reports zero training error and may predict new observations worse than a straight line.

Common misconception

That randomly assigning observations to folds is always the correct way to cross-validate, whatever the data represent. Random folds assume held-out observations are independent of those used for fitting. When observations are ordered in time, random folds place later observations in the training set and earlier ones in the held-out fold, so the model is fitted with information that would not have been available when the held-out observation occurred. The resulting error estimate is optimistic, and the same failure arises for repeated measurements on one subject or observations clustered within a group.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.