Estimating Out-of-Sample Error

What you will be able to do

The learner can estimate a predictive model's performance on unseen data by an appropriate resampling scheme, attribute the remaining error to bias, variance or irreducible noise, and state what a reported error figure does and does not license.

Orientation

The number a model reports about itself

A fitted model can always be asked how well it does on the data it was fitted to. That number is easy to obtain and answers a question nobody asked: how well the model reproduces observations whose answers were already available.

The question that matters is what it will predict for an observation it has never seen. Those two numbers differ, and the difference is not small. A polynomial of high enough degree passes exactly through every training point and reports zero error, while predicting new observations worse than a straight line would.

This unit covers how to estimate the second number: the resampling procedure that withholds data and asks, the decomposition that says where the remaining error comes from, and the two situations where the standard procedure gives an answer that is still too optimistic, data ordered in time, and data reused to choose the model.

Definition

Reading the three terms, and what each depends on

Why the expectation in the decomposition runs over training sets. Bias and variance are not properties of one fitted model. Both are statements about what happens when the same procedure is applied to a fresh sample: bias asks where those fits sit on average relative to the truth, variance asks how far they scatter around that average. Neither can be computed from the single sample in hand, which is why they are reasoned about rather than measured.

TermDepends onChanges when
squared biasthe model's capacity to represent f the model family changes
variancecapacity relative to the sample sizeeither the model or n changes
σ 2 the data-generating processnever, for any model

Flexibility moves the first two in opposite directions. A more flexible model follows the true f more closely, lowering bias, and responds more to the particular sample it was given, raising variance. Their sum therefore has a minimum at some intermediate flexibility, which is why the best model is not the most flexible one available. The minimum's location depends on n : more data lowers variance at a given flexibility, so the optimum moves toward more flexible models as the sample grows.

A floor that no model clears. Because σ 2 enters every model's test error identically, it bounds expected error from below. A reported test error below the plausible noise level is evidence that information about the held-out observations reached the fit, rather than evidence of an unusually good model.

What CV k estimates, precisely. Not the error of the model fitted to all n observations, but the error of the procedure applied to training sets of size about n ( k − 1 ) / k . For large n the distinction is immaterial. For small n each fold fits on visibly less data than the full sample, so the estimate is mildly pessimistic. A bias that shrinks as k grows and vanishes at k = n .

Loss functions other than squared error. The three-term decomposition is exact for squared-error loss and does not carry over unchanged to classification error or log-loss. The qualitative trade between bias and variance survives; the clean additive separation does not.

Intuition

Training error, optimism, and sampling variability

Fitting a model means choosing parameters that make it agree with a particular set of observations. Those observations carry noise, and the fitting procedure cannot separate noise from signal. It minimises disagreement with whatever it was given. Some of the agreement it achieves is therefore agreement with accidents of that sample, which will not recur. Asking the model about those same observations counts that agreement as success.

The more flexible the model, the more of the sample's accidents it can accommodate, so training error keeps falling while the agreement it achieves becomes less transferable.

Bias and variance are answers to two different questions about a procedure. Imagine refitting the same model on many independent samples from the same population.

  • Where do the fits sit on average, compared with the truth? A systematic offset is bias. A straight line fitted to genuinely curved data misses the same way in every sample, however many observations each contains.
  • How much do the fits differ from one another? That spread is variance. A model with many parameters relative to the data produces a visibly different fit each time, because each sample's noise pulls it somewhere else.

A rigid model has high bias and low variance; a flexible one has the reverse. Neither quantity is observable from a single sample, which is why they are reasoned about rather than measured directly.

The irreducible term explains why error has a floor. If two observations share the same x but differ in y , no function of x can predict both. That disagreement enters every model's test error identically. A reported test error below σ 2 therefore indicates that information about the held-out observations reached the fit, not that the model is unusually good.

Cross-validation withholds repeatedly rather than once. A single split gives one estimate, and which observations landed in the held-out portion affects it; an unlucky split can make a good model look poor. Rotating the held-out fold means every observation is tested on exactly once, and the spread across folds shows how much that choice mattered.

Example

Error estimates that answer different questions

Each case below reports an error figure. What separates them is not the arithmetic but which quantity the figure estimates.

A figure that estimates nothing about future performance. A shop fits a model to last year's transactions, predicts those same transactions, and reports 94% accuracy. The number is correct and describes how well the model reproduces answers it was given. It places no bound on next month's accuracy, and would be near 100% for a model that had memorised the file.

A figure that estimates performance on a fresh draw from the same population. The same shop holds out a random 20% of last year's transactions, fits on the rest, and reports 87% on the held-out portion. This estimates accuracy on another transaction from last year, useful if the population is stable, and silent about drift.

A figure that estimates next month's performance. The shop fits on January–October and evaluates on November–December, reporting 81%. The drop from 87% is the part of the earlier figure that came from knowing the future. This is the number to quote when the model will be used to predict forward.

A figure made optimistic by selection. The shop tries 40 feature sets, cross-validating each, and reports the best score of 89%. That value is the minimum of 40 noisy estimates and sits below what the chosen feature set would score on fresh data. An outer split not used in the search would give the honest figure.

A figure that reveals leakage. A model predicting whether a customer will return within 30 days reports 99.2% accuracy. Inspection shows one predictor is the date of the customer's next visit, recorded after the outcome it predicts. An error far below what the problem plausibly allows is evidence about the data pipeline rather than about the model.

---

The first and last are not estimates of test error at all. The middle three are, and they estimate different populations: another transaction from the same period, a transaction from a later period, and the performance of a whole selection procedure. Quoting one where another is meant is the most common way an honest calculation supports a claim it does not license.

Procedure

Running a cross-validation, and reporting what it gives

To estimate test error by k -fold cross-validation.

  1. Check what the observations are before splitting. Ordered in time, repeated on the same subject, or clustered within a group each forbid random folds. Step 5 covers those cases.
  2. Choose k . Five or ten are the usual choices. Larger k fits on more data, so each fit is closer to the one the full sample would give, at proportionally more computation; k = n is leave-one-out.
  3. Partition the observations into k folds of roughly equal size, assigned at random.
  4. For each fold i : fit the model on the other k − 1 folds, predict the observations in fold i , and record the loss L i . Every step of fitting, including any scaling, imputation or variable selection, must happen inside this loop, using only the k − 1 training folds.
  5. Report the mean CV k = 1 k ∑ i L i together with the spread of the L i , usually their standard deviation divided by k .

When the observations are ordered in time. Use a temporal holdout: fit on observations up to some time t and evaluate on observations after it. Repeating this at several cut points gives several estimates. A fold must never contain an observation earlier than one used to fit it.

When observations are grouped. Assign whole groups to folds rather than individual observations, so no subject or cluster appears on both sides of a split.

To choose a tuning parameter and still report an honest error. Cross-validation used to select cannot also estimate: the winning value was chosen because it scored well on those folds. Use a nested scheme.

  1. Split into outer folds.
  2. Within each outer training set, run a complete inner cross-validation to choose the tuning parameter.
  3. Fit with the chosen value on the whole outer training set, and evaluate once on the outer held-out fold.
  4. Average the outer fold losses. That average is the estimate of the procedure's test error, tuning included.

Checks. Compare the cross-validation figure with training error: the first should be larger, and a training error near zero alongside a much larger held-out error indicates a model flexible enough to interpolate. If the fold losses vary by more than the difference between two models under comparison, the comparison does not distinguish them.

Worked example

Equal training error, different held-out error

Twelve observations, x = 1 , … , 12 , with

y = 2.1 ,   4.3 ,   5.9 ,   8.4 ,   9.8 ,   12.5 ,   13.9 ,   16.4 ,   17.8 ,   20.6 ,   21.9 ,   24.3 .

The underlying relationship is close to y = 2 x . Two candidates: a straight line, and a degree-5 polynomial. Both fitted by least squares.

Step 1: fit each on all twelve and record training error.

ModelTraining MSE
degree 1 0.0770
degree 5 0.0761

On this evidence the two are indistinguishable, and the polynomial is marginally ahead. Training error cannot separate them, and it never favours the simpler model: adding parameters can only reduce the quantity being minimised.

Step 2: partition into 4 folds. Observations are assigned by position, { 1 , 5 , 9 } , { 2 , 6 , 10 } , { 3 , 7 , 11 } , { 4 , 8 , 12 } , so each fold spans the range rather than occupying one end of it.

Step 3: for each fold, fit on the other nine and score the held-out three.

Folddegree 1degree 5
1 0.1362 0.2034
2 0.1840 0.2543
3 0.1261 0.1568
4 0.0794 0.5755

Step 4: average, and report the spread.

Model CV 4 fold s.d.s.e. of the mean
degree 1 0.1314 0.0429 0.0215
degree 5 0.2975 0.1896 0.0948

Reading the result.

The line's held-out error is 0.1314 against a training error of 0.0770 : larger, as expected, since the training figure includes agreement with this sample's noise.

The polynomial's held-out error is 0.2975 , 3.9 times its training error of 0.0761 and 2.3 times the line's held-out error. The extra parameters gained 0.0009 of training error and cost 0.1661 of test error.

The difference between the two, 0.1661 , is larger than either standard error, so the ranking is not an artefact of which observations landed in which fold.

The fourth fold is the informative one. The polynomial scores 0.5755 there against 0.1568 on fold 3, a factor of nearly four between folds of the same size. Fold 4 contains observation 12, the largest x , and a degree-5 fit on the remaining nine extrapolates badly at the edge of the range. That variability is the variance term: the same procedure on a slightly different sample produces a visibly different fit. Its fold s.d. of 0.1896 against the line's 0.0429 measures the same thing.

Reporting 0.2975 and 0.1314 gives the ranking but not the reason. The fold losses show the polynomial is not uniformly worse, on fold 3 it is close to the line, but unreliable, with its average driven by one fold where extrapolation failed. A model that is usually adequate and occasionally far wrong is a different proposition from one that is uniformly mediocre, and only the spread distinguishes them.

Contrast

Splits that differ by one decision

Random folds against a temporal holdout, on daily sales.

random 5-foldtemporal holdout
held-out dayany daythe last 20% of days
fitted ondays before and after itdays before it only
answershow well the model fills a gaphow well it forecasts
typical errorlowerhigher

Both are correct procedures; they estimate different quantities. Random folds let a model learn from Tuesday and Thursday to predict Wednesday, which is unavailable when Wednesday is genuinely next. The lower error is not a better model but a different, easier question.

Cross-validation for selection against cross-validation for estimation.

selectingestimating
questionwhich candidate is best?how will the chosen one perform?
quantity usedthe ranking of the scoresthe value of one score
valid on the same foldsyesno

The ranking is comparatively robust: shared fold noise affects all candidates together. The winner's value is not, because selecting the minimum selects partly for favourable noise.

Scaling inside the fold loop against scaling before it.

Centre and scale using the whole dataset, and each training fit has used the held-out fold's mean. The leak is small for a mean over many observations and can be large for a variable-selection step on few. The distinction is not the size of the effect but whether the procedure can be described honestly: a figure obtained with outside information does not estimate performance on data the model has not seen.

k = 5 against k = n .

5-foldleave-one-out
fits required5 n
training set size 0.8 n n − 1
bias of the estimateslightly pessimisticnearly unbiased
variance of the estimatelowerhigher

Leave-one-out fits on almost the full sample, so it estimates the error of the model actually being fitted. Its n training sets overlap in all but one observation, making the held-out losses strongly dependent and their average more variable. Five or ten folds trade a little bias for a steadier estimate.

In each, one decision changes what quantity is being estimated rather than how precisely. That is why the question to ask about a validation scheme is which quantity it estimates, before asking whether the number it produced is good.

Warning

Reusing cross-validation for selection and reporting

Cross-validation is usually run many times: once per candidate tuning value, or per feature set, or per model family. The best score among those runs is then reported as the estimate of test error. That figure is optimistic, and the reason is the same one that makes training error optimistic, one level up.

Why selecting on a score biases it. Each candidate's cross-validation score is an estimate carrying its own random error from the particular fold assignment. Taking the minimum over many candidates selects partly for genuinely lower error and partly for a favourable draw. The winner's score is therefore below what that candidate would achieve on fresh data, and the gap grows with the number of candidates tried.

The effect is not small when the candidate set is large. Comparing forty tuning values on twelve folds means taking a minimum over forty noisy estimates, and the minimum of forty draws sits well below their common mean.

What this looks like in practice. A tuning curve is computed, its lowest point is read off, and that value is quoted as the expected error. Nothing in the procedure signals a problem: the folds were held out, no observation was evaluated by a model fitted to it, and the arithmetic is correct. The reuse is at the level of the selection, which the fold structure does not protect.

The fix is an outer split. Reserve data the selection never touches, or nest the cross-validations as the procedure block describes. The outer estimate then measures the whole procedure, tuning included, rather than the tuned model alone. It will be larger than the inner best score, and that is the honest number.

A related failure with the same shape. Any step fitted on all the data before splitting leaks in the same way: scaling by the overall mean and standard deviation, imputing with a column median, selecting variables by their correlation with the response. Each uses held-out observations to shape the fit later evaluated on them. Variable selection on the full dataset is the most damaging, because it can produce a respectable cross-validation score from predictors that are pure noise.

What survives the warning. Cross-validation remains the right tool for choosing among candidates; comparing scores computed the same way is what it is for. The caution concerns the second use, reporting the winner's score as an estimate of future performance.

Application

Where the estimate is the deliverable

Clinical prediction. A model estimating a patient's risk is evaluated on patients from hospitals that contributed no training data, because performance falls when case mix or measurement practice differs. A model validated only by cross-validation within one hospital's records has been shown to work where it was built, which is the weakest claim of interest.

Credit scoring. Applicants are scored over time and the population shifts, so evaluation uses a temporal holdout: fit on applications up to a date, evaluate on applications after it. A random split would report the error of filling gaps in a known period rather than the error of scoring next quarter's applicants.

Recommendation. Users and items both recur, so a random split over interactions can place a user's later interaction in the training set and an earlier one in the held-out fold. Evaluation is therefore usually split by user and by time together, holding out each user's most recent interactions.

Competition leaderboards. A public leaderboard scored on a fixed held-out set is a shared validation set, and participants who submit repeatedly are selecting on it. Final standings are computed on a second, private set precisely because the public score stops estimating test error once it has been optimised against. The selection bias of the warning block, at the scale of a whole community.

Scientific reporting. Where a model supports a claim rather than a product, what is reported is the estimate together with how it was obtained: the splitting scheme, whether tuning was nested, and the spread across folds. Two papers reporting the same accuracy on the same data are not making the same claim if one tuned on its evaluation folds.

---

The common shape. Each case asks the same question the unit began with, how will this perform on data it has not seen, and each answers it by making the held-out data resemble the unseen data in the way that matters: a different hospital, a later date, a different user, a second private set. The choice of scheme follows from what "unseen" will actually mean when the model is used, which is a question about the application rather than about statistics.

Next step

Practice Estimating Out-of-Sample Error

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.