Fitting Without a Formula

Estimating a response at a point by averaging nearby observations rather than by fitting a formula, how the neighbourhood size or bandwidth trades bias against variance, why training error cannot choose it, and what disappears when no functional form is assumed.

Definition

A nonparametric regression estimates E [ y ∣ x ] at a query point from observations near it, without committing to a functional form.

k -nearest-neighbour regression takes the k observations closest to the query point x 0 and averages their responses:

f ^ ( x 0 ) = 1 k ∑ i ∈ N k ( x 0 ) y i ,

where N k ( x 0 ) is the index set of those k observations. Weights are equal inside the neighbourhood and zero outside it, and the neighbourhood's width changes with where the data are dense.

Kernel regression replaces the sharp neighbourhood with weights that fall continuously with distance:

f ^ ( x 0 ) = ∑ i K h ( x 0 , x i ) y i ∑ i K h ( x 0 , x i ) , K h ( x 0 , x ) = exp ( − 1 2 ( x − x 0 h ) 2 )

for the Gaussian kernel. The bandwidth h plays the role k plays above: larger h spreads weight over more observations.

The smoothing parameter controls flexibility. Small k or small h produces a fit that follows the data closely, low bias, high variance. Large k or large h produces a flatter fit, high bias, low variance. At k = 1 , with each observation counted as its own nearest neighbour and predictor values distinct, the fit passes through every training observation and its training error is exactly zero. Where two observations share a predictor value and differ in response, the fit cannot pass through both and the training error is positive.

A smoothing spline reaches the same trade from the other direction, minimising ∑ i ( y i − f ( x i ) ) 2 + λ ∫ f ″ ( t ) 2 d t over all sufficiently smooth f . The penalty parameter λ controls flexibility as k and h do, with λ → ∞ forcing a straight line.

Assumptions and scope

  • Local averaging assumes the response varies smoothly enough near a query point for neighbouring observations to be informative about it. At a genuine discontinuity the method averages across the jump and is wrong on both sides.

  • Training error cannot select the smoothing parameter: it falls monotonically as the neighbourhood shrinks and reaches zero at k = 1 . Selection requires held-out error, and the same reuse caution applies as for any tuned parameter.

  • Distances depend on the units of the predictors, so with several predictors the variables must be scaled before a neighbourhood means anything.

  • The data needed for a given precision grows rapidly with the number of predictors, because a neighbourhood holding a fixed fraction of the observations must span almost the full range of each variable. Local averaging is a method for few predictors.

  • These methods do not extrapolate a trend. Outside the observed range the nearest observations are still the ones at the boundary, so predictions there are controlled by the same edge data however far out the query lies and cannot adapt to a relationship that continues changing. What is absent beyond the boundary is local data support, not neighbours.

Worked material

Example

Where local averaging is efficient, and where it is not

Local averaging estimates E [ Y ∣ X = x ] from observations near x , without assuming a functional form. Under regularity conditions, including a smooth regression function, a bandwidth shrinking at an appropriate rate, and enough data near x , the estimate is consistent. The cases below differ in what that costs.

A smooth curve of unknown form. A response rising and then levelling off, with no theory fixing whether the curve is logarithmic, a saturating exponential, or something else. Averaging neighbours follows the shape without committing to a form. This is the case the method is designed for.

A relationship known to be linear. A local average can approximate a line, but it re-estimates the fit at every point from the observations in a neighbourhood, while least squares uses all n observations to estimate two parameters. For comparable precision the local method needs more data, and it returns no slope coefficient to report.

A discontinuity. Where the regression function has a jump, any neighbourhood spanning the threshold averages across it, so the fitted curve shows a transition whose width is roughly the neighbourhood width rather than a step. The output contains no indication that this has happened. A regression tree can place a split at the threshold, so a step function is within the class of functions a tree can represent; whether a tree fitted to a particular sample locates the threshold accurately depends on the sample and on how the tree is grown and pruned.

Several predictors. With one predictor a neighbourhood holding 10 % of the observations spans about 10 % of the range. With ten predictors, holding 10 % of the observations requires spanning about 0.10 1 / 10 ≈ 80 % of each variable's range, so the neighbourhood is no longer local. Local averaging suits problems with few predictors.

Sparse regions. Where observations thin out, a k -nearest-neighbour estimate draws on points far from x and reports an average whose stated precision does not reflect that distance, while a fixed-bandwidth kernel estimate is visibly unstable because it uses few points. The second makes the data shortage apparent in the output.

Local averaging is efficient where the regression function is smooth, its form is unknown, and the predictors are few. It is less efficient than a parametric fit when the form is known, and it is biased near discontinuities and at the edges of the observed range, in both cases without a signal in the output.

Contrast

Settings and criteria that look comparable

Training error against held-out error, on the same fits.

k Training RSSLOOCV MSE
1 0.0000 14.6640
3 79.5056 8.6377
4— 8.4585
9 220.8447 15.9562
15 439.9587 27.8433

The two columns rank k = 1 oppositely: best by the first, close to worst by the second. Training error is monotone in k and held-out error is U-shaped, so only one of them contains the information needed to choose. The disagreement is largest exactly at the setting a reader is most tempted by.

k -nearest neighbours against a kernel, at matched flexibility.

k -NN, k = 4 Gaussian kernel, h = 1.0
effective observations 4 3.54
estimate at x 0 = 10 ≈ 24.6 24.9411
fitted curvediscontinuouscontinuous
where data are sparsereaches further, silentlyuses very few, visibly

At comparable flexibility the two agree closely on the estimate and differ in what they do when the data thin out. Neither behaviour is safer in general; the kernel's instability is at least legible.

Bandwidth against neighbourhood size.

Both control smoothing, and they respond to the data's density oppositely. k fixes the number of observations and lets the width float; h fixes the width and lets the count float. In a dense region they behave alike. In a sparse one, k -NN quietly averages distant points while the kernel returns a noisy estimate from two or three.

A local average against a regression tree.

Both are local and neither assumes a form. A tree partitions by searching for splits, so it represents a genuine threshold exactly and a smooth trend as a staircase. Local averaging assumes smoothness, so it represents a smooth trend well and a threshold as a ramp whose width is the neighbourhood's. The two fail on each other's easy cases.

A smoothing spline at large λ against ordinary least squares.

These converge: as λ → ∞ the curvature penalty forces f ″ = 0 and the spline becomes the least-squares straight line. The nonparametric family contains the parametric fit as its most heavily smoothed member, which is the precise sense in which assuming a form is the strongest smoothing available, and why the choice between them is a matter of degree rather than of kind.

Common errors

Common misconception

That a local fit using fewer neighbours is more accurate, because it uses only the observations closest to the query point and so is less contaminated by distant ones. Shrinking the neighbourhood lowers bias and raises variance, and the sum has a minimum at some intermediate size. At k = 1 the fit passes through every training observation and its training error is exactly zero, which looks like perfect accuracy and is instead the point of maximum variance: the prediction at any query point is one noisy observation, so a resample of the data moves it by the full noise standard deviation. Training error cannot reveal this because it falls monotonically as the neighbourhood shrinks, which is why the parameter is chosen by held-out error.

Common misconception

That a nonparametric fit makes no assumptions, so it is the safe choice whenever the shape of a relationship is unknown. What is dropped is the assumption of a particular functional form; several assumptions remain, and one of them is strong. Local averaging assumes the response varies smoothly enough that observations near a query point are informative about it, which is why the method fails at a genuine discontinuity and why it degrades near the boundary of the observed range, where the available neighbours lie on one side only. It also gives up what a parametric form provides: coefficients that can be reported, and any basis for predicting outside the observed range, where there are no neighbours to average. The choice is a trade rather than a removal of assumptions.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.