Fitting Without a Formula
Estimating a response at a point by averaging nearby observations rather than by fitting a formula, how the neighbourhood size or bandwidth trades bias against variance, why training error cannot choose it, and what disappears when no functional form is assumed.
Definition
A nonparametric regression estimates
where
Kernel regression replaces the sharp neighbourhood with weights that fall continuously with distance:
for the Gaussian kernel. The bandwidth
The smoothing parameter controls flexibility. Small
A smoothing spline reaches the same trade from the other direction, minimising
Assumptions and scope
Local averaging assumes the response varies smoothly enough near a query point for neighbouring observations to be informative about it. At a genuine discontinuity the method averages across the jump and is wrong on both sides.
Training error cannot select the smoothing parameter: it falls monotonically as the neighbourhood shrinks and reaches zero at
. Selection requires held-out error, and the same reuse caution applies as for any tuned parameter.Distances depend on the units of the predictors, so with several predictors the variables must be scaled before a neighbourhood means anything.
The data needed for a given precision grows rapidly with the number of predictors, because a neighbourhood holding a fixed fraction of the observations must span almost the full range of each variable. Local averaging is a method for few predictors.
These methods do not extrapolate a trend. Outside the observed range the nearest observations are still the ones at the boundary, so predictions there are controlled by the same edge data however far out the query lies and cannot adapt to a relationship that continues changing. What is absent beyond the boundary is local data support, not neighbours.
Worked material
Example
Where local averaging is efficient, and where it is not
Local averaging estimates
A smooth curve of unknown form. A response rising and then levelling off, with no theory fixing whether the curve is logarithmic, a saturating exponential, or something else. Averaging neighbours follows the shape without committing to a form. This is the case the method is designed for.
A relationship known to be linear. A local average can approximate a line, but it re-estimates the fit at every point from the observations in a neighbourhood, while least squares uses all
A discontinuity. Where the regression function has a jump, any neighbourhood spanning the threshold averages across it, so the fitted curve shows a transition whose width is roughly the neighbourhood width rather than a step. The output contains no indication that this has happened. A regression tree can place a split at the threshold, so a step function is within the class of functions a tree can represent; whether a tree fitted to a particular sample locates the threshold accurately depends on the sample and on how the tree is grown and pruned.
Several predictors. With one predictor a neighbourhood holding
Sparse regions. Where observations thin out, a
Local averaging is efficient where the regression function is smooth, its form is unknown, and the predictors are few. It is less efficient than a parametric fit when the form is known, and it is biased near discontinuities and at the edges of the observed range, in both cases without a signal in the output.
Contrast
Settings and criteria that look comparable
Training error against held-out error, on the same fits.
| Training RSS | LOOCV MSE | |
|---|---|---|
| 1 | ||
| 3 | ||
| 4 | — | |
| 9 | ||
| 15 |
The two columns rank
| Gaussian kernel, | ||
|---|---|---|
| effective observations | ||
| estimate at | ||
| fitted curve | discontinuous | continuous |
| where data are sparse | reaches further, silently | uses very few, visibly |
At comparable flexibility the two agree closely on the estimate and differ in what they do when the data thin out. Neither behaviour is safer in general; the kernel's instability is at least legible.
Bandwidth against neighbourhood size.
Both control smoothing, and they respond to the data's density oppositely.
A local average against a regression tree.
Both are local and neither assumes a form. A tree partitions by searching for splits, so it represents a genuine threshold exactly and a smooth trend as a staircase. Local averaging assumes smoothness, so it represents a smooth trend well and a threshold as a ramp whose width is the neighbourhood's. The two fail on each other's easy cases.
A smoothing spline at large
These converge: as
Common errors
Common misconception
That a local fit using fewer neighbours is more accurate, because it uses only the observations closest to the query point and so is less contaminated by distant ones. Shrinking the neighbourhood lowers bias and raises variance, and the sum has a minimum at some intermediate size. At
Common misconception
That a nonparametric fit makes no assumptions, so it is the safe choice whenever the shape of a relationship is unknown. What is dropped is the assumption of a particular functional form; several assumptions remain, and one of them is strong. Local averaging assumes the response varies smoothly enough that observations near a query point are informative about it, which is why the method fails at a genuine discontinuity and why it degrades near the boundary of the observed range, where the available neighbours lie on one side only. It also gives up what a parametric form provides: coefficients that can be reported, and any basis for predicting outside the observed range, where there are no neighbours to average. The choice is a trade rather than a removal of assumptions.
Related units
Requires
Connected
- Trees and Ensembles of Them (related)