Fitting Without a Formula
What you will be able to do
The learner can produce a fitted value from a nearest-neighbour or kernel-weighted local average, explain how the neighbourhood size or bandwidth moves bias against variance, choose that parameter by held-out error, and say what a nonparametric fit gives up in exchange for assuming no functional form.
Orientation
Declining to name a shape
The previous unit chose among shapes: a line, a polynomial, an exponential, a rate model. Each choice was made before looking closely at the data, on grounds of what the response is and what the mechanism suggests.
This unit declines the choice. To predict at a point, look at the observations near that point and average them. Do it again at the next point. No formula is ever written down, and no assumption about curvature or growth has to be defended.
One decision remains, and it carries everything the discarded assumptions used to: how near is near. Consult too few neighbours and the fit chases noise. Consult too many and it flattens the structure it was meant to find. The whole unit is about that parameter, what it trades, why the training data cannot choose it, and where the method is weakest.
The last point, stated in advance: local averaging is at its worst at the edges of the observed range, where a neighbourhood can only be one-sided, and it offers nothing at all outside that range, where there are no neighbours to average. The method that needs no assumptions about shape also provides no basis for saying what happens beyond the data.
Definition
How the three smoothing parameters correspond
The canonical statements above give the three estimators. What follows is how their parameters line up and where each differs.
They control one quantity under three names.
| Method | Parameter | More flexible when | Weights |
|---|---|---|---|
| equal inside the neighbourhood, zero outside | |||
| kernel | fall continuously, never exactly zero | ||
| spline | implied by the penalty |
Where the differences bite. A
A
What
Why the training criterion is useless here. Each of the three has a setting that interpolates:
Intuition
Why zero training error is the warning sign
Set
The reason is that each prediction is now a single observation. An observation is the true value plus noise, so the prediction inherits the whole of that noise. Collect the data again and every fitted value moves. The curve has recorded the sample rather than estimated the relationship, and the criterion that reports zero is measuring agreement with data the fit was handed.
Raising
What this looks like on a curve. At
The boundary is a systematic problem, not a random one. Near the middle of the data a neighbourhood straddles the query point, and the observations below and above contribute errors of opposite sign that partly cancel. At the left edge there is nothing below, so every neighbour lies to the right and every one has a systematically different expected response. The estimate is pulled inward. This is bias, present at every sample size, and it grows with the neighbourhood.
Which is why extrapolation is not merely unreliable but unavailable. Beyond the observed range there are no neighbours at all. A
Example
Where local averaging is efficient, and where it is not
Local averaging estimates
A smooth curve of unknown form. A response rising and then levelling off, with no theory fixing whether the curve is logarithmic, a saturating exponential, or something else. Averaging neighbours follows the shape without committing to a form. This is the case the method is designed for.
A relationship known to be linear. A local average can approximate a line, but it re-estimates the fit at every point from the observations in a neighbourhood, while least squares uses all
A discontinuity. Where the regression function has a jump, any neighbourhood spanning the threshold averages across it, so the fitted curve shows a transition whose width is roughly the neighbourhood width rather than a step. The output contains no indication that this has happened. A regression tree can place a split at the threshold, so a step function is within the class of functions a tree can represent; whether a tree fitted to a particular sample locates the threshold accurately depends on the sample and on how the tree is grown and pruned.
Several predictors. With one predictor a neighbourhood holding
Sparse regions. Where observations thin out, a
Local averaging is efficient where the regression function is smooth, its form is unknown, and the predictors are few. It is less efficient than a parametric fit when the form is known, and it is biased near discontinuities and at the edges of the observed range, in both cases without a signal in the output.
Procedure
Producing a fitted value, and choosing the smoothing
To compute a
- Measure the distance from
to every observation. With several predictors, scale them first. A neighbourhood computed from unscaled variables is decided by whichever has the largest units. - Take the
smallest distances. - Average their responses. That is the prediction. Ties in distance are broken arbitrarily; if a tie changes the answer noticeably,
is too small.
To compute a kernel estimate at
- Weight each observation by
, falling with distance. - Form the weighted average
. Dividing by the sum of weights is what makes it an average rather than a sum that shrinks toward zero far from the data. - Report the effective number of observations,
. An estimate resting on two effective observations should not be presented like one resting on twenty.
To choose the smoothing parameter.
- Fix a grid of candidates, for
-NN, integers from 1 up to a sizeable fraction of ; for a kernel, bandwidths spanning an order of magnitude. - For each candidate, compute held-out error. Leave-one-out is natural here because refitting costs nothing: predicting an observation from its neighbours simply excludes itself.
- Take the minimising candidate, and look at the whole curve rather than the single winner. A flat minimum means several values are equivalent and the choice does not matter; a sharp one means it does.
- Never select on training error. It falls monotonically as the fit grows more local and is exactly zero at
, so its optimum is always the degenerate fit.
If the same data must also report performance, the selection is itself fitting, so the honest estimate needs an outer split that the tuning never saw.
To decide whether to use this family at all.
- How many predictors? Beyond a few, neighbourhoods stop being local.
- Is a functional form genuinely known from the subject matter? If so, assuming it uses the data far more efficiently.
- Is the relationship expected to be smooth? A known threshold argues for a tree.
- Is a coefficient needed for reporting? This family returns none.
Checks. Plot the fit against the data at the chosen setting and at neighbouring settings; the chosen one should look neither jagged nor visibly flattened at a peak. Compare the held-out error against a straight-line fit: if the line wins, the extra flexibility is not earning its cost.
Worked example
Two criteria that disagree about the best neighbourhood
Twenty observations,
Step 1: one prediction, several neighbourhoods. At
| Neighbours used | ||
|---|---|---|
| 1 | ||
| 3 | ||
| 5 | ||
| 9 | ||
| 15 | ||
| 20 | all |
The estimate drifts downward as
Step 2: training error across the whole fit.
| Training RSS | Roughness | |
|---|---|---|
| 1 | ||
| 3 | ||
| 5 | ||
| 9 | ||
| 15 |
Training error is minimised at
Step 3: leave-one-out cross-validation. Each observation is predicted from the others, so no fit sees the point it is scored on.
| LOOCV MSE | LOOCV MSE | |||
|---|---|---|---|---|
| 1 | 6 | |||
| 2 | 7 | |||
| 3 | 9 | |||
| 4 | 12 | |||
| 5 | 15 |
Step 4: read the disagreement. The two criteria rank the candidates almost oppositely at the flexible end.
The gap between
Step 5: the same trade with continuous weights. A Gaussian kernel at
| Effective | ||
|---|---|---|
At
Step 6: the boundary. With
| Neighbours | Left | Right | |
|---|---|---|---|
| 1 | 0 | 4 | |
| 2 | 1 | 3 | |
| 10 | 2 | 2 | |
| 19 | 3 | 1 | |
| 20 | 4 | 0 |
At
Contrast
Settings and criteria that look comparable
Training error against held-out error, on the same fits.
| Training RSS | LOOCV MSE | |
|---|---|---|
| 1 | ||
| 3 | ||
| 4 | — | |
| 9 | ||
| 15 |
The two columns rank
| Gaussian kernel, | ||
|---|---|---|
| effective observations | ||
| estimate at | ||
| fitted curve | discontinuous | continuous |
| where data are sparse | reaches further, silently | uses very few, visibly |
At comparable flexibility the two agree closely on the estimate and differ in what they do when the data thin out. Neither behaviour is safer in general; the kernel's instability is at least legible.
Bandwidth against neighbourhood size.
Both control smoothing, and they respond to the data's density oppositely.
A local average against a regression tree.
Both are local and neither assumes a form. A tree partitions by searching for splits, so it represents a genuine threshold exactly and a smooth trend as a staircase. Local averaging assumes smoothness, so it represents a smooth trend well and a threshold as a ramp whose width is the neighbourhood's. The two fail on each other's easy cases.
A smoothing spline at large
These converge: as
Warning
The failures that do not announce themselves
A local-averaging fit returns a number at every query point, whatever is happening underneath. Three situations produce confident output from an estimate that is not supported, and none of them raises an error.
Beyond the observed range there are no neighbours. A
At the edges the neighbourhood is one-sided. With
A discontinuity is rendered as a ramp. Any neighbourhood spanning a genuine step averages across it, so the fit shows a gradual transition whose width is the neighbourhood's width rather than a property of the data. The output looks like evidence of a smooth change. If a threshold is suspected, a method that splits rather than averages will represent it, and comparing the two on held-out error is the check.
---
And one that announces itself as a success. Setting
---
With several predictors, "local" stops meaning local. To capture a tenth of the observations in
Application
Where the shape is the question
Growth charts. A child's height-for-age reference is a set of smooth curves through population data, and no formula relates height to age. The curves are produced by smoothing, with the bandwidth chosen so that the result follows real features of growth without following sampling noise. What is published is the curve itself, since there is no coefficient anyone would quote.
Dose-response screening. Before a parametric model is committed to, a smooth fit shows what shape the data suggest: monotone, saturating, or non-monotone. Here the nonparametric fit is a diagnostic rather than the deliverable, and it is used precisely because it will not impose the shape that is being investigated.
Environmental time series. Seasonal and long-term components are separated by local smoothing at two different bandwidths, the narrow one following the annual cycle and the wide one the trend beneath it. The bandwidths encode what counts as season and what counts as trend, which is a decision about the science rather than a technical parameter.
Spatial interpolation. Predicting a pollutant concentration between monitoring stations is a weighted average of nearby readings. The boundary problem is concrete here: a station at the edge of the network has neighbours on one side only, so its estimates are pulled inward, and the network's edge is exactly where extrapolation is most tempting.
Calibration curves in the laboratory. Where an instrument's response is monotone but not of any known form, a spline through calibration standards converts a reading into a concentration. The fitted curve is valid only across the range of the standards, and a reading outside it is reported as out of range rather than converted. The same rule the warning block gives, enforced as laboratory practice.
---
What decides the choice in each case. No formula was available and none was wanted: in the growth charts and the calibration curve because the curve is the product, and in the dose-response screen because imposing a form would prejudge the question being asked. Where a form is known from the subject matter, these methods are the wrong instrument, since they spend data discovering what was already known.