Fitting Without a Formula

What you will be able to do

The learner can produce a fitted value from a nearest-neighbour or kernel-weighted local average, explain how the neighbourhood size or bandwidth moves bias against variance, choose that parameter by held-out error, and say what a nonparametric fit gives up in exchange for assuming no functional form.

Orientation

Declining to name a shape

The previous unit chose among shapes: a line, a polynomial, an exponential, a rate model. Each choice was made before looking closely at the data, on grounds of what the response is and what the mechanism suggests.

This unit declines the choice. To predict at a point, look at the observations near that point and average them. Do it again at the next point. No formula is ever written down, and no assumption about curvature or growth has to be defended.

One decision remains, and it carries everything the discarded assumptions used to: how near is near. Consult too few neighbours and the fit chases noise. Consult too many and it flattens the structure it was meant to find. The whole unit is about that parameter, what it trades, why the training data cannot choose it, and where the method is weakest.

The last point, stated in advance: local averaging is at its worst at the edges of the observed range, where a neighbourhood can only be one-sided, and it offers nothing at all outside that range, where there are no neighbours to average. The method that needs no assumptions about shape also provides no basis for saying what happens beyond the data.

Definition

How the three smoothing parameters correspond

The canonical statements above give the three estimators. What follows is how their parameters line up and where each differs.

They control one quantity under three names. k , h and λ all set how many observations effectively contribute to each prediction. A useful common currency is the effective number of observations behind an estimate, for kernel weights w i , the quantity ( ∑ i w i ) 2 / ∑ i w i 2 . For k -nearest neighbours it is exactly k ; for a Gaussian kernel it grows with h .

MethodParameterMore flexible whenWeights
k -NN k k smallequal inside the neighbourhood, zero outside
kernel h h smallfall continuously, never exactly zero
spline λ λ smallimplied by the penalty

Where the differences bite. A k -NN fit is discontinuous: as the query point moves, an observation enters the neighbourhood and another leaves, and the prediction jumps. A kernel fit is continuous, because a distant observation's weight approaches zero smoothly rather than being cut off. This is why a k -NN curve looks ragged and a kernel curve does not, at comparable flexibility.

A k -NN neighbourhood adapts to the data's density and a fixed bandwidth does not. Where observations are sparse, the k nearest ones are far away and the estimate quietly averages over a wide region; a fixed h instead averages very few observations there and returns a noisy estimate. Neither behaviour is right in general: the first hides that it is extrapolating locally, the second announces it by being unstable.

What λ → ∞ gives. The spline penalty ∫ f ″ ( t ) 2 d t measures curvature, so driving λ up drives curvature to zero and the fit becomes the least-squares straight line. The parametric model is the limiting case of the nonparametric one, which is the precise sense in which assuming a form is the strongest possible smoothing.

Why the training criterion is useless here. Each of the three has a setting that interpolates: k = 1 , h → 0 , λ = 0 . At that setting the fitted curve passes through every training observation and the training error is exactly zero. A criterion whose optimum is always the same degenerate fit carries no information about flexibility, which is why selection uses held-out error in every case.

Intuition

Why zero training error is the warning sign

Set k = 1 and something suspicious happens: the fitted curve passes through every training point and the training error is exactly zero. No model can do better by that measure, and the model is at its worst.

The reason is that each prediction is now a single observation. An observation is the true value plus noise, so the prediction inherits the whole of that noise. Collect the data again and every fitted value moves. The curve has recorded the sample rather than estimated the relationship, and the criterion that reports zero is measuring agreement with data the fit was handed.

Raising k trades one error for another. Averaging k observations divides the noise contribution by k , so variance falls. It also pulls in observations further away, whose expected responses genuinely differ from the query point's, so bias rises. The two move in opposite directions and their sum has a minimum somewhere in between. The same shape as the flexibility trade in every other unit, with k running backwards, since large k means less flexible.

What this looks like on a curve. At k = 1 the fit is jagged, visiting every point. As k grows the fit smooths, first revealing the underlying shape and then beginning to flatten it: a peak gets clipped, a trough filled in, because the average at those points includes observations from both slopes. At k = n the fit is a horizontal line at the overall mean.

The boundary is a systematic problem, not a random one. Near the middle of the data a neighbourhood straddles the query point, and the observations below and above contribute errors of opposite sign that partly cancel. At the left edge there is nothing below, so every neighbour lies to the right and every one has a systematically different expected response. The estimate is pulled inward. This is bias, present at every sample size, and it grows with the neighbourhood.

Which is why extrapolation is not merely unreliable but unavailable. Beyond the observed range there are no neighbours at all. A k -NN fit returns the same k boundary observations for every query out there, so it predicts a constant, forever. A straight-line model asked the same question at least answers in a direction; the local average answers flat, and a flat answer is easily mistaken for a considered one.

Example

Where local averaging is efficient, and where it is not

Local averaging estimates E [ Y ∣ X = x ] from observations near x , without assuming a functional form. Under regularity conditions, including a smooth regression function, a bandwidth shrinking at an appropriate rate, and enough data near x , the estimate is consistent. The cases below differ in what that costs.

A smooth curve of unknown form. A response rising and then levelling off, with no theory fixing whether the curve is logarithmic, a saturating exponential, or something else. Averaging neighbours follows the shape without committing to a form. This is the case the method is designed for.

A relationship known to be linear. A local average can approximate a line, but it re-estimates the fit at every point from the observations in a neighbourhood, while least squares uses all n observations to estimate two parameters. For comparable precision the local method needs more data, and it returns no slope coefficient to report.

A discontinuity. Where the regression function has a jump, any neighbourhood spanning the threshold averages across it, so the fitted curve shows a transition whose width is roughly the neighbourhood width rather than a step. The output contains no indication that this has happened. A regression tree can place a split at the threshold, so a step function is within the class of functions a tree can represent; whether a tree fitted to a particular sample locates the threshold accurately depends on the sample and on how the tree is grown and pruned.

Several predictors. With one predictor a neighbourhood holding 10 % of the observations spans about 10 % of the range. With ten predictors, holding 10 % of the observations requires spanning about 0.10 1 / 10 ≈ 80 % of each variable's range, so the neighbourhood is no longer local. Local averaging suits problems with few predictors.

Sparse regions. Where observations thin out, a k -nearest-neighbour estimate draws on points far from x and reports an average whose stated precision does not reflect that distance, while a fixed-bandwidth kernel estimate is visibly unstable because it uses few points. The second makes the data shortage apparent in the output.

Local averaging is efficient where the regression function is smooth, its form is unknown, and the predictors are few. It is less efficient than a parametric fit when the form is known, and it is biased near discontinuities and at the edges of the observed range, in both cases without a signal in the output.

Procedure

Producing a fitted value, and choosing the smoothing

To compute a k -nearest-neighbour prediction at x 0 .

  1. Measure the distance from x 0 to every observation. With several predictors, scale them first. A neighbourhood computed from unscaled variables is decided by whichever has the largest units.
  2. Take the k smallest distances.
  3. Average their responses. That is the prediction. Ties in distance are broken arbitrarily; if a tie changes the answer noticeably, k is too small.

To compute a kernel estimate at x 0 .

  1. Weight each observation by K h ( x 0 , x i ) , falling with distance.
  2. Form the weighted average ∑ i w i y i / ∑ i w i . Dividing by the sum of weights is what makes it an average rather than a sum that shrinks toward zero far from the data.
  3. Report the effective number of observations, ( ∑ i w i ) 2 / ∑ i w i 2 . An estimate resting on two effective observations should not be presented like one resting on twenty.

To choose the smoothing parameter.

  1. Fix a grid of candidates, for k -NN, integers from 1 up to a sizeable fraction of n ; for a kernel, bandwidths spanning an order of magnitude.
  2. For each candidate, compute held-out error. Leave-one-out is natural here because refitting costs nothing: predicting an observation from its neighbours simply excludes itself.
  3. Take the minimising candidate, and look at the whole curve rather than the single winner. A flat minimum means several values are equivalent and the choice does not matter; a sharp one means it does.
  4. Never select on training error. It falls monotonically as the fit grows more local and is exactly zero at k = 1 , so its optimum is always the degenerate fit.

If the same data must also report performance, the selection is itself fitting, so the honest estimate needs an outer split that the tuning never saw.

To decide whether to use this family at all.

  • How many predictors? Beyond a few, neighbourhoods stop being local.
  • Is a functional form genuinely known from the subject matter? If so, assuming it uses the data far more efficiently.
  • Is the relationship expected to be smooth? A known threshold argues for a tree.
  • Is a coefficient needed for reporting? This family returns none.

Checks. Plot the fit against the data at the chosen setting and at neighbouring settings; the chosen one should look neither jagged nor visibly flattened at a peak. Compare the held-out error against a straight-line fit: if the line wins, the extra flexibility is not earning its cost.

Worked example

Two criteria that disagree about the best neighbourhood

Twenty observations, x = 1 , … , 20 :

y = 3.9 ,   15.0 ,   14.3 ,   17.1 ,   19.4 ,   20.2 ,   22.4 ,   22.0 ,   24.8 ,   24.4 , 26.6 ,   25.1 ,   27.3 ,   26.8 ,   27.6 ,   27.0 ,   28.4 ,   27.1 ,   28.0 ,   26.6 .

Step 1: one prediction, several neighbourhoods. At x 0 = 10 :

k Neighbours used f ^ ( 10 )
1 { 10 } 24.4000
3 { 9 , 10 , 11 } 25.2667
5 { 8 , … , 12 } 24.5800
9 { 6 , … , 14 } 24.4000
15 { 3 , … , 17 } 23.5600
20all 22.7000

The estimate drifts downward as k grows, because the neighbourhood reaches back into the region where the response is genuinely lower. That drift is bias arriving.

Step 2: training error across the whole fit.

k Training RSSRoughness
1 0.0000 170.0700
3 79.5056 36.8344
5 108.2684 22.7256
9 220.8447 11.9183
15 439.9587 4.6563

Training error is minimised at k = 1 , exactly zero, and rises monotonically with k . Roughness, the sum of squared differences between successive fitted values, runs the other way, from 170.07 down to 4.66 . By the training criterion alone the answer is always k = 1 , which is the criterion failing rather than an answer.

Step 3: leave-one-out cross-validation. Each observation is predicted from the others, so no fit sees the point it is scored on.

k LOOCV MSE k LOOCV MSE
1 14.6640 6 11.0568
2 8.9444 7 12.8085
3 8.6377 9 15.9562
4 8.4585 12 21.2341
5 10.1506 15 27.8433

Step 4: read the disagreement. The two criteria rank the candidates almost oppositely at the flexible end. k = 1 is best by training error at 0.0000 and among the worst by held-out error at 14.6640 , worse than k = 6 . The held-out criterion falls to a minimum of 8.4585 at k = 4 and climbs steadily after, giving the U-shape the bias-variance decomposition predicts.

The gap between 0.0000 and 14.6640 at the same setting is the whole argument for held-out evaluation, in one row.

Step 5: the same trade with continuous weights. A Gaussian kernel at x 0 = 10 , reporting the effective number of observations ( ∑ w i ) 2 / ∑ w i 2 :

h f ^ ( 10 ) Effective n
0.5 24.6763 1.56
1.0 24.9411 3.54
2.0 24.6323 7.09
5.0 23.5505 16.20

At h = 1.0 the estimate rests on about 3.5 observations, close to the k = 4 the cross-validation chose, and the two estimates agree to within about 0.5 . The parameters have different names and units and set the same thing.

Step 6: the boundary. With k = 5 , which observations does each query point actually use?

x 0 NeighboursLeftRight
1 1 , 2 , 3 , 4 , 5 04
2 1 , 2 , 3 , 4 , 5 13
10 8 , 9 , 10 , 11 , 12 22
19 16 , 17 , 18 , 19 , 20 31
20 16 , 17 , 18 , 19 , 20 40

At x 0 = 10 the neighbourhood is balanced and the errors from either side partly cancel. At x 0 = 1 every neighbour lies to the right, in a region where the response is higher, so the estimate is pulled up. Nothing about this improves with more data: collecting a hundred observations on the same range leaves the leftmost query point with all its neighbours still on one side.

Contrast

Settings and criteria that look comparable

Training error against held-out error, on the same fits.

k Training RSSLOOCV MSE
1 0.0000 14.6640
3 79.5056 8.6377
4— 8.4585
9 220.8447 15.9562
15 439.9587 27.8433

The two columns rank k = 1 oppositely: best by the first, close to worst by the second. Training error is monotone in k and held-out error is U-shaped, so only one of them contains the information needed to choose. The disagreement is largest exactly at the setting a reader is most tempted by.

k -nearest neighbours against a kernel, at matched flexibility.

k -NN, k = 4 Gaussian kernel, h = 1.0
effective observations 4 3.54
estimate at x 0 = 10 ≈ 24.6 24.9411
fitted curvediscontinuouscontinuous
where data are sparsereaches further, silentlyuses very few, visibly

At comparable flexibility the two agree closely on the estimate and differ in what they do when the data thin out. Neither behaviour is safer in general; the kernel's instability is at least legible.

Bandwidth against neighbourhood size.

Both control smoothing, and they respond to the data's density oppositely. k fixes the number of observations and lets the width float; h fixes the width and lets the count float. In a dense region they behave alike. In a sparse one, k -NN quietly averages distant points while the kernel returns a noisy estimate from two or three.

A local average against a regression tree.

Both are local and neither assumes a form. A tree partitions by searching for splits, so it represents a genuine threshold exactly and a smooth trend as a staircase. Local averaging assumes smoothness, so it represents a smooth trend well and a threshold as a ramp whose width is the neighbourhood's. The two fail on each other's easy cases.

A smoothing spline at large λ against ordinary least squares.

These converge: as λ → ∞ the curvature penalty forces f ″ = 0 and the spline becomes the least-squares straight line. The nonparametric family contains the parametric fit as its most heavily smoothed member, which is the precise sense in which assuming a form is the strongest smoothing available, and why the choice between them is a matter of degree rather than of kind.

Warning

The failures that do not announce themselves

A local-averaging fit returns a number at every query point, whatever is happening underneath. Three situations produce confident output from an estimate that is not supported, and none of them raises an error.

Beyond the observed range there are no neighbours. A k -NN fit asked about a point past the largest observation returns the average of the same k boundary observations, and returns it for every point further out. The prediction is flat to infinity. A straight-line model at least answers in a direction that can be argued about; the local average answers with a constant that reads like a considered estimate. Record the range the fit was estimated on and treat any query outside it as unanswered rather than answered flatly.

At the edges the neighbourhood is one-sided. With k = 5 on observations at x = 1 , … , 20 , the query point x 0 = 1 uses { 1 , 2 , 3 , 4 , 5 } , four neighbours above it and none below. If the response is rising, every one of them has a higher expected value, so the estimate is pulled up. This is bias, it is present at every sample size, and it grows with the neighbourhood. Collecting more observations over the same range does not fix it, because the leftmost point still has no neighbours to its left.

A discontinuity is rendered as a ramp. Any neighbourhood spanning a genuine step averages across it, so the fit shows a gradual transition whose width is the neighbourhood's width rather than a property of the data. The output looks like evidence of a smooth change. If a threshold is suspected, a method that splits rather than averages will represent it, and comparing the two on held-out error is the check.

---

And one that announces itself as a success. Setting k = 1 produces a training error of exactly zero. On the worked data that setting has leave-one-out error 14.6640 , worse than every k from 2 to 8, and worse than k = 6 at 11.0568 . A model reporting perfect agreement with the data it was given has told you nothing about a new observation, and the criterion that reports it is monotone in the parameter being chosen, so it cannot choose.

---

With several predictors, "local" stops meaning local. To capture a tenth of the observations in p dimensions a neighbourhood must span about 0.1 1 / p of each variable's range: 10 % at p = 1 , 63 % at p = 5 , 80 % at p = 10 . At that point the average includes observations from across the whole dataset, and the method has quietly become a global mean while still being described as local.

Application

Where the shape is the question

Growth charts. A child's height-for-age reference is a set of smooth curves through population data, and no formula relates height to age. The curves are produced by smoothing, with the bandwidth chosen so that the result follows real features of growth without following sampling noise. What is published is the curve itself, since there is no coefficient anyone would quote.

Dose-response screening. Before a parametric model is committed to, a smooth fit shows what shape the data suggest: monotone, saturating, or non-monotone. Here the nonparametric fit is a diagnostic rather than the deliverable, and it is used precisely because it will not impose the shape that is being investigated.

Environmental time series. Seasonal and long-term components are separated by local smoothing at two different bandwidths, the narrow one following the annual cycle and the wide one the trend beneath it. The bandwidths encode what counts as season and what counts as trend, which is a decision about the science rather than a technical parameter.

Spatial interpolation. Predicting a pollutant concentration between monitoring stations is a weighted average of nearby readings. The boundary problem is concrete here: a station at the edge of the network has neighbours on one side only, so its estimates are pulled inward, and the network's edge is exactly where extrapolation is most tempting.

Calibration curves in the laboratory. Where an instrument's response is monotone but not of any known form, a spline through calibration standards converts a reading into a concentration. The fitted curve is valid only across the range of the standards, and a reading outside it is reported as out of range rather than converted. The same rule the warning block gives, enforced as laboratory practice.

---

What decides the choice in each case. No formula was available and none was wanted: in the growth charts and the calibration curve because the curve is the product, and in the dose-response screen because imposing a form would prejudge the question being asked. Where a form is known from the subject matter, these methods are the wrong instrument, since they spend data discovering what was already known.

Next step

Practice Fitting Without a Formula

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.