Practice: Fitting Without a Formula

Direct application

Observations are recorded at x = 1 , … , 20 . The responses at x = 9 , 10 and 11 are 24.8 , 24.4 and 26.6 ; the neighbouring values at x = 8 and x = 12 are 22.0 and 25.1 . Give the 3 -nearest-neighbour prediction at x 0 = 10 , to four decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

The three closest observations to x 0 = 10 are at x = 9 , 10 and 11 .

Hint 2: Next step

Average their responses with equal weight.

Prediction

A k -nearest-neighbour fit with k = 1 is computed on twenty observations, and its predictions are then compared against those same twenty observations. What is the resulting training residual sum of squares?

Enter the value. It is checked against the answer and the precision this task asks for.

1 hint available, least help first.

Hint 1: Retrieval cue

For a training observation, which observation is its own nearest neighbour?

Method selection

A k -nearest-neighbour fit is evaluated two ways on the same twenty observations:

k Training RSSLeave-one-out MSE
1 0.0000 14.6640
3 79.5056 8.6377
4 103.1 8.4585
9 220.8447 15.9562
15 439.9587 27.8433

Which value of k should be used, and on what grounds?

Error diagnosis

Observations sit at x = 1 , … , 20 and the response rises with x . A 5 -nearest-neighbour fit uses { 1 , 2 , 3 , 4 , 5 } at the query point x 0 = 1 and { 8 , 9 , 10 , 11 , 12 } at x 0 = 10 . The fit is noticeably too high at x 0 = 1 . What explains this, and would collecting more observations over the same range fix it?

Interpretation

A Gaussian kernel fit reports, at one query point, an estimate of 24.6763 with an effective sample size of 1.56 ; at another, an estimate of 23.5505 with an effective sample size of 16.20 . Both are printed to four decimal places. What should a reader take from this?

Transfer · Evaluation

A service predicts a user's rating of an item by averaging the ratings of the twenty most similar users. For items with many ratings the predictions are good. For newly listed items, rated by only a handful of users, every prediction comes out close to the same middling value, and the service reports these with the same confidence as the rest. Which account identifies the mechanism?

Construction · Evaluation · Explanation

A laboratory has 200 calibration readings relating an instrument's raw output to a known concentration, covering concentrations from 5 to 90 units. The relationship is monotone and smooth but matches no formula anyone has proposed. A colleague fits a k -nearest-neighbour curve, selects k by minimising training error, obtains k = 1 , and proposes using the curve to convert future readings, including readings that imply concentrations of around 120 units, outside the calibrated range.

Write an assessment. Address all of the following.

  1. The fitted value. Describe how a prediction is produced at a query concentration, for both k -nearest neighbours and a kernel fit, and say what the kernel reports that the nearest-neighbour version does not.
  2. The colleague's selection. State what k = 1 does to the training error and why that criterion produced this answer. Say what the resulting curve would do if the 200 readings were collected again.
  3. Choosing k properly. Give the criterion you would use, and say what shape you expect the resulting curve of error against k to have and why.
  4. The ends of the range. Say what happens to the fit near 5 and near 90 units, whether collecting more readings between 5 and 90 would fix it, and why.
  5. The reading at 120 units. Say what the fitted curve will return there and what should be reported instead.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

For part 2, ask which observation is nearest to a training observation.

Hint 2: Concept cue

For part 4, ask on which side of the lowest query point its neighbours can possibly lie.

Hint 3: Strategy cue

For part 5, work out which observations the neighbourhood contains for a query far above every reading, and then for one further still.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

1. The fitted value. For k -nearest neighbours: measure the distance from the query point to every calibration reading, take the k smallest, and average their responses. Weights are equal inside that neighbourhood and zero outside it, so the fitted curve jumps as the query point moves and an observation enters or leaves. For a kernel fit: weight every reading by K h , a function falling with distance, and form ∑ i w i y i / ∑ i w i . Weights never reach zero, so the curve is continuous. The kernel version also reports the effective number of observations behind each estimate, ( ∑ i w i ) 2 / ∑ i w i 2 . That is the quantity the nearest-neighbour fit conceals: it always supplies exactly k neighbours, whether they are close by or scattered across the range. 2. The colleague's selection. At k = 1 every training reading is its own nearest neighbour, so each prediction equals the observed value and the training residual sum of squares is exactly zero. No other k can match that, so minimising training error selects k = 1 every time, on any dataset. The criterion is monotone in the parameter being chosen and therefore cannot choose it. What the curve has done is record the 200 readings, including their measurement noise. Collect the 200 readings again and the entire curve moves, because each fitted value is a single noisy measurement rather than an estimate combining several. That is the point of maximum variance, reached by the criterion that reports perfection. 3. Choosing k properly. Cross-validated error, and leave-one-out is natural here because excluding an observation from its own neighbourhood costs nothing. Compute the held-out error across a grid of k and take the minimum. The curve should be U-shaped. At small k the variance term dominates and the error is high; as k grows the noise is averaged down and the error falls; past some point the neighbourhood reaches into regions of genuinely different concentration, bias rises, and the error climbs again. I would also look at how flat the minimum is: a flat one means several values are equivalent and the choice does not matter much, a sharp one means it does. 4. The ends of the range. Near 5 units every neighbour lies above the query point, and near 90 every neighbour lies below. Since the relationship is monotone, those one-sided neighbourhoods have systematically different expected responses, so the fit is pulled toward the interior at both ends, too high at the bottom and too low at the top. More readings between 5 and 90 would not fix it. The lowest query point still has no readings below it however many are collected, so the neighbourhood remains one-sided. This is bias with a fixed direction, not noise that averages away. What would help is calibration standards below 5 and above 90, which changes the range rather than the density. 5. The reading at 120 units. There are no readings above 90, so the k nearest to any query beyond it are the same k readings at the top of the calibrated range. The fit returns their average, the same constant for 120, for 200, for any value out there, and returns it with no indication that the query is outside the data. That value should not be reported as a concentration. The reading is out of calibrated range and should be reported as such, exactly as a laboratory reports an out-of-range result rather than converting it. This is the cost of assuming no functional form: a parametric calibration curve would at least extrapolate in a direction that could be argued about from the instrument's physics, while local averaging has nothing to extrapolate from. Extending the range needs standards at higher concentrations, not a different fit.

A complete answer does each of these:

  • computes local average
  • relates smoothing to error
  • selects parameter by holdout
  • identifies boundary degradation
  • states parametric tradeoff
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.