Course
Statistical Learning
Estimating how a predictive model will perform on data it was not fitted to, and what a reported error figure does and does not support.
Finishing this course means you have demonstrated the required skills with the level of support this course currently assesses.
- Modules
- 8
- Lessons
- 8
- Skills
- 9
- Starting here
- Assumes 3 prior topics
The route
Module 1: Evaluation
The gap between training error and test error, the resampling that estimates the second, the three sources the remaining error comes from, and the two situations where the standard procedure still reports a figure that is too good: observations ordered in time, and data reused to choose the model.
- The learner can estimate a predictive model's performance on unseen data by an appropriate resampling scheme, attribute the remaining error to bias, variance or irreducible noise, and state what a reported error figure does and does not license.
Module 2: Model form
What the response variable and the shape of the relationship each rule out before any curve is drawn, why a polynomial fit is still a linear model, and why a coefficient fitted on the log scale has to be converted before it can be reported.
- The learner can choose a regression form appropriate to the response variable and the shape of the relationship, fit and interpret polynomial, exponential and Poisson models, and say what each assumes about the response that a linear model does not.
Module 3: Trees and ensembles
A model that fixes no functional form: how a tree chooses splits by exhaustive search, why its structure is unstable under resampling, and how averaging many trees reduces variance, including the correlation floor that limits the reduction.
- The learner can construct a regression tree by recursive binary splitting, read the prediction it makes for a given input, explain why a single deep tree has high variance, and say what bagging and random forests each change about that variance and at what cost.
Module 4: Local averaging
Estimating a response by averaging nearby observations instead of fitting a shape, the one parameter that decides how near is near, why the training data cannot choose it, and why the method is weakest at the edges of the range and unavailable beyond them.
- The learner can produce a fitted value from a nearest-neighbour or kernel-weighted local average, explain how the neighbourhood size or bandwidth moves bias against variance, choose that parameter by held-out error, and say what a nonparametric fit gives up in exchange for assuming no functional form.
Module 5: Clustering
Partitioning data with no response variable and no held-out error. What convergence guarantees, why a decreasing within-cluster sum of squares does not identify the number of clusters, and when the clustering objective does not match the structure in the data.
- Given a small dataset and an initialisation, the learner can carry out k-means assignment and update steps to convergence, and can recognise that the returned partition is locally optimal.
- Given a clustering result, the learner can read a within-cluster sum of squares curve without inferring a cluster count the data does not carry, and can attribute a failure to the objective's notion of a cluster rather than to the search.
Module 6: Shrinkage penalties
Estimation when predictors carry overlapping information and least-squares coefficients are unstable. Adding a penalty to the estimator, the difference between shrinking a coefficient and setting it to zero, and why the penalty strength requires held-out error.
- The learner can explain why least squares becomes unstable when predictors are nearly collinear, compute ridge and lasso solutions on a small design, state what each penalty does differently to a coefficient that duplicates another, and justify a penalty strength by held-out error rather than by the fit it produces on the data it was estimated from.
Module 7: Association rules
Mining co-occurrence statements from transaction data. The three rule measures and what each is a proportion of, the anti-monotone property that prunes the candidate search, and the judgements the algorithm does not make.
- The learner can compute support, confidence and lift for a rule over a small transaction set, explain why the anti-monotone property lets Apriori skip most of the candidate lattice, and decide whether a rule with high confidence describes an association at all, recognising that confidence above a threshold is compatible with the antecedent making the consequent less likely.
Module 8: Linear discriminant analysis
Constructing a decision boundary for labelled classes. Why the discriminant direction differs from the line joining the class means, and the configurations no linear boundary separates.
- The learner can compute a linear discriminant direction from class means and the pooled within-class scatter, explain why that direction generally differs from the line joining the class means, classify observations by projecting onto it, and identify when the equal-covariance assumption fails badly enough that a single linear boundary cannot separate the classes at all.
What finishing means
Finishing this course means you have demonstrated the required skills with the level of support this course currently assesses.
9 required skills. If you reach a lesson without the background it assumes, you are pointed at the prerequisite first, and returned here afterwards.
How progress is measured
Progress is inferred from evidence you produce, not from pages you have opened. Each required skill moves through states as evidence accumulates: met, practicing with help, performed unassisted, then performed again after a delay.
This course counts a skill as finished atguided. Where the system cannot admit evidence for a stronger claim — for instance when the only available scoring is your own judgment of your written answer — the skill stays at the state the evidence supports, and the reason is shown rather than hidden.