Subject
Statistical Learning
Building models that predict, and estimating how well they will. Starts from the gap between the error a model reports on the data it was fitted to and the error it will make on data it has not seen, since every later judgement, which model to prefer, how much flexibility to allow, whether a reported improvement is real, depends on estimating the second rather than the first.
Learning paths
Statistical Learning
Estimating how a predictive model will perform on data it was not fitted to, and what a reported error figure does and does not support.
What this subject develops
Say what a fitted model will do on data it has not seen
Estimate a predictive model's out-of-sample error by a resampling scheme the data structure permits, attribute the remaining error to bias, variance or irreducible noise, and state what the reported figure supports.
Match a model to the quantity being predicted
Choose a regression form the response variable and the shape of the relationship permit, fit it, and report its coefficients on the scale they belong to.
Build a regression tree and account for variance reduction
Construct a regression tree by recursive binary splitting, read the piecewise-constant fit it produces, and state how averaging an ensemble of trees affects the variance of a prediction.
Fit a curve without assuming its shape
Produce a fitted value by local averaging, set the neighbourhood size or bandwidth, and state what assuming no functional form gives up.
Partition unlabelled data and account for the criterion
Partition unlabelled data by k-means, and account for its dependence on initialisation and for the cluster shapes its objective favours.
Accept bias deliberately, and state what it gains
Recognise when a least-squares fit is unstable because predictors carry overlapping information, apply ridge or lasso shrinkage, and select the penalty strength using held-out error.
Find co-occurrence, and tell it from association
Mine frequent itemsets from transaction data, compute support, confidence and lift, and judge which rules describe an association rather than a frequency.
Construct a linear discriminant and test its assumption
Construct a linear decision boundary from class means and pooled within-class scatter, and determine whether the shared-covariance assumption holds.
See detailed outcomes
Say what a fitted model will do on data it has not seen
- The learner can estimate a predictive model's performance on unseen data by an appropriate resampling scheme, attribute the remaining error to bias, variance or irreducible noise, and state what a reported error figure does and does not license.
Match a model to the quantity being predicted
- The learner can choose a regression form appropriate to the response variable and the shape of the relationship, fit and interpret polynomial, exponential and Poisson models, and say what each assumes about the response that a linear model does not.
Build a regression tree and account for variance reduction
- The learner can construct a regression tree by recursive binary splitting, read the prediction it makes for a given input, explain why a single deep tree has high variance, and say what bagging and random forests each change about that variance and at what cost.
Fit a curve without assuming its shape
- The learner can produce a fitted value from a nearest-neighbour or kernel-weighted local average, explain how the neighbourhood size or bandwidth moves bias against variance, choose that parameter by held-out error, and say what a nonparametric fit gives up in exchange for assuming no functional form.
Partition unlabelled data and account for the criterion
- Given a small dataset and an initialisation, the learner can carry out k-means assignment and update steps to convergence, and can recognise that the returned partition is locally optimal.
- Given a clustering result, the learner can read a within-cluster sum of squares curve without inferring a cluster count the data does not carry, and can attribute a failure to the objective's notion of a cluster rather than to the search.
Accept bias deliberately, and state what it gains
- The learner can explain why least squares becomes unstable when predictors are nearly collinear, compute ridge and lasso solutions on a small design, state what each penalty does differently to a coefficient that duplicates another, and justify a penalty strength by held-out error rather than by the fit it produces on the data it was estimated from.
Find co-occurrence, and tell it from association
- The learner can compute support, confidence and lift for a rule over a small transaction set, explain why the anti-monotone property lets Apriori skip most of the candidate lattice, and decide whether a rule with high confidence describes an association at all, recognising that confidence above a threshold is compatible with the antecedent making the consequent less likely.
Construct a linear discriminant and test its assumption
- The learner can compute a linear discriminant direction from class means and the pooled within-class scatter, explain why that direction generally differs from the line joining the class means, classify observations by projecting onto it, and identify when the equal-covariance assumption fails badly enough that a single linear boundary cannot separate the classes at all.
Browse Statistical Learning reference · Practice Statistical Learning