Subject

Statistical Learning

Building models that predict, and estimating how well they will. Starts from the gap between the error a model reports on the data it was fitted to and the error it will make on data it has not seen, since every later judgement, which model to prefer, how much flexibility to allow, whether a reported improvement is real, depends on estimating the second rather than the first.

Start Statistical Learning

  1. Statistical Learning

    Estimating how a predictive model will perform on data it was not fitted to, and what a reported error figure does and does not support.

  • Say what a fitted model will do on data it has not seen

    Estimate a predictive model's out-of-sample error by a resampling scheme the data structure permits, attribute the remaining error to bias, variance or irreducible noise, and state what the reported figure supports.

  • Match a model to the quantity being predicted

    Choose a regression form the response variable and the shape of the relationship permit, fit it, and report its coefficients on the scale they belong to.

  • Build a regression tree and account for variance reduction

    Construct a regression tree by recursive binary splitting, read the piecewise-constant fit it produces, and state how averaging an ensemble of trees affects the variance of a prediction.

  • Fit a curve without assuming its shape

    Produce a fitted value by local averaging, set the neighbourhood size or bandwidth, and state what assuming no functional form gives up.

  • Partition unlabelled data and account for the criterion

    Partition unlabelled data by k-means, and account for its dependence on initialisation and for the cluster shapes its objective favours.

  • Accept bias deliberately, and state what it gains

    Recognise when a least-squares fit is unstable because predictors carry overlapping information, apply ridge or lasso shrinkage, and select the penalty strength using held-out error.

  • Find co-occurrence, and tell it from association

    Mine frequent itemsets from transaction data, compute support, confidence and lift, and judge which rules describe an association rather than a frequency.

  • Construct a linear discriminant and test its assumption

    Construct a linear decision boundary from class means and pooled within-class scatter, and determine whether the shared-covariance assumption holds.

See detailed outcomes

Say what a fitted model will do on data it has not seen

  • The learner can estimate a predictive model's performance on unseen data by an appropriate resampling scheme, attribute the remaining error to bias, variance or irreducible noise, and state what a reported error figure does and does not license.

Match a model to the quantity being predicted

  • The learner can choose a regression form appropriate to the response variable and the shape of the relationship, fit and interpret polynomial, exponential and Poisson models, and say what each assumes about the response that a linear model does not.

Build a regression tree and account for variance reduction

  • The learner can construct a regression tree by recursive binary splitting, read the prediction it makes for a given input, explain why a single deep tree has high variance, and say what bagging and random forests each change about that variance and at what cost.

Fit a curve without assuming its shape

  • The learner can produce a fitted value from a nearest-neighbour or kernel-weighted local average, explain how the neighbourhood size or bandwidth moves bias against variance, choose that parameter by held-out error, and say what a nonparametric fit gives up in exchange for assuming no functional form.

Partition unlabelled data and account for the criterion

  • Given a small dataset and an initialisation, the learner can carry out k-means assignment and update steps to convergence, and can recognise that the returned partition is locally optimal.
  • Given a clustering result, the learner can read a within-cluster sum of squares curve without inferring a cluster count the data does not carry, and can attribute a failure to the objective's notion of a cluster rather than to the search.

Accept bias deliberately, and state what it gains

  • The learner can explain why least squares becomes unstable when predictors are nearly collinear, compute ridge and lasso solutions on a small design, state what each penalty does differently to a coefficient that duplicates another, and justify a penalty strength by held-out error rather than by the fit it produces on the data it was estimated from.

Find co-occurrence, and tell it from association

  • The learner can compute support, confidence and lift for a rule over a small transaction set, explain why the anti-monotone property lets Apriori skip most of the candidate lattice, and decide whether a rule with high confidence describes an association at all, recognising that confidence above a threshold is compatible with the antecedent making the consequent less likely.

Construct a linear discriminant and test its assumption

  • The learner can compute a linear discriminant direction from class means and the pooled within-class scatter, explain why that direction generally differs from the line joining the class means, classify observations by projecting onto it, and identify when the equal-covariance assumption fails badly enough that a single linear boundary cannot separate the classes at all.

Browse Statistical Learning reference · Practice Statistical Learning

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.