Linear Discriminants, and the Covariance They Assume
Why the best direction for separating two labelled classes is generally not the line joining their means, how dividing by the pooled within-class scatter rotates it toward directions in which the classes are internally tight, and what happens when the single-covariance assumption behind that pooling fails. A case where the method returns a boundary, reports nothing amiss, and classifies barely better than chance.
Definition
Linear discriminant analysis separates labelled classes by projecting the predictors onto a single direction and thresholding the result.
For two classes with means
and let
An observation is assigned by computing
Why the inverse is there.
The assumption. Pooling the two scatters into one matrix asserts that both classes share a common covariance. Under that assumption, and with normal classes, the optimal boundary is linear and this rule attains it. Where the covariances differ, the pooled matrix describes neither class, and allowing each class its own covariance gives a quadratic boundary instead.
What the method does not report. It returns a direction and a threshold for any input whatever. Nothing in the output indicates that the classes overlap, that their means nearly coincide, or that the shared-covariance assumption has failed.
Assumptions and scope
Pooling the class scatters assumes the classes share a covariance and differ only in mean. Under that assumption with normal classes the optimal boundary is linear; where it fails, a per-class covariance rule gives a quadratic boundary and the linear one may be far worse.
The midpoint threshold between projected class means assumes the classes are equally likely a priori and that the two kinds of error cost the same. Unequal prior probabilities or unequal costs shift the threshold, and the shift is a decision rather than a computation.
must be invertible. With more predictors than observations, or with a predictor that is an exact combination of others, it is singular and the direction is undefined. This is the same near-singularity that motivates a penalty in regression. The direction is determined only up to scale: multiplying
by any nonzero constant rescales every projection and leaves the classification unchanged, so a reported direction that differs from a reference by a positive factor may still be correct. Classification error measured on the same observations used to compute the means and scatter is optimistic, for the reason any training error is. The figures in this unit are stated as training errors on small datasets and are used to compare directions on identical data, not to estimate performance.
The figures come from exact rational arithmetic on the two stated datasets, with the discriminant direction obtained by inverting the pooled scatter symbolically and the error counts from projecting every observation.
Worked material
Example
Four configurations and what the discriminant does with each
Elongated and correlated — the rotation case. The unit's twelve-point dataset: within-class correlation
This is where the inverse earns its place. The classes are stretched along the same diagonal, so the direction joining their centres is precisely the direction in which each class is most spread out.
Round and equal — the case where the naive answer is right. If both classes have the same spread in every direction and no internal correlation,
The discriminant points exactly along the line joining the centres, and the rotation is zero. Worth knowing, because it explains why the naive answer feels correct: it is correct, for round classes, and most textbook diagrams draw round classes.
A ring around a cloud — no linear boundary exists. Class
The scatter traces are
The method returns a direction anyway and misclassifies
The same ring, with per-class covariances. Assign each observation to the class under whose own spread it is least surprising. A quadratic rather than linear boundary. On the identical eighteen points this misclassifies
The data was always separable. What was not available was a straight boundary.
---
The first two bracket the linear method: it improves substantially on the naive direction when classes are elongated, and reduces to that direction when they are round. The third shows the method failing while giving every appearance of having worked. A direction, a threshold, a classification, and no indication that the assumption behind them was false. The fourth shows the failure was in the shape of boundary permitted, not in the fitting or the data.
Contrast
Pairs that differ in one respect
The discriminant direction against the difference of means, same data.
| direction | ||
| angle | ||
| gap between projected means | ||
| Fisher criterion | ||
| errors |
The naive direction produces a gap more than ten times larger and separates worse. That single row is the argument for the inverse: a gap is not a measure of separation unless the spread beside it is known.
A rotation that matters against one that does not.
When the within-class correlation is
So the correction's size is a property of the data rather than of the method, and inspecting
A failure of fitting against a failure of form.
On the ring dataset the linear rule misclassifies
The per-class covariance rule misclassifies
Pooled covariance against per-class covariance.
Pooling assumes the classes share a shape and gives a linear boundary, fewer parameters, more stable on small samples, optimal when the assumption holds. Allowing each class its own covariance gives a quadratic boundary, more flexible, and on the ring data the difference between
The cost is parameters: two covariance matrices instead of one, estimated from the observations of a single class each. With few observations per class the extra flexibility can hurt, which is why the pooled version is not simply obsolete.
A direction against its negative.
Multiplying
Training error against what it estimates.
Every error count in this unit was computed on the observations that produced the means and the scatter. The
Common errors
Common misconception
That the best direction for separating two classes is the line joining their means, so the pooled scatter matrix is a normalising detail that does not change the answer. It changes the answer, and often drastically. On the two-class dataset of this unit the class means are
Common misconception
That a poor classification rate means the classifier was fitted badly, so more data or a better-chosen threshold would fix it. Some class configurations admit no linear boundary at all, and the discriminant will return a direction regardless. It has no way to report that no linear boundary separates the classes. On the second dataset of this unit, class 0 is a compact cloud at the origin and class 1 is a ring surrounding it. Their means nearly coincide, at
Related units
Requires
Connected
- Binary Outcome Models for Experimental Research (contrasts with)
- Clustering, and What the Objective Assumes (contrasts with)
- Shrinkage, and the Trade It Makes (related)