Linear Discriminants, and the Covariance They Assume

Why the best direction for separating two labelled classes is generally not the line joining their means, how dividing by the pooled within-class scatter rotates it toward directions in which the classes are internally tight, and what happens when the single-covariance assumption behind that pooling fails. A case where the method returns a boundary, reports nothing amiss, and classifies barely better than chance.

Definition

Linear discriminant analysis separates labelled classes by projecting the predictors onto a single direction and thresholding the result.

For two classes with means x ¯ 0 and x ¯ 1 , let each class's scatter about its own mean be

S k = ∑ i ∈ C k ( x i − x ¯ k ) ( x i − x ¯ k ) T ,

and let S w = S 0 + S 1 be the pooled within-class scatter. The discriminant direction is

w = S w − 1 ( x ¯ 1 − x ¯ 0 ) .

An observation is assigned by computing w T x and comparing it against a threshold, conventionally the midpoint between the two projected class means when the classes are equally likely and equally costly to confuse.

Why the inverse is there. w maximises the Fisher criterion, the ratio of squared between-class separation to within-class spread along the projection. Without S w − 1 the direction would be x ¯ 1 − x ¯ 0 , which maximises separation alone and ignores how far the classes spread along it.

The assumption. Pooling the two scatters into one matrix asserts that both classes share a common covariance. Under that assumption, and with normal classes, the optimal boundary is linear and this rule attains it. Where the covariances differ, the pooled matrix describes neither class, and allowing each class its own covariance gives a quadratic boundary instead.

What the method does not report. It returns a direction and a threshold for any input whatever. Nothing in the output indicates that the classes overlap, that their means nearly coincide, or that the shared-covariance assumption has failed.

Assumptions and scope

  • Pooling the class scatters assumes the classes share a covariance and differ only in mean. Under that assumption with normal classes the optimal boundary is linear; where it fails, a per-class covariance rule gives a quadratic boundary and the linear one may be far worse.

  • The midpoint threshold between projected class means assumes the classes are equally likely a priori and that the two kinds of error cost the same. Unequal prior probabilities or unequal costs shift the threshold, and the shift is a decision rather than a computation.

  • S w must be invertible. With more predictors than observations, or with a predictor that is an exact combination of others, it is singular and the direction is undefined. This is the same near-singularity that motivates a penalty in regression.

  • The direction is determined only up to scale: multiplying w by any nonzero constant rescales every projection and leaves the classification unchanged, so a reported direction that differs from a reference by a positive factor may still be correct.

  • Classification error measured on the same observations used to compute the means and scatter is optimistic, for the reason any training error is. The figures in this unit are stated as training errors on small datasets and are used to compare directions on identical data, not to estimate performance.

  • The figures come from exact rational arithmetic on the two stated datasets, with the discriminant direction obtained by inverting the pooled scatter symbolically and the error counts from projecting every observation.

Worked material

Example

Four configurations and what the discriminant does with each

Elongated and correlated — the rotation case. The unit's twelve-point dataset: within-class correlation 0.897150 , difference of means at 63.4349 degrees, discriminant at 127.1467 degrees. Rotation of 63.7117 degrees, 0 errors against 1 , separation ratio 2.043788 better.

This is where the inverse earns its place. The classes are stretched along the same diagonal, so the direction joining their centres is precisely the direction in which each class is most spread out.

Round and equal — the case where the naive answer is right. If both classes have the same spread in every direction and no internal correlation, S w is a multiple of the identity. Its inverse is then also a multiple of the identity, and

w = S w − 1 ( x ¯ 1 − x ¯ 0 ) ∝ x ¯ 1 − x ¯ 0 .

The discriminant points exactly along the line joining the centres, and the rotation is zero. Worth knowing, because it explains why the naive answer feels correct: it is correct, for round classes, and most textbook diagrams draw round classes.

A ring around a cloud — no linear boundary exists. Class 0 is nine points clustered at the origin; class 1 is nine points on a ring around it. The class means are ( 0 , 0 ) and ( 0.7 7 ― , 0.1 1 ― ) , nearly the same point, so the difference of means is almost nothing.

The scatter traces are 24 and 3496 9 , a ratio of 16.1852 : the classes differ enormously in spread, and the shared-covariance assumption is false.

The method returns a direction anyway and misclassifies 7 of 18 , or 38.89 % , against a chance rate of 50 % for two balanced classes. It is barely better than guessing, and no threshold on that projection does better, because a ring cannot be cut from its own centre by a straight line.

The same ring, with per-class covariances. Assign each observation to the class under whose own spread it is least surprising. A quadratic rather than linear boundary. On the identical eighteen points this misclassifies 0 .

The data was always separable. What was not available was a straight boundary.

---

The first two bracket the linear method: it improves substantially on the naive direction when classes are elongated, and reduces to that direction when they are round. The third shows the method failing while giving every appearance of having worked. A direction, a threshold, a classification, and no indication that the assumption behind them was false. The fourth shows the failure was in the shape of boundary permitted, not in the fitting or the data.

Contrast

Pairs that differ in one respect

The discriminant direction against the difference of means, same data.

S w − 1 ( x ¯ 1 − x ¯ 0 ) x ¯ 1 − x ¯ 0
direction ( − 0.568182 , 0.750000 ) ( 2 , 4 )
angle 127.1467 ∘ 63.4349 ∘
gap between projected means 1.863636 20
Fisher criterion 1.863636 0.911854
errors 0 of 12 1 of 12

The naive direction produces a gap more than ten times larger and separates worse. That single row is the argument for the inverse: a gap is not a measure of separation unless the spread beside it is known.

A rotation that matters against one that does not.

When the within-class correlation is 0.897150 , the inverse rotates the direction by 63.7117 degrees. When the classes are round, S w a multiple of the identity, it rotates by nothing at all, and the discriminant coincides with the difference of means.

So the correction's size is a property of the data rather than of the method, and inspecting S w tells you in advance whether it will matter.

A failure of fitting against a failure of form.

On the ring dataset the linear rule misclassifies 7 of 18 . The temptation is to look for a fitting error or to seek more data. Neither would help: the class means nearly coincide, so the construction's numerator carries almost no information, and no straight line separates a ring from its centre.

The per-class covariance rule misclassifies 0 on the identical points. The data was separable throughout; the linear boundary was the constraint.

Pooled covariance against per-class covariance.

Pooling assumes the classes share a shape and gives a linear boundary, fewer parameters, more stable on small samples, optimal when the assumption holds. Allowing each class its own covariance gives a quadratic boundary, more flexible, and on the ring data the difference between 7 errors and 0 .

The cost is parameters: two covariance matrices instead of one, estimated from the observations of a single class each. With few observations per class the extra flexibility can hurt, which is why the pooled version is not simply obsolete.

A direction against its negative.

Multiplying w by − 1 produces the same boundary with the class labels exchanged; multiplying by any positive constant changes nothing at all. So a reported direction that disagrees with a reference by a positive factor is the same answer, and one disagreeing by sign is the same boundary described from the other side. Only the line matters, not the arrow.

Training error against what it estimates.

Every error count in this unit was computed on the observations that produced the means and the scatter. The 0 of 12 is a training figure and is optimistic, as any training error is. It is used here to compare two directions on identical data, which it supports, and not to estimate how the rule would perform on new observations, which it does not.

Common errors

Common misconception

That the best direction for separating two classes is the line joining their means, so the pooled scatter matrix is a normalising detail that does not change the answer. It changes the answer, and often drastically. On the two-class dataset of this unit the class means are ( 3.5 , 3.1 6 ― ) and ( 5.5 , 7.1 6 ― ) , so their difference is ( 2 , 4 ) , pointing at 63.4349 degrees from the x -axis. The pooled within-class scatter is S w = ( 11 11 11 41 / 3 ) , whose off-diagonal entry reflects a within-class correlation of 0.897150 , and the discriminant direction S w − 1 ( x ¯ 1 − x ¯ 0 ) is ( − 25 / 44 , 3 / 4 ) ≈ ( − 0.568182 , 0.750000 ) , pointing at 127.1467 degrees. A difference of 63.7117 degrees, with a negative first component that no inspection of the class means would suggest. The consequence is measurable: projecting onto the discriminant direction separates the twelve observations with 0 errors, while projecting onto the difference of means misclassifies 1 , and the Fisher separation ratio is 1.863636 against 0.911854 , better by a factor of 2.043788 . The inverse is doing the work: it discounts separation measured along directions in which the classes are internally spread out, and when the classes are elongated and correlated, that discounting rotates the answer.

Common misconception

That a poor classification rate means the classifier was fitted badly, so more data or a better-chosen threshold would fix it. Some class configurations admit no linear boundary at all, and the discriminant will return a direction regardless. It has no way to report that no linear boundary separates the classes. On the second dataset of this unit, class 0 is a compact cloud at the origin and class 1 is a ring surrounding it. Their means nearly coincide, at ( 0 , 0 ) and ( 0.7 7 ― , 0.1 1 ― ) , so the difference of means carries almost no information, and the class scatters differ by a factor of 16.1852 in trace, violating the equal-covariance assumption outright. Linear discriminant analysis still produces a direction and still classifies: it misclassifies 7 of 18 observations, or 38.89 % , against a chance level of 50 % for two balanced classes. No threshold on that projection does better, because the classes are not linearly separable in any direction. A rule allowing each class its own covariance, assigning each point to the class whose own spread makes it least surprising, misclassifies 0 of 18 on the same data. The failure is therefore in the shape of the boundary the method is permitted to draw, not in the fitting, and no amount of additional data from the same process would change it.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.