Module 8 of 8 · Lesson 1 of 1
Linear Discriminants, and the Covariance They Assume
Linear discriminant analysis and the equal-covariance assumption.
What you will be able to do
The learner can compute a linear discriminant direction from class means and the pooled within-class scatter, explain why that direction generally differs from the line joining the class means, classify observations by projecting onto it, and identify when the equal-covariance assumption fails badly enough that a single linear boundary cannot separate the classes at all.
Orientation
The direction that points the wrong way
Twelve observations, two classes, two measurements each. Class
To separate them, aim at the line joining the centres. That difference is
The direction that actually separates these classes is
The reason is that both classes are stretched along the same diagonal, with a within-class correlation of
The arithmetic confirms it. Projecting onto the discriminant direction classifies all twelve observations correctly; projecting onto the difference of means misclassifies one. The Fisher separation ratio is
Getting there takes one matrix inverse: divide the difference of means by the pooled within-class scatter, and the direction rotates itself.
And then the harder question. That pooling assumes both classes have the same shape and differ only in position. Where the assumption holds, the resulting straight boundary is the best one available. Where it fails, the method gives no sign: it returns a direction, applies a threshold, and reports a classification exactly as it does when everything is fine.
This unit covers the construction, why the inverse rotates the answer, and a dataset where one class surrounds the other, their means nearly coincident, their spreads differing by a factor of
Definition
What each factor in the direction contributes
The canonical statement above gives
The two factors answer different questions.
| Factor | Supplies | Ignores |
|---|---|---|
| where the classes differ | how wide each class is | |
| how wide each class is, per direction | where the classes differ |
Each alone is insufficient, and multiplying them is what produces separation relative to spread rather than separation alone.
Why scatter and not covariance.
The Fisher criterion is what is being maximised. For a direction
the squared separation of the projected means over the within-class spread along
Scale invariance, and what it means for an answer. Replacing
What the threshold encodes. The midpoint between projected class means is the right choice only when the two classes are equally likely in the population and the two kinds of error cost the same. Both are assumptions about the world rather than facts about the data, and relaxing either moves the threshold, which is a decision to be recorded, not a computation to be performed.
Where the construction breaks down.
Intuition
Separation is a ratio, not a gap
Imagine two long thin clouds of points, both tilted the same way, one offset from the other along that same tilt. Their centres are far apart. Aim at the line joining the centres and project: the projected centres are indeed far apart, and each cloud, being long in exactly that direction, smears across a wide interval. The two intervals overlap, and points near the facing ends land in the wrong one.
Now aim across the tilt instead. The projected centres are closer together, which sounds worse. But each cloud is thin in this direction, so each projects to a narrow interval, and the two intervals do not touch.
A smaller gap and a much smaller spread can separate better than a large gap and a huge spread. That is the entire idea: the gap relative to the spread is what counts, and any method optimising the gap alone will choose badly whenever the classes are elongated.
How the inverse expresses this.
When the classes have no internal correlation and equal spread in every direction,
On this unit's data the correction is drastic. The within-class correlation is
Now the assumption. Adding the two scatters into one matrix says the classes have the same shape and differ only in where they sit. When that is true, a straight boundary is genuinely the best available, and this rule finds it.
When it is false,
The method does not notice. It inverts, multiplies, thresholds and reports, misclassifying
Example
Four configurations and what the discriminant does with each
Elongated and correlated — the rotation case. The unit's twelve-point dataset: within-class correlation
This is where the inverse earns its place. The classes are stretched along the same diagonal, so the direction joining their centres is precisely the direction in which each class is most spread out.
Round and equal — the case where the naive answer is right. If both classes have the same spread in every direction and no internal correlation,
The discriminant points exactly along the line joining the centres, and the rotation is zero. Worth knowing, because it explains why the naive answer feels correct: it is correct, for round classes, and most textbook diagrams draw round classes.
A ring around a cloud — no linear boundary exists. Class
The scatter traces are
The method returns a direction anyway and misclassifies
The same ring, with per-class covariances. Assign each observation to the class under whose own spread it is least surprising. A quadratic rather than linear boundary. On the identical eighteen points this misclassifies
The data was always separable. What was not available was a straight boundary.
---
The first two bracket the linear method: it improves substantially on the naive direction when classes are elongated, and reduces to that direction when they are round. The third shows the method failing while giving every appearance of having worked. A direction, a threshold, a classification, and no indication that the assumption behind them was false. The fourth shows the failure was in the shape of boundary permitted, not in the fitting or the data.
Procedure
Constructing the discriminant, and testing what it assumed
To compute the direction.
- Compute each class mean separately. Not the overall mean. The construction needs the two centres and the difference between them.
- Compute each class's scatter about its own mean,
. Deviations are taken from the class mean, not the grand mean; using the grand mean would fold the between-class separation into the within-class term and defeat the whole construction. - Add them:
. - Read the off-diagonal before going further. Divide it by
to get the within-class correlation. A value near zero means the naive direction will be close to the discriminant; a large value means expect a substantial rotation. - Solve
. In two dimensions invert directly; in general solve the system rather than forming the inverse. - Do not normalise unless you need the magnitude. The direction is determined up to positive scale and the classification is unaffected.
To classify.
- Project every observation: compute
. - Compute the two projected class means.
- Set the threshold at their midpoint, and only if the classes are equally likely and the two errors cost the same. Otherwise move it, and record that you did and why.
- Assign by comparison against the threshold, and count the errors.
- Report the projected means alongside the threshold, so a reader sees the margin and not merely the count.
To check the assumption the pooling made.
- Compare the two class scatters before pooling. Compare their traces, and compare their shapes. A ratio near one in both respects supports the pooling.
- If the traces differ by an order of magnitude, the assumption has failed. Say so, and expect the linear boundary to be poor regardless of how it was fitted.
- Check whether the class means are close. If
is near zero, the numerator of the construction carries almost nothing and no linear rule will separate the classes, whatever the scatter looks like. - Demonstrate the consequence rather than asserting it. Classify with the linear rule and count the errors; compare against the chance rate for the class balance at hand.
- Compare against a rule using each class's own covariance. If it does substantially better, the cost of the shared-covariance assumption is that difference, stated as a number.
Checks. Confirm
Worked example
Twelve points, two directions, one inverse
The data. Six observations per class, two measurements each.
| class | class |
|---|---|
Step 1: class means.
So
Step 2: scatter within each class. Summing
Read the off-diagonal. The within-class correlation is
so both classes are strongly elongated along a shared diagonal. That is the fact the next step exploits.
Step 3: invert and multiply. The determinant is
For a
Numerically
Step 4: project and threshold. Computing
| class | class | ||
|---|---|---|---|
Projected class means are
Every class
Step 5: the comparison that justifies the inverse. Repeat with the naive direction
| class | class | ||
|---|---|---|---|
Projected means
The separation ratios. Evaluating the Fisher criterion on each direction:
The naive direction produces a larger gap between projected means,
Contrast
Pairs that differ in one respect
The discriminant direction against the difference of means, same data.
| direction | ||
| angle | ||
| gap between projected means | ||
| Fisher criterion | ||
| errors |
The naive direction produces a gap more than ten times larger and separates worse. That single row is the argument for the inverse: a gap is not a measure of separation unless the spread beside it is known.
A rotation that matters against one that does not.
When the within-class correlation is
So the correction's size is a property of the data rather than of the method, and inspecting
A failure of fitting against a failure of form.
On the ring dataset the linear rule misclassifies
The per-class covariance rule misclassifies
Pooled covariance against per-class covariance.
Pooling assumes the classes share a shape and gives a linear boundary, fewer parameters, more stable on small samples, optimal when the assumption holds. Allowing each class its own covariance gives a quadratic boundary, more flexible, and on the ring data the difference between
The cost is parameters: two covariance matrices instead of one, estimated from the observations of a single class each. With few observations per class the extra flexibility can hurt, which is why the pooled version is not simply obsolete.
A direction against its negative.
Multiplying
Training error against what it estimates.
Every error count in this unit was computed on the observations that produced the means and the scatter. The
Warning
Errors in fitting and interpreting a linear discriminant
Using the difference of means as the direction. It maximises the gap between projected class means and ignores the spread around them. On this unit's data it gives a gap more than ten times larger than the discriminant's and separates worse,
Citing the equal-covariance assumption instead of testing it. The assumption is checkable in one step: compare the two class scatters before pooling them. On the ring data their traces are
Taking a returned boundary as evidence that one exists. The method inverts, multiplies and thresholds for any input. On the ring data it produces a direction, a threshold and a classification with no error, warning or diagnostic, and misclassifies
Diagnosing a bad error rate as a fitting problem. When the class means nearly coincide, the construction's numerator carries almost nothing, and no linear rule separates the classes however it is fitted. More data from the same process would not change this. The per-class rule reaching
Reporting the training error as performance. Every error count here was computed on the observations that produced the means and scatter, so it is optimistic for the reason any training error is. Comparing two directions on identical data is a legitimate use; estimating future performance is not, and needs held-out or resampled error.
Accepting the midpoint threshold without asking. The midpoint is right when the classes are equally likely and the two kinds of error cost the same. In screening problems neither typically holds, a missed case and a false alarm rarely cost alike, and the threshold that follows from the actual costs can be far from the midpoint. It is a decision, and it belongs in the report.
---
Two that follow from the matrix being inverted.
Ignoring near-singular
Computing scatter about the grand mean. Deviations must be taken from each class's own mean. Using the overall mean folds the between-class separation into the within-class term, so the matrix being inverted contains the very signal the numerator is trying to isolate, and the direction is wrong in a way that is easy to miss because the arithmetic still completes.
---
And one about what the method is for. It builds a boundary; it does not judge whether a boundary is the right description. If the classes overlap because the measurements do not distinguish them, a discriminant reports that overlap as an error rate and cannot say whether better measurements, a different boundary shape, or a different question is the remedy.
Application
Settings suited to a linear boundary
Dimension reduction before another method. With
Screening and diagnosis, where the threshold is the decision. The direction is fitted from data; the threshold encodes what an error costs. A missed case and a false alarm almost never cost the same in medical screening, so the midpoint is the wrong choice by default. Moving the threshold trades the two error types against each other along a curve, and choosing a point on that curve is a policy decision that the fitting cannot make.
Small samples with many predictors. The pooled covariance is estimated from all observations at once, so it is far more stable than two separate per-class estimates when each class is small. This is the practical argument for the linear method over the quadratic one even where the assumption is imperfect: a slightly wrong shared covariance, well estimated, can beat two right ones estimated from a handful of points each.
Where predictors nearly duplicate one another,
Quality control with drifting classes. A boundary fitted on historical data assumes the classes keep their shape and position. A process that drifts moves one class relative to the other, and the boundary degrades without any signal from the method. The error rate rises only if labels keep arriving to measure it.
Where the linear form is the wrong tool. Classes nested one inside the other, or separated by a curved frontier, admit no linear boundary at all. The ring example in this unit is the extreme:
---
What decides in each case. Whether the classes plausibly share a shape, how much data supports each class separately, and what the two errors cost. The method supplies a direction; every one of those three is supplied by the analyst, and none is visible in the output.