Linear Discriminants, and the Covariance They Assume

What you will be able to do

The learner can compute a linear discriminant direction from class means and the pooled within-class scatter, explain why that direction generally differs from the line joining the class means, classify observations by projecting onto it, and identify when the equal-covariance assumption fails badly enough that a single linear boundary cannot separate the classes at all.

Orientation

The direction that points the wrong way

Twelve observations, two classes, two measurements each. Class 0 centres on ( 3.5 , 3.1 6 ― ) and class 1 on ( 5.5 , 7.1 6 ― ) .

To separate them, aim at the line joining the centres. That difference is ( 2 , 4 ) , pointing up and to the right at 63.4349 degrees. This is the direction most readers name first.

The direction that actually separates these classes is ( − 0.568182 , 0.750000 ) , pointing up and to the left, at 127.1467 degrees. The two differ by 63.7117 degrees, and the sign of the first component is reversed. Nothing about the two class centres suggests it.

The reason is that both classes are stretched along the same diagonal, with a within-class correlation of 0.897150 . Projecting onto the line joining the centres spreads each class over a wide interval, so the two overlap. Projecting onto the perpendicular-ish direction gives a smaller gap between the centres and a much smaller spread within each class, and separation is a matter of the ratio, not the gap.

The arithmetic confirms it. Projecting onto the discriminant direction classifies all twelve observations correctly; projecting onto the difference of means misclassifies one. The Fisher separation ratio is 1.863636 against 0.911854 , better by a factor of 2.043788 .

Getting there takes one matrix inverse: divide the difference of means by the pooled within-class scatter, and the direction rotates itself.

And then the harder question. That pooling assumes both classes have the same shape and differ only in position. Where the assumption holds, the resulting straight boundary is the best one available. Where it fails, the method gives no sign: it returns a direction, applies a threshold, and reports a classification exactly as it does when everything is fine.

This unit covers the construction, why the inverse rotates the answer, and a dataset where one class surrounds the other, their means nearly coincident, their spreads differing by a factor of 16.1852 , on which the method misclassifies 7 of 18 while a rule permitting each class its own shape gets every one right.

Definition

What each factor in the direction contributes

The canonical statement above gives w = S w − 1 ( x ¯ 1 − x ¯ 0 ) and the threshold rule. What follows is what each piece decides, since those are the parts a reader of a fitted discriminant needs back.

The two factors answer different questions.

FactorSuppliesIgnores
x ¯ 1 − x ¯ 0 where the classes differhow wide each class is
S w − 1 how wide each class is, per directionwhere the classes differ

Each alone is insufficient, and multiplying them is what produces separation relative to spread rather than separation alone.

Why scatter and not covariance. S w sums squared deviations without dividing by a count. Dividing by n − 2 gives the pooled covariance estimate, and since the direction is determined only up to scale, either may be used. The classification is unchanged. Reporting which was used matters only when the magnitude of w is quoted.

The Fisher criterion is what is being maximised. For a direction v , form

J ( v ) = ( v T ( x ¯ 1 − x ¯ 0 ) ) 2 v T S w v ,

the squared separation of the projected means over the within-class spread along v . Maximising J gives w = S w − 1 ( x ¯ 1 − x ¯ 0 ) , which is why the inverse appears rather than being an arbitrary correction. J also provides the natural way to compare two candidate directions, since it is the quantity one of them maximises.

Scale invariance, and what it means for an answer. Replacing w by c w for any c > 0 multiplies every projection by c and moves the threshold by the same factor, leaving every classification intact. So a direction differing from a reference by a positive scalar is the same answer; one differing by a negative scalar is the same boundary with the class labels exchanged.

What the threshold encodes. The midpoint between projected class means is the right choice only when the two classes are equally likely in the population and the two kinds of error cost the same. Both are assumptions about the world rather than facts about the data, and relaxing either moves the threshold, which is a decision to be recorded, not a computation to be performed.

Where the construction breaks down. S w must be invertible. With more predictors than observations, or with one predictor an exact combination of others, it is singular and w is undefined. Near-singularity is worse than outright failure, because a direction is still returned and is highly sensitive to the particular sample. The same instability that motivates a penalty in regression, and the reason shrunken discriminants exist.

Intuition

Separation is a ratio, not a gap

Imagine two long thin clouds of points, both tilted the same way, one offset from the other along that same tilt. Their centres are far apart. Aim at the line joining the centres and project: the projected centres are indeed far apart, and each cloud, being long in exactly that direction, smears across a wide interval. The two intervals overlap, and points near the facing ends land in the wrong one.

Now aim across the tilt instead. The projected centres are closer together, which sounds worse. But each cloud is thin in this direction, so each projects to a narrow interval, and the two intervals do not touch.

A smaller gap and a much smaller spread can separate better than a large gap and a huge spread. That is the entire idea: the gap relative to the spread is what counts, and any method optimising the gap alone will choose badly whenever the classes are elongated.

How the inverse expresses this. S w records how far the classes spread in each direction, including how those directions are correlated. Multiplying by S w − 1 divides the difference of means by that spread, direction by direction, so directions along which the classes are internally stretched get discounted, and directions in which they are tight get promoted.

When the classes have no internal correlation and equal spread in every direction, S w is a multiple of the identity, its inverse is too, and w points along the difference of means after all. That is the case where the naive answer happens to be right, and it is the reason the naive answer seems plausible: it is correct exactly when the classes are round.

On this unit's data the correction is drastic. The within-class correlation is 0.897150 , so the classes are strongly elongated along a shared diagonal. The difference of means points at 63.4349 degrees; the discriminant points at 127.1467 degrees, 63.7117 degrees away, with the first component changing sign. The result is 0 errors against 1 , and a separation ratio of 1.863636 against 0.911854 .

Now the assumption. Adding the two scatters into one matrix says the classes have the same shape and differ only in where they sit. When that is true, a straight boundary is genuinely the best available, and this rule finds it.

When it is false, S w is an average describing neither class, and the direction built from it can be worthless. The sharpest version: one class a tight cloud, the other a ring around it. Their means nearly coincide, so x ¯ 1 − x ¯ 0 , the numerator of the whole construction, is almost nothing, and no direction admits a threshold separating a ring from its own centre.

The method does not notice. It inverts, multiplies, thresholds and reports, misclassifying 7 of 18 where chance would give about 9 . A rule allowing each class its own covariance gets all 18 right on the same points. The output of the failed fit is indistinguishable in form from the output of a successful one, which is why the assumption is worth testing rather than reciting.

Example

Four configurations and what the discriminant does with each

Elongated and correlated — the rotation case. The unit's twelve-point dataset: within-class correlation 0.897150 , difference of means at 63.4349 degrees, discriminant at 127.1467 degrees. Rotation of 63.7117 degrees, 0 errors against 1 , separation ratio 2.043788 better.

This is where the inverse earns its place. The classes are stretched along the same diagonal, so the direction joining their centres is precisely the direction in which each class is most spread out.

Round and equal — the case where the naive answer is right. If both classes have the same spread in every direction and no internal correlation, S w is a multiple of the identity. Its inverse is then also a multiple of the identity, and

w = S w − 1 ( x ¯ 1 − x ¯ 0 ) ∝ x ¯ 1 − x ¯ 0 .

The discriminant points exactly along the line joining the centres, and the rotation is zero. Worth knowing, because it explains why the naive answer feels correct: it is correct, for round classes, and most textbook diagrams draw round classes.

A ring around a cloud — no linear boundary exists. Class 0 is nine points clustered at the origin; class 1 is nine points on a ring around it. The class means are ( 0 , 0 ) and ( 0.7 7 ― , 0.1 1 ― ) , nearly the same point, so the difference of means is almost nothing.

The scatter traces are 24 and 3496 9 , a ratio of 16.1852 : the classes differ enormously in spread, and the shared-covariance assumption is false.

The method returns a direction anyway and misclassifies 7 of 18 , or 38.89 % , against a chance rate of 50 % for two balanced classes. It is barely better than guessing, and no threshold on that projection does better, because a ring cannot be cut from its own centre by a straight line.

The same ring, with per-class covariances. Assign each observation to the class under whose own spread it is least surprising. A quadratic rather than linear boundary. On the identical eighteen points this misclassifies 0 .

The data was always separable. What was not available was a straight boundary.

---

The first two bracket the linear method: it improves substantially on the naive direction when classes are elongated, and reduces to that direction when they are round. The third shows the method failing while giving every appearance of having worked. A direction, a threshold, a classification, and no indication that the assumption behind them was false. The fourth shows the failure was in the shape of boundary permitted, not in the fitting or the data.

Procedure

Constructing the discriminant, and testing what it assumed

To compute the direction.

  1. Compute each class mean separately. Not the overall mean. The construction needs the two centres and the difference between them.
  2. Compute each class's scatter about its own mean, S k = ∑ i ∈ C k ( x i − x ¯ k ) ( x i − x ¯ k ) T . Deviations are taken from the class mean, not the grand mean; using the grand mean would fold the between-class separation into the within-class term and defeat the whole construction.
  3. Add them: S w = S 0 + S 1 .
  4. Read the off-diagonal before going further. Divide it by S 11 S 22 to get the within-class correlation. A value near zero means the naive direction will be close to the discriminant; a large value means expect a substantial rotation.
  5. Solve S w w = x ¯ 1 − x ¯ 0 . In two dimensions invert directly; in general solve the system rather than forming the inverse.
  6. Do not normalise unless you need the magnitude. The direction is determined up to positive scale and the classification is unaffected.

To classify.

  1. Project every observation: compute w T x .
  2. Compute the two projected class means.
  3. Set the threshold at their midpoint, and only if the classes are equally likely and the two errors cost the same. Otherwise move it, and record that you did and why.
  4. Assign by comparison against the threshold, and count the errors.
  5. Report the projected means alongside the threshold, so a reader sees the margin and not merely the count.

To check the assumption the pooling made.

  1. Compare the two class scatters before pooling. Compare their traces, and compare their shapes. A ratio near one in both respects supports the pooling.
  2. If the traces differ by an order of magnitude, the assumption has failed. Say so, and expect the linear boundary to be poor regardless of how it was fitted.
  3. Check whether the class means are close. If x ¯ 1 − x ¯ 0 is near zero, the numerator of the construction carries almost nothing and no linear rule will separate the classes, whatever the scatter looks like.
  4. Demonstrate the consequence rather than asserting it. Classify with the linear rule and count the errors; compare against the chance rate for the class balance at hand.
  5. Compare against a rule using each class's own covariance. If it does substantially better, the cost of the shared-covariance assumption is that difference, stated as a number.

Checks. Confirm S w is symmetric and its determinant positive; a non-positive determinant means the scatter was computed wrongly or the classes are degenerate. Confirm the projected class means fall on opposite sides of the threshold, if they do not, the direction or the threshold is wrong rather than the data being hard. And before reporting an error count, note that it was computed on the observations that produced the means and scatter, so it is a training error and is optimistic.

Worked example

Twelve points, two directions, one inverse

The data. Six observations per class, two measurements each.

class 0 class 1
( 2 , 2 ) ( 4 , 6 )
( 3 , 3 ) ( 5 , 7 )
( 4 , 4 ) ( 6 , 8 )
( 5 , 5 ) ( 7 , 9 )
( 3 , 2 ) ( 5 , 6 )
( 4 , 3 ) ( 6 , 7 )

Step 1: class means.

x ¯ 0 = ( 21 6 , 19 6 ) = ( 3.5 ,   3.1 6 ― ) , x ¯ 1 = ( 33 6 , 43 6 ) = ( 5.5 ,   7.1 6 ― ) .

So x ¯ 1 − x ¯ 0 = ( 2 , 4 ) , at 63.4349 degrees from the x -axis. This is the naive direction.

Step 2: scatter within each class. Summing ( x i − x ¯ k ) ( x i − x ¯ k ) T over each class and adding:

S w = ( 11 11 11 41 3 ) .

Read the off-diagonal. The within-class correlation is

11 11 × 41 3 = 0.897150 ,

so both classes are strongly elongated along a shared diagonal. That is the fact the next step exploits.

Step 3: invert and multiply. The determinant is

det S w = 11 × 41 3 − 11 2 = 451 3 − 121 = 88 3 .

For a 2 × 2 matrix, S w − 1 = 1 det ( S 22 − S 12 − S 12 S 11 ) , so

w = S w − 1 ( 2 4 ) = 3 88 ( 41 3 ( 2 ) − 11 ( 4 ) − 11 ( 2 ) + 11 ( 4 ) ) = 3 88 ( − 50 3 22 ) = ( − 25 44 3 4 ) .

Numerically w = ( − 0.568182 ,   0.750000 ) , at 127.1467 degrees, 63.7117 degrees from the difference of means, with the first component now negative.

Step 4: project and threshold. Computing w T x for every observation:

class 0 class 1
( 2 , 2 ) 0.3636 ( 4 , 6 ) 2.2273
( 3 , 3 ) 0.5455 ( 5 , 7 ) 2.4091
( 4 , 4 ) 0.7273 ( 6 , 8 ) 2.5909
( 5 , 5 ) 0.9091 ( 7 , 9 ) 2.7727
( 3 , 2 ) − 0.2045 ( 5 , 6 ) 1.6591
( 4 , 3 ) − 0.0227 ( 6 , 7 ) 1.8409

Projected class means are 0.386364 and 2.250000 , so the midpoint threshold is 1.318182 .

Every class 0 projection is below it; every class 1 projection is above. 0 errors of 12 , and the nearest point to the boundary, 1.6591 , clears it by 0.341 .

Step 5: the comparison that justifies the inverse. Repeat with the naive direction ( 2 , 4 ) :

class 0 class 1
( 2 , 2 ) 12 ( 4 , 6 ) 32
( 3 , 3 ) 18 ( 5 , 7 ) 38
( 4 , 4 ) 24 ( 6 , 8 ) 44
( 5 , 5 ) 30 ( 7 , 9 ) 50
( 3 , 2 ) 14 ( 5 , 6 ) 34
( 4 , 3 ) 20 ( 6 , 7 ) 40

Projected means 19. 6 ― and 39. 6 ― , threshold 29. 6 ― . The class 0 observation ( 5 , 5 ) projects to 30 , above the threshold, misclassified. 1 error of 12 .

The separation ratios. Evaluating the Fisher criterion on each direction:

J ( w ) = 1.863636 , J ( x ¯ 1 − x ¯ 0 ) = 0.911854 , ratio  2.043788 .

The naive direction produces a larger gap between projected means, 20 against 1.863636 in raw units, and separates worse, because the spread grew faster than the gap. One matrix inverse converts a gap into a ratio, and on data with correlated classes that rotates the answer by more than sixty degrees.

Contrast

Pairs that differ in one respect

The discriminant direction against the difference of means, same data.

S w − 1 ( x ¯ 1 − x ¯ 0 ) x ¯ 1 − x ¯ 0
direction ( − 0.568182 , 0.750000 ) ( 2 , 4 )
angle 127.1467 ∘ 63.4349 ∘
gap between projected means 1.863636 20
Fisher criterion 1.863636 0.911854
errors 0 of 12 1 of 12

The naive direction produces a gap more than ten times larger and separates worse. That single row is the argument for the inverse: a gap is not a measure of separation unless the spread beside it is known.

A rotation that matters against one that does not.

When the within-class correlation is 0.897150 , the inverse rotates the direction by 63.7117 degrees. When the classes are round, S w a multiple of the identity, it rotates by nothing at all, and the discriminant coincides with the difference of means.

So the correction's size is a property of the data rather than of the method, and inspecting S w tells you in advance whether it will matter.

A failure of fitting against a failure of form.

On the ring dataset the linear rule misclassifies 7 of 18 . The temptation is to look for a fitting error or to seek more data. Neither would help: the class means nearly coincide, so the construction's numerator carries almost no information, and no straight line separates a ring from its centre.

The per-class covariance rule misclassifies 0 on the identical points. The data was separable throughout; the linear boundary was the constraint.

Pooled covariance against per-class covariance.

Pooling assumes the classes share a shape and gives a linear boundary, fewer parameters, more stable on small samples, optimal when the assumption holds. Allowing each class its own covariance gives a quadratic boundary, more flexible, and on the ring data the difference between 7 errors and 0 .

The cost is parameters: two covariance matrices instead of one, estimated from the observations of a single class each. With few observations per class the extra flexibility can hurt, which is why the pooled version is not simply obsolete.

A direction against its negative.

Multiplying w by − 1 produces the same boundary with the class labels exchanged; multiplying by any positive constant changes nothing at all. So a reported direction that disagrees with a reference by a positive factor is the same answer, and one disagreeing by sign is the same boundary described from the other side. Only the line matters, not the arrow.

Training error against what it estimates.

Every error count in this unit was computed on the observations that produced the means and the scatter. The 0 of 12 is a training figure and is optimistic, as any training error is. It is used here to compare two directions on identical data, which it supports, and not to estimate how the rule would perform on new observations, which it does not.

Warning

Errors in fitting and interpreting a linear discriminant

Using the difference of means as the direction. It maximises the gap between projected class means and ignores the spread around them. On this unit's data it gives a gap more than ten times larger than the discriminant's and separates worse, 1 error against 0 , separation ratio 0.911854 against 1.863636 . It is correct only when the classes are round, which is the case most diagrams draw and few datasets satisfy.

Citing the equal-covariance assumption instead of testing it. The assumption is checkable in one step: compare the two class scatters before pooling them. On the ring data their traces are 24 and 3496 9 , differing by a factor of 16.1852 , visible immediately, and decisive. An analysis that names the assumption in a caveat and never computes it has recorded awareness rather than evidence.

Taking a returned boundary as evidence that one exists. The method inverts, multiplies and thresholds for any input. On the ring data it produces a direction, a threshold and a classification with no error, warning or diagnostic, and misclassifies 7 of 18 against a chance rate of 50 % . The output of a failed fit is identical in form to the output of a successful one.

Diagnosing a bad error rate as a fitting problem. When the class means nearly coincide, the construction's numerator carries almost nothing, and no linear rule separates the classes however it is fitted. More data from the same process would not change this. The per-class rule reaching 0 errors on the identical points shows the data was separable and the shape of the boundary was the constraint.

Reporting the training error as performance. Every error count here was computed on the observations that produced the means and scatter, so it is optimistic for the reason any training error is. Comparing two directions on identical data is a legitimate use; estimating future performance is not, and needs held-out or resampled error.

Accepting the midpoint threshold without asking. The midpoint is right when the classes are equally likely and the two kinds of error cost the same. In screening problems neither typically holds, a missed case and a false alarm rarely cost alike, and the threshold that follows from the actual costs can be far from the midpoint. It is a decision, and it belongs in the report.

---

Two that follow from the matrix being inverted.

Ignoring near-singular S w . With predictors that nearly duplicate one another, S w is near-singular and the direction is highly sensitive to the sample. The instability that motivates a penalty in regression, arriving here in the same form. A direction reported from such a fit may change substantially on a resample.

Computing scatter about the grand mean. Deviations must be taken from each class's own mean. Using the overall mean folds the between-class separation into the within-class term, so the matrix being inverted contains the very signal the numerator is trying to isolate, and the direction is wrong in a way that is easy to miss because the arithmetic still completes.

---

And one about what the method is for. It builds a boundary; it does not judge whether a boundary is the right description. If the classes overlap because the measurements do not distinguish them, a discriminant reports that overlap as an error rate and cannot say whether better measurements, a different boundary shape, or a different question is the remedy.

Application

Settings suited to a linear boundary

Dimension reduction before another method. With K classes the same construction yields up to K − 1 discriminant directions, and projecting onto them reduces the predictors to a handful of coordinates chosen specifically to separate the classes. Unlike a variance-based projection, which finds directions of greatest spread regardless of labels, this uses the labels, so it can keep a low-variance direction that happens to separate the classes and discard a high-variance one that does not.

Screening and diagnosis, where the threshold is the decision. The direction is fitted from data; the threshold encodes what an error costs. A missed case and a false alarm almost never cost the same in medical screening, so the midpoint is the wrong choice by default. Moving the threshold trades the two error types against each other along a curve, and choosing a point on that curve is a policy decision that the fitting cannot make.

Small samples with many predictors. The pooled covariance is estimated from all observations at once, so it is far more stable than two separate per-class estimates when each class is small. This is the practical argument for the linear method over the quadratic one even where the assumption is imperfect: a slightly wrong shared covariance, well estimated, can beat two right ones estimated from a handful of points each.

Where predictors nearly duplicate one another, S w approaches singularity and the direction becomes unstable in exactly the way an unpenalised regression coefficient does, which is why shrunken and regularised discriminants exist, adding to the diagonal for the same reason and with the same trade.

Quality control with drifting classes. A boundary fitted on historical data assumes the classes keep their shape and position. A process that drifts moves one class relative to the other, and the boundary degrades without any signal from the method. The error rate rises only if labels keep arriving to measure it.

Where the linear form is the wrong tool. Classes nested one inside the other, or separated by a curved frontier, admit no linear boundary at all. The ring example in this unit is the extreme: 7 errors of 18 from the linear rule and 0 from a per-class covariance rule on identical data. The diagnostic is to compare class scatters and class means before fitting, near-coincident means with very different spreads is the signature.

---

What decides in each case. Whether the classes plausibly share a shape, how much data supports each class separately, and what the two errors cost. The method supplies a direction; every one of those three is supplied by the analyst, and none is visible in the output.

Next step

Practice Linear Discriminants, and the Covariance They Assume

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.