Practice: Linear Discriminants, and the Covariance They Assume

Direct application

Two classes of six observations each give the pooled within-class scatter

S w = ( 11 11 11 41 3 ) .

Compute det S w , to six decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

The determinant of ( a b c d ) is a d − b c .

Hint 2: Next step

11 × 41 3 = 451 3 , and 121 = 363 3 .

Direct application

With

S w = ( 11 11 11 41 3 ) , x ¯ 1 − x ¯ 0 = ( 2 4 ) ,

compute the discriminant direction w = S w − 1 ( x ¯ 1 − x ¯ 0 ) .

Report its first component, w 1 , to six decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

det S w = 88 3 .

Hint 2: Next step

The first row of the inverse is 1 det ( S 22 , − S 12 ) .

Direct application

On a two-class dataset, the Fisher separation criterion, squared separation of the projected class means divided by the within-class spread along the projection, evaluates to 1.863636 for the discriminant direction and 0.911854 for the difference of class means.

By what factor does the discriminant direction separate better? Report the ratio to six decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

1 hint available, least help first.

Hint 1: Retrieval cue

A larger Fisher criterion means better separation relative to spread.

Direct application

Projecting twelve observations onto the discriminant direction gives:

  • class 0 : 0.3636 , 0.5455 , 0.7273 , 0.9091 , − 0.2045 , − 0.0227
  • class 1 : 2.2273 , 2.4091 , 2.5909 , 2.7727 , 1.6591 , 1.8409

The projected class means are 0.386364 and 2.250000 . Using the midpoint of those two means as the threshold, assigning below it to class 0 and above it to class 1 :

How many of the twelve observations are misclassified?

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

The midpoint threshold is the average of the two projected class means.

Hint 2: Next step

Compare the largest class 0 projection and the smallest class 1 projection against that threshold.

Explanation · Comparison

On a two-class dataset, projecting onto the difference of class means ( 2 , 4 ) puts the projected class means 20 apart. Projecting onto the discriminant direction ( − 0.568182 , 0.750000 ) puts them only 1.863636 apart.

Yet the discriminant direction misclassifies 0 of 12 observations and the difference of means misclassifies 1 . The pooled within-class scatter is S w = ( 11 11 11 41 / 3 ) .

(a) Explain how a direction with a gap ten times smaller can separate better.

(b) Compute the within-class correlation from S w and say what it tells you about the shape of the classes.

(c) State the condition under which the two directions would coincide, and explain why that makes the naive answer feel right.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Two intervals overlap depending on their separation and their widths.

Hint 2: Concept cue

For (c), ask what S w − 1 does to a vector when S w is a multiple of the identity.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) Separation is a ratio, not a gap.

What decides whether two projected classes overlap is the gap between their centres relative to how widely each class spreads around its own centre. A direction can push the centres far apart and, if the classes are elongated along that same direction, spread each class over an even wider interval, so the intervals still overlap.

That is what happens here. The difference of means points along the direction in which both classes are internally stretched, so the gap of 20 comes with an enormous within-class spread. The discriminant direction gives a gap of only 1.863636 , but each class projects to a narrow interval, and the intervals do not touch.

The Fisher criterion makes this precise, dividing squared gap by within-class spread: 1.863636 for the discriminant against 0.911854 for the difference of means, better by a factor of 2.043788 . The raw gaps are not comparable quantities, because they are measured along directions with different scales.

(b) The within-class correlation.

S 12 S 11 S 22 = 11 11 × 41 3 = 11 451 3 = 0.897150 .

A correlation near 0.9 means the two measurements move together strongly within each class: each class is a long, thin, diagonally oriented cloud rather than a round blob. The two variables are nearly redundant inside a class, so the classes are stretched along one diagonal and narrow across it.

That shape is exactly what makes the naive direction fail, because the line joining the class centres runs roughly along the elongation.

(c) When the two coincide.

They coincide when S w is a multiple of the identity, equal spread in both variables and zero within-class correlation, so each class is round. Then S w − 1 is also a multiple of the identity, and

w = S w − 1 ( x ¯ 1 − x ¯ 0 ) ∝ x ¯ 1 − x ¯ 0 ,

the same direction up to scale, which leaves the classification unchanged.

This is why the naive answer feels right: it is right, for round classes. Diagrams illustrating classification almost always draw round clouds, so the case where the correction does nothing is the case most often pictured. The correction's size is a property of the data, inspect S w and you can tell in advance whether it will matter.

A complete answer does each of these:

  • explains whitening
  • computes pooled scatter

Error diagnosis · Method selection

An analyst fits a linear discriminant to eighteen observations in two balanced classes and obtains a direction, a threshold, and a classification with no warnings. It misclassifies 7 of the 18 .

They conclude the classifier needs more training data.

The two class means are ( 0 , 0 ) and ( 0.7 7 ― , 0.1 1 ― ) . The class scatter traces are 24 and 3496 9 . A rule allowing each class its own covariance misclassifies 0 of the 18 on the same observations.

(a) Compute the chance error rate for this setting and compare it with 7 of 18 .

(b) Using the means and the scatter traces, give two distinct reasons the linear rule cannot work here.

(c) Assess the recommendation to collect more data, and say what the per-class result establishes.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

What is the expected error rate if you assign at random between two balanced classes?

Hint 2: Concept cue

Look at the numerator of w = S w − 1 ( x ¯ 1 − x ¯ 0 ) when the class means nearly coincide.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) Barely better than guessing. With two balanced classes, assigning at random gives an expected error rate of 50 % , or 9 of 18 . The fitted rule misclassifies 7 of 18 , which is 38.89 % . So the classifier improves on chance by about eleven percentage points. Two observations. For a method given labelled data and a two-dimensional predictor space, that is close to no information at all, and it is the figure that should have prompted the diagnosis rather than a request for more data. (b) Two independent reasons. The class means nearly coincide. They are ( 0 , 0 ) and ( 0.7 7 ― , 0.1 1 ― ) , so x ¯ 1 − x ¯ 0 is a very short vector. That difference is the numerator of the whole construction, w = S w − 1 ( x ¯ 1 − x ¯ 0 ) : if the classes have essentially the same centre, there is no direction along which their centres separate, and every projection puts the two classes on top of one another. This alone defeats any linear rule, whatever the covariances. The covariances are grossly unequal. The scatter traces are 24 and 3496 9 ≈ 388.44 , a ratio of 16.1852 . Pooling them produces a matrix describing neither class. An average of a tight cloud and a very broad one. The linear form is optimal only when both classes share a covariance, and that premise is false by more than an order of magnitude. The two reasons are separate. Fixing the covariance disparity would not create a separation between coincident means, and separating the means would not repair the pooling. (c) More data would not help, and the per-class result shows why. Additional observations from the same process would estimate the same means and the same covariances more precisely. The means would remain nearly coincident and the covariances would remain unequal, because these are properties of the populations rather than artefacts of a small sample. The error rate would not improve. What the per-class result establishes is decisive: a rule permitting each class its own covariance misclassifies 0 of 18 on the identical observations. So the data contains all the information needed to separate the classes perfectly. The information was never missing. The constraint was the shape of boundary the method was permitted to draw. A linear rule can only cut the plane with a straight line, and these classes are arranged so that no straight line separates them: the configuration is one class surrounding the other, which a quadratic boundary handles and a linear one cannot. The correct recommendation is to change the method, not the sample size. And the general lesson is that the fitted output gave no sign of any of this. A direction, a threshold and a classification were produced exactly as they are when the assumptions hold.

A complete answer does each of these:

  • detects assumption failure
  • classifies by projection

Interpretation · Evaluation

A report reads:

"We fitted a linear discriminant to our twelve labelled samples. The direction is ( − 0.568182 , 0.750000 ) and the threshold 1.318182 . The classifier achieved 100 % accuracy. The negative first coefficient shows that variable 1 is inversely related to class membership. We recommend deploying it for screening, using the same midpoint threshold."

(a) Assess the 100 % accuracy figure.

(b) Assess the interpretation of the negative coefficient, given that both variables increase from class 0 to class 1 and their within-class correlation is 0.897150 .

(c) Assess the recommendation to deploy with the midpoint threshold for screening.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Which observations produced the direction, and which were used to measure the accuracy?

Hint 2: Concept cue

For (c), list the two assumptions the midpoint threshold encodes and ask whether a screening setting satisfies them.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) It is a training figure, and it is on twelve samples. The direction and threshold were computed from the same twelve observations the accuracy was measured on. The class means and pooled scatter that produced w came from those points, so the rule was fitted to separate exactly them. A perfect score under those conditions is the weakest evidence of performance, not the strongest. The sample size compounds it. Twelve observations, six per class, in two dimensions, perfect separation is unsurprising even for classes that overlap substantially in the population, and would be achievable for many random label assignments. What the figure does support is a comparison on this data between two candidate directions, which is how the unit uses it: 0 errors for the discriminant against 1 for the difference of means. What it does not support is any claim about future observations. That needs held-out or resampled error, and none is reported. (b) The coefficient interpretation is wrong. A discriminant coefficient is not a per-variable effect. The direction is S w − 1 ( x ¯ 1 − x ¯ 0 ) , and the inverse mixes the variables together. Each component depends on the scatter of both and on their correlation, not on one variable's relationship to the class. The data contradicts the reading directly. Both variables increase from class 0 to class 1 : the means go from ( 3.5 , 3.1 6 ― ) to ( 5.5 , 7.1 6 ― ) , so the difference ( 2 , 4 ) is positive in both components. Variable 1 is positively associated with class membership, and the discriminant's first coefficient is still negative. The reason is the within-class correlation of 0.897150 . Because the variables move together strongly inside each class, the discriminant achieves separation by taking a contrast between them, roughly, how much variable 2 exceeds what variable 1 predicts. A negative weight on variable 1 is what constructs that contrast, and it says nothing about variable 1's own direction of association. Reading individual discriminant coefficients as effects is the same error as reading individual regression coefficients as effects when predictors are collinear: the combination is determined, the split between them is not interpretable. (c) The threshold is the wrong default for screening. The midpoint between projected class means is correct only when two conditions hold: the classes are equally likely in the population, and the two kinds of error cost the same. Neither is typical in screening. The condition being screened for is usually rare, so a threshold calibrated on a balanced sample of six and six misrepresents the prior. And a missed case and a false alarm almost never carry equal cost, in most screening settings a miss is far worse, which argues for moving the threshold toward catching more cases at the price of more false alarms. So the threshold should be set from the prevalence and the relative costs, both of which are decisions about the application rather than quantities computed from the fit. The report inherits a default from a balanced twelve-sample study and proposes deploying it unchanged. Before deployment, the defensible sequence is: estimate performance on data not used in fitting, establish the prevalence in the screened population, decide the relative cost of the two errors, and set the threshold from those, reporting the resulting sensitivity and specificity rather than a single accuracy number, which conceals the trade entirely.

A complete answer does each of these:

  • classifies by projection
  • explains whitening

Transfer · Evaluation · Explanation

A factory classifies components as sound or defective from two sensor readings. An engineer fits a linear discriminant on 40 labelled components, 20 of each class, and reports 95 % accuracy on those 40 .

Six months later the rule is misclassifying about a third of components. The team's diagnostics show:

  • sound components: mean ( 10.0 , 10.0 ) , scatter trace 180 ;
  • defective components: mean ( 10.4 , 10.2 ) , scatter trace 2,900 ;
  • defects arise from several unrelated causes, each pushing a reading away from nominal in a different direction;
  • defects are about 2 % of production, not the 50 % in the fitting sample;
  • a missed defect costs roughly 200 times a false alarm.

The team proposes collecting 400 labelled components and refitting the same model.

Write a review covering:

(a) What the original 95 % figure established, and what it did not.

(b) What the means and scatter traces say about whether a linear boundary can work here, with the relevant ratio computed.

(c) Whether refitting on ten times the data addresses the problem.

(d) What the prevalence and cost figures imply about the threshold.

(e) What you would do instead, and what you would report.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

Compare the distance between class means against the within-class spreads.

Hint 2: Concept cue

Defects arising from several unrelated causes place the defective class where, relative to the sound class?

Hint 3: Strategy cue

Separate four questions: what the accuracy measured, whether a linear boundary exists, whether data volume is the constraint, and where the threshold belongs.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The 95 % was a training figure on a balanced sample. The direction and threshold were computed from the same 40 components the accuracy was measured on, so the rule was fitted to separate exactly those points. Training accuracy is optimistic by construction, and on 40 observations in two dimensions a high figure is easy to obtain even for classes that overlap in the population. It was also measured on a sample that was 50 % defective, against a production rate of 2 % . Accuracy on a balanced sample says nothing about accuracy on a stream where one class is rare. A rule calling everything sound would score 98 % on production and be useless. What the figure supports: a comparison of candidate directions on that data. What it does not support: any statement about performance on future components, which needs error on observations not used in fitting. (b) A linear boundary cannot work here, for two separate reasons. The class means nearly coincide. ( 10.0 , 10.0 ) against ( 10.4 , 10.2 ) , so x ¯ 1 − x ¯ 0 = ( 0.4 , 0.2 ) , tiny against within-class spreads with traces in the hundreds and thousands. That difference is the numerator of w = S w − 1 ( x ¯ 1 − x ¯ 0 ) , so the construction is built on almost no signal. Projected onto any direction, the two classes will have nearly the same centre. The covariances are grossly unequal. The scatter traces are 180 and 2,900 , a ratio of

2900 180 = 16.1 1 ― ,

over sixteen to one. Pooling them yields a matrix describing neither class. The linear boundary is optimal only when both classes share a covariance, and that premise fails by more than an order of magnitude. The mechanism explains both. Defects arise from several unrelated causes pushing readings away from nominal in different directions, so the defective class is a diffuse cloud surrounding the tight sound class, with roughly the same centre. This is the ring-around-a-cloud configuration, and no straight line separates a surrounding shell from its own centre. The 95 % was fitted to a particular sample of that shell; as different defect causes appeared over six months, the sampled shell filled out and the fitted line stopped corresponding to anything. (c) Refitting on 400 components does not address it. More data estimates the same means and the same covariances more precisely. The means would remain nearly coincident and the covariance ratio would remain about sixteen, because these are properties of the populations rather than artefacts of the sample size. The linear rule would be fitted more stably to a boundary that cannot separate the classes. The proposal also repeats the sampling error: if the 400 are again balanced, the threshold will again be calibrated for a prevalence the production line does not have. More data is worth collecting, it is needed for honest evaluation and for a per-class covariance estimate, but refitting the same model on it is the part that does not help. (d) Prevalence and costs both move the threshold, in the same direction. The midpoint between projected class means is correct only when the classes are equally likely and the two errors cost the same. Neither holds. Defects are 2 % of production, not 50 % . On its own, a rarer positive class argues for a threshold less willing to call a component defective. But a missed defect costs about 200 times a false alarm, which argues strongly the other way. Combining them, the relevant comparison weighs prevalence against cost: 0.02 × 200 = 4 against 0.98 × 1 , so the expected cost of missing a defect outweighs that of a false alarm by roughly four to one. The threshold should be set to flag substantially more components than a midpoint rule would, accepting many false alarms to avoid misses. Neither figure is computable from the fit. Both are facts about the application, and the current rule encodes the opposite of both. (e) What to do instead. Change the shape of the boundary, or the framing. Given one tight class and one diffuse class surrounding it, the natural alternatives are a rule permitting each class its own covariance, which gives a quadratic boundary and is the direct fix for unequal covariances, or, better suited to the mechanism, treating this as anomaly detection rather than classification: model the sound class alone, which is tight and well characterised, and flag components far from it. That matches how defects actually arise, since "defective" is not one coherent class but several unrelated departures from nominal. Split the estimation from the evaluation. Fit on one portion and estimate error on another never used in fitting. Report performance on a sample reflecting the 2 % prevalence, not a balanced one. Report sensitivity and specificity, never accuracy. At 2 % prevalence accuracy is dominated by the common class and conceals exactly the errors that matter. Report the miss rate and the false-alarm rate separately, with the threshold that produced them. Set the threshold from prevalence and costs explicitly, stating both, so that the choice is reviewable rather than inherited from a balanced pilot study. Monitor for drift. The failure appeared over six months, which suggests the defect population changed. Whatever rule replaces this one should have its error rates tracked on continuing labelled samples, since no fitted boundary reports its own obsolescence. What I would tell the team. The original figure was measured on the data that produced the rule and on a class balance that does not exist in production, so it never supported deployment. The deeper problem is that the defective class surrounds the sound one rather than sitting beside it, visible in the near-identical means and the sixteen-to-one spread ratio, so no straight boundary was ever going to work, and the six months of apparent success were a property of the particular defects sampled rather than of the method.

A complete answer does each of these:

  • computes pooled scatter
  • derives discriminant direction
  • explains whitening
  • classifies by projection
  • detects assumption failure
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.