Two classes of six observations each give the pooled within-class scatter
Compute , to six decimal places.
Enter the value. It is checked against the answer and the precision this task asks for.
2 hints available, least help first.
Hint 1: Retrieval cue
The determinant of is .
Hint 2: Next step
, and .
Direct application
With
compute the discriminant direction .
Report its first component, , to six decimal places.
Enter the value. It is checked against the answer and the precision this task asks for.
2 hints available, least help first.
Hint 1: Retrieval cue
.
Hint 2: Next step
The first row of the inverse is .
Direct application
On a two-class dataset, the Fisher separation criterion, squared separation of the projected class means divided by the within-class spread along the projection, evaluates to for the discriminant direction and for the difference of class means.
By what factor does the discriminant direction separate better? Report the ratio to six decimal places.
Enter the value. It is checked against the answer and the precision this task asks for.
1 hint available, least help first.
Hint 1: Retrieval cue
A larger Fisher criterion means better separation relative to spread.
Direct application
Projecting twelve observations onto the discriminant direction gives:
class : , , , , ,
class : , , , , ,
The projected class means are and . Using the midpoint of those two means as the threshold, assigning below it to class and above it to class :
How many of the twelve observations are misclassified?
Enter the value. It is checked against the answer and the precision this task asks for.
2 hints available, least help first.
Hint 1: Retrieval cue
The midpoint threshold is the average of the two projected class means.
Hint 2: Next step
Compare the largest class projection and the smallest class projection against that threshold.
Explanation · Comparison
On a two-class dataset, projecting onto the difference of class means puts the projected class means apart. Projecting onto the discriminant direction puts them only apart.
Yet the discriminant direction misclassifies of observations and the difference of means misclassifies . The pooled within-class scatter is .
(a) Explain how a direction with a gap ten times smaller can separate better.
(b) Compute the within-class correlation from and say what it tells you about the shape of the classes.
(c) State the condition under which the two directions would coincide, and explain why that makes the naive answer feel right.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Two intervals overlap depending on their separation and their widths.
Hint 2: Concept cue
For (c), ask what does to a vector when is a multiple of the identity.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) Separation is a ratio, not a gap.
What decides whether two projected classes overlap is the gap between their centres relative to how widely each class spreads around its own centre. A direction can push the centres far apart and, if the classes are elongated along that same direction, spread each class over an even wider interval, so the intervals still overlap.
That is what happens here. The difference of means points along the direction in which both classes are internally stretched, so the gap of comes with an enormous within-class spread. The discriminant direction gives a gap of only , but each class projects to a narrow interval, and the intervals do not touch.
The Fisher criterion makes this precise, dividing squared gap by within-class spread: for the discriminant against for the difference of means, better by a factor of . The raw gaps are not comparable quantities, because they are measured along directions with different scales.
(b) The within-class correlation.
A correlation near means the two measurements move together strongly within each class: each class is a long, thin, diagonally oriented cloud rather than a round blob. The two variables are nearly redundant inside a class, so the classes are stretched along one diagonal and narrow across it.
That shape is exactly what makes the naive direction fail, because the line joining the class centres runs roughly along the elongation.
(c) When the two coincide.
They coincide when is a multiple of the identity, equal spread in both variables and zero within-class correlation, so each class is round. Then is also a multiple of the identity, and
the same direction up to scale, which leaves the classification unchanged.
This is why the naive answer feels right: it is right, for round classes. Diagrams illustrating classification almost always draw round clouds, so the case where the correction does nothing is the case most often pictured. The correction's size is a property of the data, inspect and you can tell in advance whether it will matter.
A complete answer does each of these:
explains whitening
computes pooled scatter
Error diagnosis · Method selection
An analyst fits a linear discriminant to eighteen observations in two balanced classes and obtains a direction, a threshold, and a classification with no warnings. It misclassifies of the .
They conclude the classifier needs more training data.
The two class means are and . The class scatter traces are and . A rule allowing each class its own covariance misclassifies of the on the same observations.
(a) Compute the chance error rate for this setting and compare it with of .
(b) Using the means and the scatter traces, give two distinct reasons the linear rule cannot work here.
(c) Assess the recommendation to collect more data, and say what the per-class result establishes.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
What is the expected error rate if you assign at random between two balanced classes?
Hint 2: Concept cue
Look at the numerator of when the class means nearly coincide.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) Barely better than guessing. With two balanced classes, assigning at random gives an expected error rate of , or of . The fitted rule misclassifies of , which is . So the classifier improves on chance by about eleven percentage points. Two observations. For a method given labelled data and a two-dimensional predictor space, that is close to no information at all, and it is the figure that should have prompted the diagnosis rather than a request for more data. (b) Two independent reasons.The class means nearly coincide. They are and , so is a very short vector. That difference is the numerator of the whole construction, : if the classes have essentially the same centre, there is no direction along which their centres separate, and every projection puts the two classes on top of one another. This alone defeats any linear rule, whatever the covariances. The covariances are grossly unequal. The scatter traces are and , a ratio of . Pooling them produces a matrix describing neither class. An average of a tight cloud and a very broad one. The linear form is optimal only when both classes share a covariance, and that premise is false by more than an order of magnitude. The two reasons are separate. Fixing the covariance disparity would not create a separation between coincident means, and separating the means would not repair the pooling. (c) More data would not help, and the per-class result shows why. Additional observations from the same process would estimate the same means and the same covariances more precisely. The means would remain nearly coincident and the covariances would remain unequal, because these are properties of the populations rather than artefacts of a small sample. The error rate would not improve. What the per-class result establishes is decisive: a rule permitting each class its own covariance misclassifies of on the identical observations. So the data contains all the information needed to separate the classes perfectly. The information was never missing. The constraint was the shape of boundary the method was permitted to draw. A linear rule can only cut the plane with a straight line, and these classes are arranged so that no straight line separates them: the configuration is one class surrounding the other, which a quadratic boundary handles and a linear one cannot. The correct recommendation is to change the method, not the sample size. And the general lesson is that the fitted output gave no sign of any of this. A direction, a threshold and a classification were produced exactly as they are when the assumptions hold.
A complete answer does each of these:
detects assumption failure
classifies by projection
Interpretation · Evaluation
A report reads:
"We fitted a linear discriminant to our twelve labelled samples. The direction is and the threshold . The classifier achieved accuracy. The negative first coefficient shows that variable 1 is inversely related to class membership. We recommend deploying it for screening, using the same midpoint threshold."
(a) Assess the accuracy figure.
(b) Assess the interpretation of the negative coefficient, given that both variables increase from class to class and their within-class correlation is .
(c) Assess the recommendation to deploy with the midpoint threshold for screening.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Which observations produced the direction, and which were used to measure the accuracy?
Hint 2: Concept cue
For (c), list the two assumptions the midpoint threshold encodes and ask whether a screening setting satisfies them.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) It is a training figure, and it is on twelve samples. The direction and threshold were computed from the same twelve observations the accuracy was measured on. The class means and pooled scatter that produced came from those points, so the rule was fitted to separate exactly them. A perfect score under those conditions is the weakest evidence of performance, not the strongest. The sample size compounds it. Twelve observations, six per class, in two dimensions, perfect separation is unsurprising even for classes that overlap substantially in the population, and would be achievable for many random label assignments. What the figure does support is a comparison on this data between two candidate directions, which is how the unit uses it: errors for the discriminant against for the difference of means. What it does not support is any claim about future observations. That needs held-out or resampled error, and none is reported. (b) The coefficient interpretation is wrong. A discriminant coefficient is not a per-variable effect. The direction is , and the inverse mixes the variables together. Each component depends on the scatter of both and on their correlation, not on one variable's relationship to the class. The data contradicts the reading directly. Both variables increase from class to class : the means go from to , so the difference is positive in both components. Variable 1 is positively associated with class membership, and the discriminant's first coefficient is still negative. The reason is the within-class correlation of . Because the variables move together strongly inside each class, the discriminant achieves separation by taking a contrast between them, roughly, how much variable 2 exceeds what variable 1 predicts. A negative weight on variable 1 is what constructs that contrast, and it says nothing about variable 1's own direction of association. Reading individual discriminant coefficients as effects is the same error as reading individual regression coefficients as effects when predictors are collinear: the combination is determined, the split between them is not interpretable. (c) The threshold is the wrong default for screening. The midpoint between projected class means is correct only when two conditions hold: the classes are equally likely in the population, and the two kinds of error cost the same. Neither is typical in screening. The condition being screened for is usually rare, so a threshold calibrated on a balanced sample of six and six misrepresents the prior. And a missed case and a false alarm almost never carry equal cost, in most screening settings a miss is far worse, which argues for moving the threshold toward catching more cases at the price of more false alarms. So the threshold should be set from the prevalence and the relative costs, both of which are decisions about the application rather than quantities computed from the fit. The report inherits a default from a balanced twelve-sample study and proposes deploying it unchanged. Before deployment, the defensible sequence is: estimate performance on data not used in fitting, establish the prevalence in the screened population, decide the relative cost of the two errors, and set the threshold from those, reporting the resulting sensitivity and specificity rather than a single accuracy number, which conceals the trade entirely.
A complete answer does each of these:
classifies by projection
explains whitening
Transfer · Evaluation · Explanation
A factory classifies components as sound or defective from two sensor readings. An engineer fits a linear discriminant on labelled components, of each class, and reports accuracy on those .
Six months later the rule is misclassifying about a third of components. The team's diagnostics show:
sound components: mean , scatter trace ;
defective components: mean , scatter trace ;
defects arise from several unrelated causes, each pushing a reading away from nominal in a different direction;
defects are about of production, not the in the fitting sample;
a missed defect costs roughly times a false alarm.
The team proposes collecting labelled components and refitting the same model.
Write a review covering:
(a) What the original figure established, and what it did not.
(b) What the means and scatter traces say about whether a linear boundary can work here, with the relevant ratio computed.
(c) Whether refitting on ten times the data addresses the problem.
(d) What the prevalence and cost figures imply about the threshold.
(e) What you would do instead, and what you would report.
Write your answer, then compare it with the worked solution.
3 hints available, least help first.
Hint 1: Retrieval cue
Compare the distance between class means against the within-class spreads.
Hint 2: Concept cue
Defects arising from several unrelated causes place the defective class where, relative to the sound class?
Hint 3: Strategy cue
Separate four questions: what the accuracy measured, whether a linear boundary exists, whether data volume is the constraint, and where the threshold belongs.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) The was a training figure on a balanced sample. The direction and threshold were computed from the same components the accuracy was measured on, so the rule was fitted to separate exactly those points. Training accuracy is optimistic by construction, and on observations in two dimensions a high figure is easy to obtain even for classes that overlap in the population. It was also measured on a sample that was defective, against a production rate of . Accuracy on a balanced sample says nothing about accuracy on a stream where one class is rare. A rule calling everything sound would score on production and be useless. What the figure supports: a comparison of candidate directions on that data. What it does not support: any statement about performance on future components, which needs error on observations not used in fitting. (b) A linear boundary cannot work here, for two separate reasons.The class means nearly coincide. against , so , tiny against within-class spreads with traces in the hundreds and thousands. That difference is the numerator of , so the construction is built on almost no signal. Projected onto any direction, the two classes will have nearly the same centre. The covariances are grossly unequal. The scatter traces are and , a ratio of
over sixteen to one. Pooling them yields a matrix describing neither class. The linear boundary is optimal only when both classes share a covariance, and that premise fails by more than an order of magnitude. The mechanism explains both. Defects arise from several unrelated causes pushing readings away from nominal in different directions, so the defective class is a diffuse cloud surrounding the tight sound class, with roughly the same centre. This is the ring-around-a-cloud configuration, and no straight line separates a surrounding shell from its own centre. The was fitted to a particular sample of that shell; as different defect causes appeared over six months, the sampled shell filled out and the fitted line stopped corresponding to anything. (c) Refitting on components does not address it. More data estimates the same means and the same covariances more precisely. The means would remain nearly coincident and the covariance ratio would remain about sixteen, because these are properties of the populations rather than artefacts of the sample size. The linear rule would be fitted more stably to a boundary that cannot separate the classes. The proposal also repeats the sampling error: if the are again balanced, the threshold will again be calibrated for a prevalence the production line does not have. More data is worth collecting, it is needed for honest evaluation and for a per-class covariance estimate, but refitting the same model on it is the part that does not help. (d) Prevalence and costs both move the threshold, in the same direction. The midpoint between projected class means is correct only when the classes are equally likely and the two errors cost the same. Neither holds. Defects are of production, not . On its own, a rarer positive class argues for a threshold less willing to call a component defective. But a missed defect costs about times a false alarm, which argues strongly the other way. Combining them, the relevant comparison weighs prevalence against cost: against , so the expected cost of missing a defect outweighs that of a false alarm by roughly four to one. The threshold should be set to flag substantially more components than a midpoint rule would, accepting many false alarms to avoid misses. Neither figure is computable from the fit. Both are facts about the application, and the current rule encodes the opposite of both. (e) What to do instead.Change the shape of the boundary, or the framing. Given one tight class and one diffuse class surrounding it, the natural alternatives are a rule permitting each class its own covariance, which gives a quadratic boundary and is the direct fix for unequal covariances, or, better suited to the mechanism, treating this as anomaly detection rather than classification: model the sound class alone, which is tight and well characterised, and flag components far from it. That matches how defects actually arise, since "defective" is not one coherent class but several unrelated departures from nominal. Split the estimation from the evaluation. Fit on one portion and estimate error on another never used in fitting. Report performance on a sample reflecting the prevalence, not a balanced one. Report sensitivity and specificity, never accuracy. At prevalence accuracy is dominated by the common class and conceals exactly the errors that matter. Report the miss rate and the false-alarm rate separately, with the threshold that produced them. Set the threshold from prevalence and costs explicitly, stating both, so that the choice is reviewable rather than inherited from a balanced pilot study. Monitor for drift. The failure appeared over six months, which suggests the defect population changed. Whatever rule replaces this one should have its error rates tracked on continuing labelled samples, since no fitted boundary reports its own obsolescence. What I would tell the team. The original figure was measured on the data that produced the rule and on a class balance that does not exist in production, so it never supported deployment. The deeper problem is that the defective class surrounds the sound one rather than sitting beside it, visible in the near-identical means and the sixteen-to-one spread ratio, so no straight boundary was ever going to work, and the six months of apparent success were a property of the particular defects sampled rather than of the method.
A complete answer does each of these:
computes pooled scatter
derives discriminant direction
explains whitening
classifies by projection
detects assumption failure
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.