Practice: Clustering, and What the Objective Assumes
Question
Direct application
Six points:
(a) Carry out the assignment step, showing which centroid each point is nearest.
(b) Carry out the update step, giving both new centroids.
(c) Run the assignment step again and state whether the algorithm has converged, saying what condition you checked.
(d) Report the final within-cluster sum of squares.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Compare squared distances; there is no need to take square roots, since the ordering is the same.
Hint 2: Next step
After updating, run assignment once more. The run ends when that step changes nothing.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) Assignment.
The two clusters of three sit around
Checking one explicitly, for
Labels:
(b) Update.
Cluster 0 mean:
Cluster 1 mean:
(c) Second assignment, and the stopping condition.
Each centroid moved only within its own group, and the groups remain separated by roughly
The algorithm has converged, and the condition checked is that the labels did not change between successive assignment steps, not that a fixed number of passes was completed. Stopping at a count would leave it unsettled whether this is a stable partition or a snapshot of one still moving.
(d) Final WCSS.
Cluster 0, distances to
: : :
Subtotal
Cluster 1, distances to
A complete answer does each of these:
- executes assignment update
Direct application
Six points:
Run k-means with
Report the final within-cluster sum of squares, exactly.
Enter the value. It is checked against the answer and the precision this task asks for.
2 hints available, least help first.
Hint 1: Retrieval cue
The centroid of a cluster is the coordinatewise mean of its members.
Hint 2: Next step
Once the labels stop changing, sum squared distances from each point to the centroid of its own cluster.
Direct application
On a nine-point dataset with three groups, exhaustive enumeration over all partitions establishes that the best three-cluster solution has within-cluster sum of squares
By what factor is the converged run worse than the optimum? Give the ratio to four decimal places.
Enter the value. It is checked against the answer and the precision this task asks for.
2 hints available, least help first.
Hint 1: Retrieval cue
k-means minimises WCSS, so a larger value is a worse solution.
Hint 2: Next step
The factor is the converged value divided by the optimal value.
Error diagnosis · Explanation
A colleague runs k-means with
(a) Explain how both runs can have converged, given that one is
(b) Your colleague suspects an arithmetic error in one of the runs. Say why that is the wrong diagnosis, and what actually produced the difference.
(c) State what should have been done before reporting, and what makes the resulting claim checkable by a reader.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Write down exactly what has to be true for k-means to stop.
Hint 2: Concept cue
Ask what would have to happen for the worse partition to escape, and whether the update step can arrange it.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) Both satisfy the stopping condition. Convergence in k-means means two things, and both hold in each run: no single point lowers the objective by switching groups, and every centroid sits at the mean of its assigned points. These are local conditions, statements about what one point or one centroid can do alone. The worse run reaches a partition where two centroids share the distant group at
A complete answer does each of these:
- detects local optimum
- executes assignment update
Interpretation · Comparison
Two datasets, each clustered at every
Dataset P, WCSS:
Dataset Q, WCSS:
(a) Tabulate the successive drops for each dataset, as absolute reductions and as percentages of the previous value.
(b) For each dataset, say what number of clusters the curve supports, or say that it supports none, and justify the answer from the drop sequence.
(c) Both curves decrease at every step. Explain why that fact alone carries no information about whether either dataset contains groups.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Compute each drop as a percentage of the value it fell from.
Hint 2: Concept cue
Ask what the curve would look like for data with no groups at all, and whether you could tell it apart from what you are given.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) The drops. Dataset P: |
|---|---|---|---|
| 1 |
| 2 |
| 3 |
| 4 |
| 5 |
| 6 |
|---|---|---|---|
| 1 |
| 2 |
| 3 |
| 4 |
| 5 |
| 6 |
| 7 |
A complete answer does each of these:
- reads wcss curve
Error diagnosis · Method selection
Thirty-six points lie on two concentric rings: twelve at radius
The analyst concludes that k-means converged to a local optimum and proposes running more restarts.
(a) The by-ring partition scores
(b) Explain what actually caused the result, referring to which partitions the objective scores well.
(c) Single linkage on the same points with
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Where is the mean of a ring of points?
Hint 2: Concept cue
k-means minimises WCSS. Which of the two partitions has the lower value, and what follows about restarts?
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) The diagnosis is wrong, and the two figures settle it. The returned partition scores
A complete answer does each of these:
- attributes failure to objective
- selects method for structure
Interpretation · Evaluation
A marketing team clusters customer records into
"The analysis identified three customer segments. Segment 1 (
) are younger, lower-spending occasional buyers. Segment 2 ( ) are mid-career, steady, mid-value. Segment 3 ( ) are older, high-value, infrequent. The segmentation was stable: repeated runs returned the same three groups. We recommend three distinct campaigns."
The data were clustered on raw age (years), annual spend (euros) and purchase count.
(a) The report offers stability as evidence that the segments are real. Say what stability does and does not establish.
(b) Identify what the report omits that a reader would need in order to disagree with it.
(c) The team asks whether three is the right number of segments. Say what could settle that and what could not.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Euros and years are added together inside the same squared distance. What does that imply about their influence?
Hint 2: Concept cue
Ask what the report would look like if the data had no segments at all.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) What stability establishes. Stability means the algorithm reliably finds the same partition from different starting points. That is a fact about the objective's landscape, that one solution has a wide basin of attraction, and it is genuinely useful: it rules out the local-optimum problem, where different runs return solutions of different quality. It establishes nothing about whether the segments correspond to real distinctions among customers. A method given
A complete answer does each of these:
- reads wcss curve
- attributes failure to objective
Direct application · Interpretation
Four customers, recorded as (age in years, annual income in euros):
| Age | Income | |
|---|---|---|
| A | ||
| B | ||
| C | ||
| D |
(a) Clustering with
(b) Standardise each column to a z-score using the population standard deviation. Give the standardised coordinates, then state which partition is optimal and its WCSS.
(c) The two partitions disagree completely. Say what decided the outcome in each case, and what the analyst is actually choosing when they decide whether to standardise.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Compare the age range with the income range, then compare their squares.
Hint 2: Next step
For the z-scores, use the population standard deviation: divide by
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) Why the raw partition groups by income. Squared Euclidean distance adds the squared age difference and the squared income difference directly, as though a year and a euro were the same unit. The age range here is
The by-age partition loses by a factor of about
|---|---|---|
| A |
| B |
| C |
| D |
The by-income partition
|---|---|---|
| by-income partition |
| by-age partition |
A complete answer does each of these:
- attributes failure to objective
Transfer · Evaluation · Explanation
An engineering team monitors a fleet of pumps. Each pump reports two standardised readings once a minute: vibration and temperature. Healthy pumps cycle through a repeating operating loop, so their readings trace a closed band at a roughly constant distance from the fleet's average operating point. Genuine faults show up as readings that drift toward the middle of that loop. A pump running unusually cool and still.
The team clusters a day of readings with k-means,
Write a review covering:
(a) Whether k-means with this objective can detect the faults described, and why, supporting the answer with which partitions the objective scores well and where the healthy readings' mean lies.
(b) What the single run leaves unestablished, what would establish it, and how large the risk is in general terms.
(c) Whether "no anomalous regime detected" is a finding, given how
(d) What the team's stability claim does and does not support.
(e) What you would run instead, what its criterion scores well, and what it would cost them.
Write your answer, then compare it with the worked solution.
3 hints available, least help first.
Hint 1: Retrieval cue
Where does the mean of a closed band of readings sit relative to the readings themselves?
Hint 2: Concept cue
Ask what WCSS rewards, then ask how a point near the band's centre scores under it.
Hint 3: Strategy cue
Separate the three failures, the search, the supplied
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) It cannot, and the geometry is the reason. Healthy readings trace a closed band at a roughly constant distance from the fleet's average operating point. A ring. The mean of a ring lies in the hole at its centre, where no healthy data sits. WCSS scores best on partitions whose points lie close to their cluster's mean, so with
- New parameters that are equally consequential. DBSCAN's neighbourhood radius and minimum-points threshold replace
- No centroids to describe. Density-based clusters have no representative mean, so the tidy per-regime summaries in the current report would not survive the change.
A complete answer does each of these:
- reads wcss curve
- attributes failure to objective
- selects method for structure
Construction · Direct application · Explanation
Six points on a line:
(a) Starting from centroids at
(b) Starting instead from
(c) Both runs converged. Say what convergence establishes and what it does not, and what would detect a worse solution.
Write your answer, then compare it with the worked solution.
3 hints available, least help first.
Hint 1: Retrieval cue
Assign each point to the nearer centroid, then move each centroid to its members' mean.
Hint 2: Concept cue
Stop when the labels repeat, not after a fixed number of rounds.
Hint 3: Strategy cue
In (c), ask what the algorithm checked before halting, and what it never checked.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) From centroids
Round 1, assign. Point
Round 1, update. Centroids become
Round 2, assign. Midpoint
Round 2, update. Centroids become
Round 3, assign. Midpoint
Final partition
(b) From centroids
Both runs agree here, which is what well-separated data does. The instructive case is the one that does not: on data whose groups sit closer together, two centroids can settle inside one natural group while the other is absorbed whole, and no single point improves the objective by moving alone.
(c) What convergence establishes. Exactly two things: no single point lowers the objective by changing clusters, and every centroid is at the mean of its members. That is a local optimum.
What it does not establish is that the partition is the best available. The algorithm halts because the objective stopped falling, not because it reached the lowest value; the assignments are finite and the search follows a path determined entirely by where it started. A converged run reports a stopping condition, not a proof.
What would detect a worse solution. Multiple restarts from different initialisations, compared by objective value: the lower WCSS is the better solution found. That reduces sensitivity to initialisation and can find a better partition, but establishes no guarantee of global optimality, which would require exhaustive or global optimisation. A defensible report therefore states the objective value and the number of restarts behind it.
A complete answer does each of these:
- executes assignment update
- detects local optimum
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.