Skip to content

Risk Modelling and Survival Analysis · Elementary principles of machine learning

Unsupervised Learning: Clustering and Dimension Reduction (k-means, Hierarchical, PCA)

Updated 11 October 2026 · Fact-checked

Unsupervised learning finds structure in data with no target variable. Clustering (k-means, hierarchical) groups similar observations. Principal components analysis (PCA) replaces many correlated variables with a few uncorrelated components that capture most of the variance. To solve questions, define distance, apply the algorithm step by step, then interpret the result.

Understand Unsupervised Learning: Clustering and Dimension Reduction

In supervised learning each observation has a label or response, such as claim size or lapse or not. In unsupervised learning there is no response. You only have the features. The aim is to find structure: groups of similar policyholders, or a few combinations of variables that describe the data well.

Clustering puts observations into groups so that members of a group are close to each other and far from other groups. Closeness is measured by a distance, most often Euclidean distance. Because distance depends on scale, you usually standardise variables first. Otherwise a variable measured in rupees will swamp one measured in years.

k-means needs you to choose the number of clusters k in advance. It starts with k initial centres. It assigns each point to the nearest centre, then moves each centre to the mean of its points. It repeats until assignments stop changing. It minimises the total within-cluster sum of squares, but it can end in a local minimum, so the result depends on the starting centres. Different starts are usually tried.

Hierarchical clustering does not need k in advance. The common agglomerative version starts with each point as its own cluster and repeatedly merges the two closest clusters. Distance between clusters depends on the linkage: single (smallest distance between members), complete (largest), or average. The result is a dendrogram. Cutting it at a chosen height gives a set of clusters.

Principal components analysis (PCA) is dimension reduction. The first principal component is the direction of greatest variance in the data. The second is the direction of greatest remaining variance, at right angles to the first, and so on. Each component is a linear combination of the original variables. Components are uncorrelated. You keep the first few that explain enough of the total variance. Variables should normally be centred, and standardised if their scales differ. The components come from the eigenvectors of the covariance (or correlation) matrix, and the eigenvalues give the variance of each component.

Key rules to remember

Euclidean distance
d(x, y) = √( Σ (xᵢ − yᵢ)² )
Sum over all variables i. Standardise variables first if scales differ.
Cluster centre (centroid)
centroid = (1 ÷ n) Σ of the points in the cluster
Take the mean of each variable separately. k-means recomputes this after every assignment step.
Within-cluster sum of squares
WCSS = Σ over clusters Σ over points in cluster ‖x − centroid‖²
k-means tries to minimise this. It always falls as k rises, so do not choose k by minimum WCSS alone.
Standardisation
z = (x − mean) ÷ standard deviation
Use before clustering or PCA when variables are in different units.
Principal component
PC₁ = φ₁₁X₁ + φ₁₂X₂ + … + φ₁ₚXₚ, with Σ φ₁ⱼ² = 1
The loadings φ are the entries of the eigenvector. They are scaled to unit length.
Proportion of variance explained
PVE of component m = λₘ ÷ Σ λⱼ
λ are eigenvalues of the covariance or correlation matrix. For standardised data, Σ λⱼ = p, the number of variables.

How to solve Unsupervised Learning: Clustering and Dimension Reduction questions

Use this order for any clustering or PCA question, whether it is a calculation or a discussion.

  1. 1Identify the task: grouping observations (clustering) or reducing variables (PCA). State that there is no response variable.
  2. 2Decide whether to standardise. If variables have different units or spreads, say you will, and why.
  3. 3State the distance measure, and for hierarchical clustering the linkage. For k-means state k and the starting centres.
  4. 4Run the algorithm step by step. For k-means: assign, recompute centroids, repeat until no change. For hierarchical: merge the closest pair, update distances, record the merge height. For PCA: use the given eigenvalues or loadings.
  5. 5Check the result: stable assignments, sensible cluster sizes, cumulative variance explained.
  6. 6Interpret in context, for example risk groups of policyholders, and state limitations such as sensitivity to starting points, scale or outliers.

Quickest way: Fast k-means and PCA working

When to use it: Small numerical questions with few points, usually one or two variables, and PCA questions that give eigenvalues.

  1. For k-means, use squared distances. You do not need the square root to find the nearest centre.
  2. Write a small table: point, distance to each centre, assigned cluster. Then recompute means once per row of the table.
  3. Stop as soon as a pass leaves all assignments unchanged.
  4. For hierarchical clustering, write the distance matrix once. After each merge, update only the row and column of the new cluster.
  5. For PCA, divide each eigenvalue by the sum of eigenvalues, then add up the proportions until you reach the required level.

Common mistakes in Unsupervised Learning: Clustering and Dimension Reduction

  • Not standardising variables with different units before clustering or PCA.

    Students focus on the algorithm and forget distance and variance depend on scale.

    Fix: Always say whether you standardise and why. If scales differ, standardise, or use the correlation matrix for PCA.

  • Stopping k-means after one assignment step.

    Students forget that moving centroids can change assignments.

    Fix: After updating centroids, reassign every point. Stop only when no point changes cluster.

  • Saying k-means always finds the best clustering.

    The algorithm converges, so students assume the answer is optimal.

    Fix: State that it converges to a local minimum and depends on starting centres. Try several starts.

  • Choosing k by picking the value with the lowest WCSS.

    WCSS falls every time k increases, so it looks like improvement.

    Fix: Use the elbow idea: look for where the fall in WCSS levels off. Also use judgement about what the groups mean.

  • Treating PCA as variable selection, or components as original variables.

    Students think keeping two components means keeping two of the original variables.

    Fix: Each component is a linear combination of all the variables. PCA creates new variables rather than choosing old ones.

  • Confusing the linkage methods in hierarchical clustering.

    Single, complete and average sound alike.

    Fix: Single uses the minimum distance between members, complete the maximum, average the mean of all pairwise distances.

Worked examples

Example 1

Five observations on one variable: 1, 2, 6, 7, 10. Apply k-means with k = 2 and starting centres 1 and 10. Find the final clusters and the within-cluster sum of squares.

Show the solution
  1. Pass 1, assign to nearest centre. Centre 1 gets 1 and 2. Centre 10 gets 10. For 6: distance to 1 is 5 and to 10 is 4, so it goes to 10. For 7: distance 6 and 3, so it goes to 10. Clusters: {1, 2} and {6, 7, 10}.
  2. Recompute centroids. Cluster A: (1 + 2) ÷ 2 = 1.5. Cluster B: (6 + 7 + 10) ÷ 3 = 23 ÷ 3 = 7.667.
  3. Pass 2, reassign. Point 2: distance to 1.5 is 0.5, to 7.667 is 5.667, stays in A. Point 6: distance to 1.5 is 4.5, to 7.667 is 1.667, stays in B. No point changes, so stop.
  4. WCSS for A: (1 − 1.5)² + (2 − 1.5)² = 0.25 + 0.25 = 0.5.
  5. WCSS for B: (6 − 7.667)² + (7 − 7.667)² + (10 − 7.667)² = 2.778 + 0.444 + 5.444 = 8.667.
  6. Total WCSS = 0.5 + 8.667 = 9.167.

Answer: Final clusters {1, 2} and {6, 7, 10}, with centroids 1.5 and 7.667. Total WCSS ≈ 9.17.

Example 2

A PCA on four standardised variables gives eigenvalues of the correlation matrix of 2.4, 1.0, 0.4 and 0.2. (a) Find the proportion of variance explained by each component. (b) How many components are needed to explain at least 85% of the variance? (c) Explain why the variables were standardised.

Show the solution
  1. The sum of the eigenvalues is 2.4 + 1.0 + 0.4 + 0.2 = 4.0. This equals the number of variables, as expected for standardised data.
  2. (a) PC1: 2.4 ÷ 4 = 0.60. PC2: 1.0 ÷ 4 = 0.25. PC3: 0.4 ÷ 4 = 0.10. PC4: 0.2 ÷ 4 = 0.05.
  3. (b) Cumulative: PC1 gives 60%. PC1 and PC2 give 85%. So two components reach exactly 85%, which meets the requirement of at least 85%.
  4. (c) Variables measured in different units have different variances. Unstandardised PCA would give most weight to the variable with the largest variance, whatever its importance. Standardising puts all variables on equal footing.

Answer: Proportions are 60%, 25%, 10% and 5%. Two components explain 85% of the variance, so two are needed. Standardisation stops large-scale variables dominating the components.

Exam tips

  • Show every iteration of k-means or every merge in hierarchical clustering. Marks are given for the process, not only the final groups.
  • In discussion questions, always link the method to the purpose: grouping policyholders, finding common risk factors, or reducing the number of rating variables.
  • State assumptions: the distance measure, the linkage, whether data were standardised, and the starting centres.
  • Compare methods using clear points: k-means needs k in advance and suits large data sets; hierarchical gives a dendrogram and needs no k, but is heavier to compute on large data.
  • Expect Paper B to ask for these methods in R. Know which function does what and how to read the output, such as cluster sizes, centres and variance explained.

Practice questions from Elementary principles of machine learning

Unsupervised Learning: Clustering and Dimension Reduction in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Unsupervised Learning: Clustering and Dimension Reduction: frequently asked questions

What is the difference between hierarchical and k-means clustering?

k-means needs you to choose k first and gives one flat set of clusters by moving centroids until assignments settle. Hierarchical clustering builds a tree of merges called a dendrogram and you choose the number of clusters afterwards by cutting it. k-means copes better with large data sets.

How does k-means clustering work step by step?

Choose k and starting centres. Assign each point to its nearest centre. Recompute each centre as the mean of its assigned points. Repeat the assignment and update steps until no point changes cluster.

What is PCA in simple terms?

PCA turns many correlated variables into a smaller set of uncorrelated components. The first component captures the most variance, the second the most of what remains, and so on. You keep enough components to describe the data well.

How do I choose the number of clusters or components?

For clusters, look for an elbow in the plot of WCSS against k, or cut a dendrogram where there is a large jump in merge height. For PCA, keep components until the cumulative variance explained reaches a target, or use a scree plot. All of these need judgement.

Is unsupervised learning examined in both papers?

It sits within the machine learning part of CS2. You should be ready for short written explanations and small calculations in Paper A, and for running and interpreting the methods in R in Paper B.