Skip to content

FRM Exam Part I · Machine Learning and Prediction

K-Means and Hierarchical Clustering Explained for FRM Part I

Updated 11 October 2026 · Fact-checked

Clustering is unsupervised learning that groups similar observations without labels. K-means splits data into a preset number K of clusters by minimising within-cluster squared distance to centroids. Hierarchical clustering builds a tree (dendrogram) by merging or splitting groups. You choose K with the elbow method or silhouette score.

Understand Clustering: K-Means and Hierarchical Methods

Clustering is a type of unsupervised learning. You give the algorithm data with no labels and it finds groups of observations that look alike. In risk work you might group borrowers with similar profiles, or funds with similar return behaviour.

Everything starts with a distance measure. The most common is Euclidean distance, the straight-line distance between two points. Manhattan distance adds absolute differences instead. Because distance depends on scale, you must standardise features first. A feature measured in millions will otherwise swamp one measured in percentages.

K-means needs you to pick K in advance. It places K centroids, assigns each point to the nearest centroid, then moves each centroid to the mean of its points. It repeats until assignments stop changing. The goal is to minimise the within-cluster sum of squares (WCSS). The result can depend on the starting centroids, so it is run several times with different starts. It works best for compact, roughly round clusters and is sensitive to outliers.

Hierarchical clustering does not need K up front. Agglomerative (bottom-up) starts with every point alone and repeatedly merges the two closest clusters. Divisive (top-down) starts with one cluster and splits it. The merge history is drawn as a dendrogram. You cut the tree at a chosen height to get a number of clusters. The distance between clusters depends on the linkage: single (closest pair), complete (farthest pair), average, or Ward (smallest increase in within-cluster variance).

To choose K, use the elbow method: plot WCSS against K and pick the point where the drop flattens. WCSS always falls as K rises, so you look for diminishing returns, not the minimum. The silhouette score measures how well each point fits its own cluster versus the nearest other one. It ranges from −1 to 1, and higher is better.

Key formulas to remember

Euclidean distance
d(x, y) = √[Σ (xᵢ − yᵢ)²]
Sum over all features i. Standardise features first.
Manhattan distance
d(x, y) = Σ |xᵢ − yᵢ|
Less sensitive to a single large difference than Euclidean.
Centroid
centroid = (1 ÷ n) × Σ of the points in the cluster
The mean of each feature across the cluster's members.
Within-cluster sum of squares
WCSS = Σ over clusters Σ over points ‖x − centroid‖²
K-means minimises this. It always falls as K increases.
Silhouette coefficient
s = (b − a) ÷ max(a, b)
a = average distance to own cluster; b = average distance to nearest other cluster. Range −1 to 1.

How to solve Clustering: K-Means and Hierarchical Methods questions

Use this sequence for any clustering question, whether conceptual or numerical.

  1. 1Identify the task: no labels means unsupervised clustering. Labels would point to classification.
  2. 2Check feature scaling. If features have very different units, standardise before computing distances.
  3. 3Name the algorithm: K needed in advance means K-means; a tree or dendrogram means hierarchical.
  4. 4For K-means numbers, assign each point to the nearest centroid using the stated distance, then recompute each centroid as the mean.
  5. 5For hierarchical numbers, find the smallest distance between clusters under the stated linkage and merge those two. Repeat.
  6. 6To choose K, look for the elbow in WCSS or the highest silhouette score. Do not pick the K with the lowest WCSS.
  7. 7Read a dendrogram by height: the vertical axis is the distance at which clusters merged. Cutting at a height gives the clusters below it.
  8. 8State the limitation the question hints at: outliers, non-round clusters, or sensitivity to starting centroids.

Quickest way: Spot the algorithm and the trap

When to use it: Conceptual multiple-choice questions where you must compare methods or pick the true statement.

  1. Ask: is K fixed in advance? Yes means K-means, no means hierarchical.
  2. Ask: bottom-up or top-down? Bottom-up is agglomerative, top-down is divisive.
  3. Eliminate any option that claims clustering uses labels or predicts a target.
  4. Eliminate any option that says WCSS is minimised by the largest K.
  5. For numbers, compute only the distances you need, using squared distance when you only compare.

Common mistakes in Clustering: K-Means and Hierarchical Methods

  • Treating clustering as supervised learning.

    Groups sound like classes, so students assume labels exist.

    Fix: Clustering finds groups from the data alone. If the question mentions known labels, it is classification.

  • Choosing the K that gives the lowest WCSS.

    Lower error looks better.

    Fix: WCSS reaches zero when K equals the number of points. Use the elbow or the silhouette score.

  • Forgetting to standardise features.

    Students focus on the algorithm and ignore units.

    Fix: Large-scale features dominate distance. Standardise first whenever units differ.

  • Saying K-means gives the same answer every run.

    It looks like a deterministic procedure.

    Fix: Results depend on the initial centroids and can reach a local optimum. Run it with several starts.

  • Confusing agglomerative and divisive.

    The names are similar and rarely used in daily work.

    Fix: Agglomerative accumulates: bottom-up merging. Divisive divides: top-down splitting.

  • Reading dendrogram height as the number of points.

    Students look at branch length without checking the axis.

    Fix: Height is the distance or dissimilarity at which a merge happened. A tall merge means the clusters were far apart.

Worked examples

Example 1

A K-means run has two centroids, A = (1, 2) and B = (5, 6). A point P = (2, 4) is to be assigned using Euclidean distance. Which cluster does it join, and what is the new centroid of that cluster if it previously held only the points (1, 2) and (3, 6)?

Show the solution
  1. Distance to A: √[(2 − 1)² + (4 − 2)²] = √(1 + 4) = √5 ≈ 2.236.
  2. Distance to B: √[(2 − 5)² + (4 − 6)²] = √(9 + 4) = √13 ≈ 3.606.
  3. P is closer to A, so it joins cluster A.
  4. Cluster A now has points (1, 2), (3, 6) and (2, 4).
  5. New centroid x = (1 + 3 + 2) ÷ 3 = 2.
  6. New centroid y = (2 + 6 + 4) ÷ 3 = 4.

Answer: P joins cluster A, and the new centroid of A is (2, 4).

Example 2

Five loans are clustered by agglomerative clustering with single linkage. The one-dimensional risk scores are 1, 2, 6, 8 and 15. Find the first three merges and the merge distances.

Show the solution
  1. Start with five singleton clusters: {1}, {2}, {6}, {8}, {15}.
  2. Smallest gap is between 1 and 2, distance 1. Merge to {1, 2}.
  3. Under single linkage, the distance from {1, 2} to {6} is the closest pair: 6 − 2 = 4. The gap between 6 and 8 is 2, which is smaller.
  4. Second merge: {6} and {8} at distance 2, giving {6, 8}.
  5. Now the candidates: {1, 2} to {6, 8} is 6 − 2 = 4. {6, 8} to {15} is 15 − 8 = 7. {1, 2} to {15} is 13.
  6. Third merge: {1, 2} and {6, 8} at distance 4, giving {1, 2, 6, 8}.

Answer: Merges: {1, 2} at 1, {6, 8} at 2, then {1, 2, 6, 8} at 4. The point 15 joins last, at distance 7.

Exam tips

  • Expect conceptual questions: unsupervised versus supervised, K fixed versus not, bottom-up versus top-down.
  • If a question gives a WCSS table by K, look for where the improvement drops sharply, not the smallest value.
  • Always check whether the question says to standardise before measuring distance.
  • For hand calculations, compare squared distances to save time, since the order is the same.
  • Know the limits: K-means struggles with outliers and non-spherical clusters, and hierarchical methods become costly on large datasets.

Practice questions from Machine Learning and Prediction

Clustering: K-Means and Hierarchical Methods: frequently asked questions

What is the main difference between K-means and hierarchical clustering?

K-means needs you to set K in advance and iteratively updates centroids. Hierarchical clustering builds a full tree of merges or splits, and you pick the number of clusters afterwards by cutting the dendrogram. K-means is usually faster on large datasets.

How do I choose the number of clusters with the elbow method?

Run K-means for several values of K and plot WCSS against K. WCSS falls as K rises. Choose the K where the curve bends and further clusters add little improvement.

What is the difference between agglomerative and divisive clustering?

Agglomerative clustering starts with each point alone and merges the closest clusters step by step. Divisive clustering starts with all points in one cluster and splits it repeatedly. Agglomerative is the more common of the two.

Do I need to scale data before clustering?

Yes, in most cases. Distance measures are sensitive to units, so a feature with large values will dominate. Standardising puts features on a comparable scale.