FRM Exam Part I · Machine-Learning Methods
Unsupervised Learning: Clustering and PCA Explained for FRM Part 1
Updated 11 October 2026 · Fact-checked
Unsupervised learning finds structure in data that has no labels. Clustering (k-means, hierarchical) groups similar observations. Principal components analysis (PCA) shrinks many correlated variables into a few uncorrelated components that capture most of the variance. Clustering groups rows; PCA compresses columns.
Understand Unsupervised Learning: Clustering and PCA
In supervised learning you have a target, such as default or no default. In unsupervised learning you have only features. The goal is to find patterns, groups or simpler representations without being told the right answer.
K-means clustering splits observations into k groups. You choose k first. The algorithm picks k starting centroids, assigns each point to the nearest centroid (usually by Euclidean distance), recomputes each centroid as the mean of its points, and repeats until assignments stop changing. The result can depend on the starting centroids, so you run it several times. Features should be scaled first, because a variable with large units dominates the distance.
Choosing k. The elbow method plots the within-cluster sum of squares (inertia) against k. Inertia always falls as k rises. You pick the k where the drop flattens, the elbow. The silhouette score is another check. It compares how close a point is to its own cluster against the nearest other cluster. Values near 1 mean well separated.
Hierarchical clustering does not need k in advance. Agglomerative (bottom-up) starts with each point alone and repeatedly merges the two closest clusters. Divisive (top-down) starts with one cluster and splits it. The merges are shown in a dendrogram. The height of a merge is the distance between the clusters joined. You cut the dendrogram at a chosen height to get clusters. The linkage rule defines cluster distance: single (nearest pair), complete (farthest pair), average, or Ward (minimizes the increase in variance).
Principal components analysis is a dimension-reduction tool. It finds new axes, the principal components, that are linear combinations of the original variables. The first component captures the largest share of variance. Each next component captures the most remaining variance and is uncorrelated with the earlier ones. Each eigenvalue is the variance of its component, and eigenvalue ÷ sum of eigenvalues is the share of variance explained. You keep enough components to explain a target share, for example 90%. In risk work, PCA on yield curves typically yields level, slope and curvature factors. Scale variables first if units differ.
Key formulas to remember
- Euclidean distance
- d(x, y) = √[Σ (xᵢ − yᵢ)²]
- Standard distance for k-means. Scale features first.
- Centroid
- centroid = average of each feature across the points in the cluster
- Recomputed after every assignment step in k-means.
- Within-cluster sum of squares (inertia)
- WCSS = Σ over clusters Σ over points ‖x − centroid‖²
- K-means minimizes this. It falls as k rises, so use the elbow, not the minimum.
- Variance explained by a component
- share of component j = λⱼ ÷ Σ λ
- λ are eigenvalues of the covariance (or correlation) matrix.
- Cumulative variance explained
- (λ₁ + … + λₘ) ÷ Σ λ
- Use to decide how many components to keep.
- Principal component
- PCⱼ = w₁ⱼX₁ + w₂ⱼX₂ + … + wₙⱼXₙ
- Weights (loadings) form a unit-length eigenvector. Components are uncorrelated.
How to solve Unsupervised Learning: Clustering and PCA questions
Use this routine for any clustering or PCA question.
- 1Decide the task: grouping observations (clustering) or compressing variables (PCA).
- 2Check the data: are features on different scales? If so, standardize before distance or PCA.
- 3For k-means, compute distances to each centroid, assign each point to the nearest, then recompute centroids as means.
- 4For hierarchical clustering, find the smallest distance between clusters under the stated linkage, merge, and read heights on the dendrogram.
- 5For choosing k, look for the elbow in WCSS or the highest silhouette score.
- 6For PCA, divide each eigenvalue by the total to get variance shares, then add them until the target is reached.
- 7Check the answer: shares sum to 100%, centroids lie within the data range, and the result matches the question's wording.
Quickest way: Fast checks for clustering and PCA questions
When to use it: When a question gives eigenvalues, a small data set, or a conceptual statement to judge.
- For PCA, sum the eigenvalues once, then divide. If the variables are standardized, the total equals the number of variables.
- For one k-means step, compute only the distances needed. Compare squared distances to avoid square roots.
- For a dendrogram, the higher the merge, the less similar the clusters.
- For conceptual options, eliminate any that say PCA uses labels, k-means needs no k, or inertia rises with k.
Common mistakes in Unsupervised Learning: Clustering and PCA
Choosing k as the value that minimizes WCSS.
WCSS falls with every extra cluster, so it seems the lower the better.
Fix: WCSS is lowest when k equals the number of points. Pick the elbow or the best silhouette score.
Forgetting to scale features.
The algorithm runs without error on raw data.
Fix: Standardize first. Otherwise large-unit variables dominate distance in k-means and variance in PCA.
Treating PCA as a prediction model with labels.
PCA is often used before regressions, so it is mixed up with supervised methods.
Fix: PCA uses only the features. It is unsupervised and does not look at any target.
Reading a dendrogram by horizontal position instead of merge height.
Leaf order on the chart looks meaningful.
Fix: Only the height of a merge shows distance. Cut at a height to choose the number of clusters.
Saying k-means always gives the same answer.
Students assume the algorithm is deterministic.
Fix: Results depend on starting centroids and can reach a local optimum. Rerun with different starts.
Confusing a component's variance share with a correlation.
Both are reported as percentages.
Fix: Share explained is λⱼ ÷ Σλ. It says how much total variance the component carries.
Worked examples
Example 1
A standardized data set has four variables. The PCA eigenvalues are 2.4, 1.0, 0.4 and 0.2. What share of total variance do the first two components explain together, and how many components are needed to explain at least 90%?
Show the solution
- Total variance = 2.4 + 1.0 + 0.4 + 0.2 = 4.0, which equals the number of standardized variables.
- PC1 share = 2.4 ÷ 4.0 = 60%.
- PC2 share = 1.0 ÷ 4.0 = 25%.
- First two together = 60% + 25% = 85%.
- Add PC3: 0.4 ÷ 4.0 = 10%, so cumulative = 95%.
- 85% is below 90% and 95% is above it, so three components are needed.
Answer: The first two explain 85%; three components are needed to reach at least 90%.
Example 2
K-means with k = 2 has centroids A = (1, 1) and B = (5, 5). Point P = (2, 3) and point Q = (4, 4) are observed. Assign each point to a cluster, then compute the new centroid of the cluster that contains both P and Q if they both join the same cluster and no other points exist.
Show the solution
- Squared distance P to A = (2−1)² + (3−1)² = 1 + 4 = 5.
- Squared distance P to B = (2−5)² + (3−5)² = 9 + 4 = 13. P goes to A.
- Squared distance Q to A = (4−1)² + (4−1)² = 9 + 9 = 18.
- Squared distance Q to B = (4−5)² + (4−5)² = 1 + 1 = 2. Q goes to B.
- So P and Q do not share a cluster, and the premise fails. Recompute centroids from the actual assignment.
- Cluster A holds only P, so its new centroid is (2, 3).
- Cluster B holds only Q, so its new centroid is (4, 4).
Answer: P joins cluster A and Q joins cluster B. The new centroids are (2, 3) and (4, 4).
Exam tips
- Expect conceptual questions: unsupervised vs supervised, what k-means needs, what a dendrogram shows, why scaling matters.
- PCA numeric questions are usually eigenvalue division. Do the sum first and keep fractions simple.
- Know the elbow method cold: WCSS always falls with k, and you look for where the fall flattens.
- Remember the finance link: PCA on yield curves gives level, slope and curvature, with level explaining the most variance.
- Watch for the contrast question: clustering groups observations; PCA reduces the number of variables.
Practice questions from Machine-Learning Methods
- A node in a classification tree contains 40 loans, of which 20 defaulted. A candidate split sends 20 loans to the left child with 2 defaults…
- A node in a classification tree holds 40 observations: 30 non-defaults and 10 defaults. Using the Gini impurity, 1 minus the sum of squared …
- An analyst has 10,000 observations to forecast corporate bond downgrades. She first standardizes all features using the mean and standard de…
- A model predicting loan losses achieves a mean squared error of 0.5 on the training sample but 4.0 on a held-out validation sample. Which in…
- A risk analyst fits a single unpruned decision tree to predict loan default and finds that it classifies the training data almost perfectly …
Unsupervised Learning: Clustering and PCA in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Unsupervised Learning: Clustering and PCA: frequently asked questions
What is the difference between PCA and clustering?
Clustering groups similar observations into clusters. PCA builds a few new uncorrelated variables from many correlated ones to reduce dimensions. Both are unsupervised, but clustering works on rows and PCA works on columns.
How do I choose the number of clusters in k-means?
Plot within-cluster sum of squares against k and pick the elbow where improvement flattens. You can also compare silhouette scores. Business judgment on usefulness also matters.
Does hierarchical clustering need the number of clusters in advance?
No. You build the full dendrogram and then cut it at a chosen height to get the clusters. A lower cut gives more clusters and a higher cut gives fewer.
Why must I scale data before PCA or k-means?
Both methods are driven by variance or distance. A variable measured in large units would dominate the result. Standardizing puts all variables on equal footing.