CFA Level II Exam · Machine Learning
PCA and Clustering: Unsupervised Learning for CFA Level II
Updated 7 October 2026 · Fact-checked
Unsupervised learning finds structure in data with no target variable. PCA reduces many correlated features to a few uncorrelated principal components that keep most of the variance. Clustering groups similar observations: k-means needs k set in advance, while hierarchical clustering builds a dendrogram by merging (agglomerative) or splitting (divisive).
Understand Unsupervised Learning: PCA and Clustering
In unsupervised learning there is no labelled outcome to predict. The algorithm looks only at the features and finds patterns. Two tasks matter for the exam: dimension reduction (PCA) and clustering (k-means and hierarchical).
Principal components analysis (PCA) takes many correlated features and builds new variables called principal components. Each component is a linear combination of the original features. The components are uncorrelated with each other (orthogonal). The first component (PC1) captures the largest share of total variance. PC2 captures the most of what remains, and so on. You keep only the first few components and drop the rest. This gives a simpler dataset with less noise and a lower risk of overfitting.
Two ideas help you read PCA output. Eigenvectors define the direction of each component, and eigenvalues measure how much variance each component explains. The scree plot shows the proportion of variance explained by each component. You choose how many components to keep by looking for the point where extra components add little, or by reaching a target cumulative variance. The drawback is interpretability: components are blends of features and often have no clear economic meaning. PCA is applied to features, so it is not a prediction method on its own.
Clustering groups observations so that those in the same cluster are similar and those in different clusters are different. Similarity is measured by a distance, most often Euclidean distance (straight-line distance). Features should usually be on comparable scales, or large-scale features will dominate the distance.
K-means splits the data into k non-overlapping clusters. You choose k first. The algorithm picks k starting centroids, assigns each point to the nearest centroid, recomputes each centroid as the mean of its points, and repeats until assignments stop changing. Results depend on the starting centroids, so you often run it several times. It is fast and suits large datasets.
Hierarchical clustering builds a nested tree of clusters shown as a dendrogram. Agglomerative (bottom-up) starts with each observation as its own cluster and merges the closest pairs step by step. Divisive (top-down) starts with one cluster and splits it. You do not choose the number of clusters in advance. You cut the dendrogram at a chosen height. Agglomerative methods are less computationally heavy than divisive in practice and are more common, but both are slower than k-means on big data.
Key formulas to remember
- Proportion of variance explained by a component
- Proportion for PCj = eigenvalue of PCj ÷ Σ of all eigenvalues
- Eigenvalues sum to total variance. The proportions across all components sum to 100%.
- Cumulative variance explained
- Cumulative proportion for first m components = Σ (proportions of PC1 to PCm)
- Use this to decide how many components to keep for a target level such as 90%.
- Euclidean distance between two points
- d = √[(x1 − y1)² + (x2 − y2)² + … + (xn − yn)²]
- The usual similarity measure in clustering. Smaller distance means more similar.
- K-means centroid update
- New centroid = average of each feature across all points assigned to the cluster
- Repeat assignment and update until no point changes cluster.
- PCA properties
- Components are uncorrelated; PC1 variance ≥ PC2 variance ≥ PC3 variance …
- Eigenvalues are ordered from largest to smallest.
How to solve Unsupervised Learning: PCA and Clustering questions
Use this order for any PCA or clustering item. Read the vignette for what the analyst is trying to do, then match the tool.
- 1Identify the task: reduce many features (PCA) or group observations (clustering). Check there is no target variable, which confirms unsupervised learning.
- 2For PCA, find the eigenvalues or the scree plot data in the exhibit. Compute each proportion as eigenvalue ÷ total.
- 3Add the proportions in order from PC1 to get cumulative variance. Keep the fewest components that reach the stated target.
- 4For k-means, check that k is given and note the starting centroids. Compute Euclidean distance from each point to each centroid and assign to the nearest.
- 5Recompute centroids as feature means of each cluster, and say whether assignments would change. If they do not, the algorithm has converged.
- 6For hierarchical clustering, read the dendrogram from the bottom for agglomerative merges and from the top for divisive splits. Cutting at a lower height gives more clusters.
- 7Match the answer to the stated limitation or advantage: PCA components are hard to interpret; k-means needs k in advance; hierarchical suits smaller datasets.
Quickest way: Four-second tool match and variance check
When to use it: Use when the question asks which method fits, which statement is correct, or how many components to keep.
- Reduce features, uncorrelated outputs, variance: PCA.
- Fixed k, centroids, large dataset: k-means.
- Dendrogram, no preset number of clusters, merging or splitting: hierarchical.
- For components, divide each eigenvalue by the sum, then keep adding from the largest until the target is reached.
- Eliminate options that say PCA is supervised, components are correlated, or k-means chooses k itself.
Common mistakes in Unsupervised Learning: PCA and Clustering
Treating PCA as a method that selects the best original features.
Dimension reduction sounds like feature selection.
Fix: PCA creates new variables that are combinations of all original features. It does not drop original features one by one.
Forgetting that principal components are uncorrelated.
Students link PCA with correlated inputs and assume the outputs are correlated too.
Fix: Inputs are correlated; components are orthogonal and uncorrelated with each other.
Using the eigenvalue instead of the proportion when asked how much variance a component explains.
Rushing past the exhibit.
Fix: Divide the eigenvalue by the sum of all eigenvalues. Then check if the question wants a cumulative figure.
Saying k-means finds the number of clusters on its own.
Confusing k-means with a dendrogram, where you choose the cut afterwards.
Fix: In k-means you set k before running. Hierarchical clustering does not require k in advance.
Mixing up agglomerative and divisive.
Both build a hierarchy and the names are similar.
Fix: Agglomerative is bottom-up: many clusters merge into one. Divisive is top-down: one cluster splits into many.
Ignoring feature scale when computing distance.
Students assume the numbers are directly comparable.
Fix: A feature measured in large units dominates Euclidean distance. If the vignette mentions different scales, scaling is the concern.
Worked examples
Example 1
An analyst applies PCA to six correlated valuation ratios. The eigenvalues of the components, from largest to smallest, are 3.0, 1.5, 0.6, 0.45, 0.3 and 0.15. Q1: What proportion of total variance does PC1 explain? Q2: How many components are needed to explain at least 85% of total variance? Q3: Which statement about the components is correct? A) PC2 is correlated with PC1. B) The components are uncorrelated with each other. C) PCA requires a labelled target variable.
Show the solution
- Total variance = 3.0 + 1.5 + 0.6 + 0.45 + 0.3 + 0.15 = 6.0.
- Q1: PC1 proportion = 3.0 ÷ 6.0 = 50%.
- Q2: PC2 proportion = 1.5 ÷ 6.0 = 25%, so cumulative for two components = 75%. PC3 = 0.6 ÷ 6.0 = 10%, so cumulative for three = 85%.
- Three components reach exactly 85%, which meets 'at least 85%'. Two components give only 75%, which is too low.
- Q3: PCA components are orthogonal, so they are uncorrelated. PCA is unsupervised, so C is wrong. A is wrong.
Answer: Q1: 50%. Q2: Three components. Q3: B, the components are uncorrelated with each other.
Example 2
A portfolio manager clusters four stocks using two standardised features, (earnings growth, leverage): A (1, 1), B (2, 1), C (6, 5), D (7, 5). K-means is run with k = 2 and starting centroids at A (1, 1) and D (7, 5). Q1: Which cluster does C join first? Q2: What are the centroids after the first update? Q3: Will assignments change in the next pass?
Show the solution
- Q1: Distance from C (6, 5) to centroid A (1, 1) = √(25 + 16) = √41 ≈ 6.40. Distance to centroid D (7, 5) = √(1 + 0) = 1. C is nearer D, so C joins D's cluster.
- Assign the others: B (2, 1) to A: √(1 + 0) = 1. B to D: √(25 + 16) ≈ 6.40. B joins A's cluster. A stays with A, D stays with D.
- Clusters: {A, B} and {C, D}.
- Q2: Centroid 1 = ((1 + 2) ÷ 2, (1 + 1) ÷ 2) = (1.5, 1). Centroid 2 = ((6 + 7) ÷ 2, (5 + 5) ÷ 2) = (6.5, 5).
- Q3: Check B: distance to (1.5, 1) = 0.5; to (6.5, 5) = √(20.25 + 16) ≈ 6.02. B stays. Check C: distance to (6.5, 5) = 0.5; to (1.5, 1) = √(20.25 + 16) ≈ 6.02. C stays. A and D are even clearer, so nothing changes.
Answer: Q1: C joins D's cluster. Q2: Centroids are (1.5, 1) and (6.5, 5). Q3: No, assignments do not change, so the algorithm has converged.
Exam tips
- Expect a scree plot or an eigenvalue table. Practise turning eigenvalues into proportions and cumulative percentages quickly.
- Know the contrast list cold: PCA is unsupervised and produces uncorrelated components; k-means needs k in advance; hierarchical gives a dendrogram with no preset k.
- For dendrograms, decide first whether the question is agglomerative (bottom-up) or divisive (top-down), then read the heights.
- Watch for interpretability: PCA components are hard to explain economically, which is its main stated weakness.
- There is no penalty for wrong answers, so never leave an item blank. Eliminate options that call PCA or clustering supervised.
Unsupervised Learning: PCA and Clustering in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Unsupervised Learning: PCA and Clustering: frequently asked questions
What is the main purpose of PCA at CFA Level II?
PCA reduces a large set of correlated features into a smaller set of uncorrelated components that keep most of the variance. This lowers noise and the risk of overfitting. The trade-off is that components are hard to interpret.
What is the difference between k-means and hierarchical clustering?
K-means needs you to set the number of clusters k first and assigns points to the nearest centroid, repeating until stable. Hierarchical clustering builds a dendrogram of nested clusters and you choose the number of clusters by cutting it. K-means suits large datasets better.
What is the difference between agglomerative and divisive clustering?
Agglomerative clustering is bottom-up. Each observation starts alone and the closest clusters merge until one remains. Divisive clustering is top-down. All observations start together and the group is split repeatedly.
How do I decide how many principal components to keep?
Use the scree plot or the cumulative proportion of variance. Keep the fewest components that reach the target stated in the vignette, or stop where extra components add little explained variance.