FRM Part I · FRM Exam Part I · Machine Learning and Prediction
A risk manager clusters customers on two variables: annual transaction volume (in USD, ranging 0 to 2,000,000) and number of late payments (ranging 0 to 10), using K-means with Euclidean distance on the raw data. What is the most likely problem, and the appropriate remedy?
Transaction volume will dominate the Euclidean distance because its scale is vastly larger than the late-payment count, so clusters mostly reflect volume. The remedy is to standardize or rescale the variables before clustering so each has comparable influence on distances.
- ATransaction volume will dominate the distance calculation; standardize the variables before clusteringCorrect
- BLate payments will dominate the distance calculation; multiply them by 1,000 before clustering
- CK-means cannot be used with two variables; switch to a regression
- DThe centroids will be undefined; replace Euclidean distance with class labels
Explanation
Euclidean distance is sensitive to scale. A variable measured in millions swamps one measured on a 0-10 scale, so clusters reflect mainly transaction volume. Standardizing (e.g., z-scores) or rescaling gives each variable comparable influence. Inflating late payments would simply reverse the bias.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- Compared with a regression using all original predictors, a principal components regression (PCR) that uses the first few components has whi…
- A risk team has transaction data for 50,000 corporate clients with no labels. They run k-means to group clients with similar trading behavio…
- A risk analyst fits a linear model to predict loan losses using 60 correlated explanatory variables and only 120 observations. The analyst w…
- A single neuron receives inputs x1 = 2 and x2 = -1 with weights w1 = 0.5 and w2 = 1.5 and bias b = 0.5. It uses a ReLU activation, f(z) = ma…
- A node in a classification tree holds 100 observations: 50 defaults and 50 non-defaults. A candidate split sends 40 observations left (35 de…
- A risk analyst fits a single, very deep classification tree to predict loan default. It classifies the training data almost perfectly but pe…