Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A risk manager clusters customers on two variables: annual transaction volume (in USD, ranging 0 to 2,000,000) and number of late payments (ranging 0 to 10), using K-means with Euclidean distance on the raw data. What is the most likely problem, and the appropriate remedy?

Transaction volume will dominate the Euclidean distance because its scale is vastly larger than the late-payment count, so clusters mostly reflect volume. The remedy is to standardize or rescale the variables before clustering so each has comparable influence on distances.

  1. ATransaction volume will dominate the distance calculation; standardize the variables before clusteringCorrect
  2. BLate payments will dominate the distance calculation; multiply them by 1,000 before clustering
  3. CK-means cannot be used with two variables; switch to a regression
  4. DThe centroids will be undefined; replace Euclidean distance with class labels

Explanation

Euclidean distance is sensitive to scale. A variable measured in millions swamps one measured on a 0-10 scale, so clusters reflect mainly transaction volume. Standardizing (e.g., z-scores) or rescaling gives each variable comparable influence. Inflating late payments would simply reverse the bias.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions