Skip to content

IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning

A analyst applies K-means clustering to policyholder data containing annual premium (in rupees, values in the lakhs) and age (in years, values 20 to 70) without any rescaling. What is the most likely consequence?

Premium will dominate. Euclidean distance is not scale invariant, so a variable measured in lakhs of rupees swamps age measured in years, and the clusters will mostly reflect premium. Standardising the variables before clustering avoids this distortion.

  1. ADistances will be dominated by premium, so clusters will mainly reflect premium and largely ignore ageCorrect
  2. BDistances will be dominated by age, so premium will be ignored
  3. CThe algorithm will fail to converge for any data set
  4. DThe clusters will be unaffected, since Euclidean distance is scale invariant
  5. The number of clusters will automatically be reduced to one

Explanation

Euclidean distance sums squared differences, so the variable with the larger numerical spread contributes far more. Premium in lakhs dwarfs age differences of tens. Standardising the variables first removes this problem.

Did you get it right without looking?

One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.

More Elementary principles of machine learning questions