IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning
A analyst applies K-means clustering to policyholder data containing annual premium (in rupees, values in the lakhs) and age (in years, values 20 to 70) without any rescaling. What is the most likely consequence?
Premium will dominate. Euclidean distance is not scale invariant, so a variable measured in lakhs of rupees swamps age measured in years, and the clusters will mostly reflect premium. Standardising the variables before clustering avoids this distortion.
- ADistances will be dominated by premium, so clusters will mainly reflect premium and largely ignore ageCorrect
- BDistances will be dominated by age, so premium will be ignored
- CThe algorithm will fail to converge for any data set
- DThe clusters will be unaffected, since Euclidean distance is scale invariant
- The number of clusters will automatically be reduced to one
Explanation
Euclidean distance sums squared differences, so the variable with the larger numerical spread contributes far more. Premium in lakhs dwarfs age differences of tens. Standardising the variables first removes this problem.
Did you get it right without looking?
One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.
More Elementary principles of machine learning questions
- In principal components analysis (PCA) of a standardised data set with 5 variables, the eigenvalues of the correlation matrix are 2.5, 1.3, …
- A pricing team fits a very flexible model to claim data and finds that its error on the training data is very small but its error on a separ…
- A data scientist at an Indian insurer tunes the regularisation parameter of a lasso regression. She tries many values and picks the one with…
- A binary classifier for insurance claim fraud is tested on 200 claims. It flags 40 as fraud, of which 30 are truly fraud. In total 50 of the…
- Which statement about the K-means algorithm is correct?
- An analyst uses 5-fold cross-validation on 1,000 policy records to choose between candidate models. Which description of the procedure is co…