Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

A risk analyst is building a model to predict loan defaults using features such as annual income (in thousands of dollars, ranging 20 to 500) and debt-to-income ratio (ranging 0 to 1). She plans to use a K-nearest neighbors algorithm. Which preprocessing step is MOST appropriate before fitting the model?

Standardize or rescale both features. K-nearest neighbors uses distance calculations, so a feature measured in large units like income would dominate a ratio between 0 and 1. Scaling gives each feature comparable influence, whereas distance-based methods are clearly not scale invariant.

  1. AStandardize or rescale both features so that they are on comparable scalesCorrect
  2. BLeave the features unscaled because distance-based methods are scale invariant
  3. CDrop the debt-to-income ratio because its range is small
  4. DApply scaling only to the target variable

Explanation

K-nearest neighbors relies on distances, so a feature with a large numeric range such as income would dominate one with a small range such as the ratio. Standardizing or rescaling puts the features on comparable scales. Distance-based methods are not scale invariant, so leaving them unscaled is wrong.

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions