FRM Part I · FRM Exam Part I · Machine-Learning Methods
A risk analyst is building a model to predict loan defaults using features such as annual income (in thousands of dollars, ranging 20 to 500) and debt-to-income ratio (ranging 0 to 1). She plans to use a K-nearest neighbors algorithm. Which preprocessing step is MOST appropriate before fitting the model?
Standardize or rescale both features. K-nearest neighbors uses distance calculations, so a feature measured in large units like income would dominate a ratio between 0 and 1. Scaling gives each feature comparable influence, whereas distance-based methods are clearly not scale invariant.
- AStandardize or rescale both features so that they are on comparable scalesCorrect
- BLeave the features unscaled because distance-based methods are scale invariant
- CDrop the debt-to-income ratio because its range is small
- DApply scaling only to the target variable
Explanation
K-nearest neighbors relies on distances, so a feature with a large numeric range such as income would dominate one with a small range such as the ratio. Standardizing or rescaling puts the features on comparable scales. Distance-based methods are not scale invariant, so leaving them unscaled is wrong.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- A LASSO model minimizes the sum of squared residuals plus lambda times the sum of absolute coefficients. A model has coefficients 2.0, -1.5,…
- A ridge regression is fitted with a single standardized predictor and no intercept. The ordinary least squares slope is 0.80, the sum of squ…
- A risk team has a training sample of 1,000 observations for a credit scoring model. A feature 'missing_income' is present for 10% of records…
- For a given prediction model, the squared bias at a point is 0.04, the variance of the model's prediction is 0.09, and the irreducible error…
- A bank's credit model classifies 200 loan applicants as likely defaulters or non-defaulters. The confusion matrix shows 40 true positives (d…
- A hidden node in a neural network receives three inputs x1 = 2, x2 = -1 and x3 = 3 with weights 0.5, 2.0 and -0.4 respectively, and a bias o…