FRM Part I · FRM Exam Part I · Machine-Learning Methods
A bank has a dataset of 1,000 customers with a categorical feature 'region' that has four unordered categories: North, South, East and West. The model is a logistic regression with an intercept. Which encoding is MOST appropriate?
Create three dummy variables and omit one category as the baseline. With an intercept in the model, using all four dummies causes perfect multicollinearity, while coding the regions 1 to 4 in one column wrongly imposes an order and equal spacing between unordered categories.
- ACreate three binary dummy variables, omitting one category as the baselineCorrect
- BEncode the regions as 1, 2, 3 and 4 in a single numeric column
- CCreate four binary dummy variables and keep the intercept
- DDrop the feature because categorical variables cannot be used in regression
Explanation
Unordered categories should be one-hot encoded. With an intercept, k categories need k-1 dummies to avoid perfect multicollinearity (the dummy variable trap). A single 1-4 column imposes a false ordering and spacing. Four dummies with an intercept are perfectly collinear.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- A risk analyst is building a model to predict loan defaults using features such as annual income (in thousands of dollars, ranging from 20 t…
- A data scientist splits data into training, validation and test sets to build a default-prediction model. Which use of the three sets is cor…
- A risk analyst fits a linear model to predict loan losses using 60 correlated predictors and finds that many coefficients are very large wit…
- A modeler has a strongly right-skewed feature, transaction size, with values ranging from 100 to 10,000,000 dollars. Extreme values are vali…
- A bank's data set includes a categorical feature, 'region', with four values: North, South, East and West. The analyst wants to use it in a …
- A risk team evaluates a fraud classifier where only 1% of transactions are fraudulent. A model that labels every transaction as non-fraud ac…