Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

A bank has a dataset of 1,000 customers with a categorical feature 'region' that has four unordered categories: North, South, East and West. The model is a logistic regression with an intercept. Which encoding is MOST appropriate?

Create three dummy variables and omit one category as the baseline. With an intercept in the model, using all four dummies causes perfect multicollinearity, while coding the regions 1 to 4 in one column wrongly imposes an order and equal spacing between unordered categories.

  1. ACreate three binary dummy variables, omitting one category as the baselineCorrect
  2. BEncode the regions as 1, 2, 3 and 4 in a single numeric column
  3. CCreate four binary dummy variables and keep the intercept
  4. DDrop the feature because categorical variables cannot be used in regression

Explanation

Unordered categories should be one-hot encoded. With an intercept, k categories need k-1 dummies to avoid perfect multicollinearity (the dummy variable trap). A single 1-4 column imposes a false ordering and spacing. Four dummies with an intercept are perfectly collinear.

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions