Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

A bank's data set includes a categorical feature, 'region', with four values: North, South, East and West. The analyst wants to use it in a regression-based model that includes an intercept. Which treatment is MOST appropriate?

Create three dummy variables and omit one region as the base category. With an intercept, including all four dummies makes them sum to the constant column, causing perfect multicollinearity. Integer coding would wrongly impose an order on nominal categories, and categorical data can be used once encoded.

  1. ACreate three binary dummy variables, omitting one category as the base, to avoid perfect multicollinearityCorrect
  2. BCreate four binary dummy variables and include all of them along with the intercept
  3. CEncode the regions as 1, 2, 3, 4 so the model treats them as an ordered numeric variable
  4. DDelete the feature because categorical variables cannot be used in regression

Explanation

With an intercept, including a dummy for every category produces perfect multicollinearity (the dummy variable trap), because the dummies sum to one. Using three dummies with one omitted base category avoids this. Integer coding imposes an artificial ordering on a nominal variable.

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions