Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

A risk analyst fits a single unpruned decision tree to predict loan default and finds that it classifies the training data almost perfectly but performs poorly on a hold-out sample. Which action is most consistent with addressing this problem?

Pruning the tree or setting a minimum number of observations per leaf is correct. The tree is overfitting, fitting noise in training data, and limiting complexity reduces variance and improves out-of-sample performance. Deeper trees or zero training error would worsen overfitting.

  1. AGrow the tree to greater depth so that each leaf contains one observation
  2. BPrune the tree or impose a minimum number of observations per leafCorrect
  3. CRemove the hold-out sample and retrain on all data
  4. DAdd more splits until training error is exactly zero

Explanation

Near-perfect training fit with poor hold-out performance signals overfitting. Pruning or requiring a minimum leaf size limits tree complexity and reduces variance. Growing deeper or driving training error to zero worsens overfitting, and discarding the hold-out set removes the means of detecting it.

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions