Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A risk analyst fits a single, very deep classification tree to predict loan default. It classifies the training data almost perfectly but performs poorly on a hold-out sample. Which statement best describes the problem and the most appropriate remedy?

The deep tree has overfit the training data, giving high variance and poor out-of-sample results. Pruning the tree or limiting its depth reduces complexity and improves generalization. Adding more splits would make the overfitting worse, since the model already fits training noise almost perfectly.

  1. AThe tree has high bias; adding more splits would fix it
  2. BThe tree has overfit the training data; pruning or limiting depth would reduce varianceCorrect
  3. CThe tree suffers from multicollinearity; dropping correlated predictors is the standard cure
  4. DThe tree is underfit; the training sample should be reduced

Explanation

Near-perfect training accuracy with weak out-of-sample accuracy is the signature of overfitting, i.e., high variance. Pruning, limiting depth or requiring a minimum leaf size reduces complexity. Adding more splits (option A) would worsen the overfitting.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions