Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A risk analyst fits a single classification tree to predict loan default and lets it grow until every terminal node contains only one class. Training accuracy is 100%, but out-of-sample accuracy is much lower. Which statement best describes the problem and a standard remedy?

The tree is overfitted with high variance, since it memorizes training data and generalizes poorly. Pruning or limiting its depth reduces complexity and is the standard remedy. Standardizing features does not help because tree splits do not depend on feature scale.

  1. AThe tree is underfitted with high bias; adding more features to the root node is the remedy
  2. BThe tree is overfitted with high variance; pruning or limiting depth is a standard remedyCorrect
  3. CThe tree suffers from multicollinearity; dropping correlated predictors is the remedy
  4. DThe tree is biased because of unscaled inputs; standardizing the features is the remedy

Explanation

A fully grown tree memorizes the training data, giving perfect in-sample fit but poor generalization, which is overfitting (high variance). Pruning, limiting depth, or requiring a minimum node size reduces complexity. Scaling is irrelevant because tree splits are invariant to monotonic feature transformations.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions