FRM Part I · FRM Exam Part I · Machine-Learning Methods
A risk analyst fits a single classification tree to predict loan default and lets it grow until every training observation sits in a pure leaf. The tree has near-zero training error but performs poorly on a hold-out sample. Which action is the most appropriate way to address this problem?
Prune the tree or limit its depth or minimum leaf size, selecting the setting using validation performance. A fully grown tree overfits the training data, giving low training error but high variance out of sample. Regularizing the tree reduces complexity and improves generalization.
- APrune the tree or impose a minimum leaf size or maximum depth, choosing the setting by validation performanceCorrect
- BAdd more splits so the tree captures remaining training noise
- CRemove the hold-out sample and rely on training error to select the tree
- DStandardize the features, since scaling controls tree depth
Explanation
A fully grown tree overfits: low bias but high variance. Pruning or limiting depth or leaf size reduces complexity, and the tuning is chosen using validation data. More splits worsen overfitting, training error is a biased guide, and trees are insensitive to feature scaling.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- A risk analyst is building a model to predict loan defaults using features such as annual income (in thousands of dollars, ranging from 20 t…
- A bagged ensemble averages B identical-variance tree predictions. Each tree has prediction variance 0.40, and the pairwise correlation betwe…
- A data scientist splits data into training, validation and test sets to build a default-prediction model. Which use of the three sets is cor…
- A risk analyst has a dataset of 20,000 past retail loans, each labelled with whether the borrower defaulted within 12 months, and wants a mo…
- A risk analyst fits a linear model to predict loan losses using 60 correlated predictors and finds that many coefficients are very large wit…
- A modeler has a strongly right-skewed feature, transaction size, with values ranging from 100 to 10,000,000 dollars. Extreme values are vali…