Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

In a random forest, each tree is trained on a bootstrap sample, leaving roughly one third of observations out of that tree's sample. How are these out-of-bag observations typically used?

Out-of-bag observations give an estimate of generalization error. Each observation is predicted using only the trees whose bootstrap samples excluded it, so the prediction is effectively out-of-sample. This provides a validation-style error estimate without needing a separate hold-out set.

  1. ATo estimate generalization error by predicting each observation using only trees that did not train on itCorrect
  2. BTo replace the test set for tuning the number of predictors only, never for error
  3. CTo add additional splits to each tree after growth
  4. DTo compute the bias of the model by removing them from all trees

Explanation

Each observation is out-of-bag for about a third of trees. Averaging predictions from those trees gives a validation-style prediction, and the resulting out-of-bag error estimates generalization error without a separate hold-out set.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions