FRM Part I · FRM Exam Part I · Machine Learning and Prediction
In a random forest, each tree is trained on a bootstrap sample, leaving roughly one third of observations out of that tree's sample. How are these out-of-bag observations typically used?
Out-of-bag observations give an estimate of generalization error. Each observation is predicted using only the trees whose bootstrap samples excluded it, so the prediction is effectively out-of-sample. This provides a validation-style error estimate without needing a separate hold-out set.
- ATo estimate generalization error by predicting each observation using only trees that did not train on itCorrect
- BTo replace the test set for tuning the number of predictors only, never for error
- CTo add additional splits to each tree after growth
- DTo compute the bias of the model by removing them from all trees
Explanation
Each observation is out-of-bag for about a third of trees. Averaging predictions from those trees gives a validation-style prediction, and the resulting out-of-bag error estimates generalization error without a separate hold-out set.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- In a K-means run with K = 2, five one-dimensional observations are 1, 2, 4, 9 and 10. The initial centroids are 2 and 9. After assigning eac…
- A bank's credit-default classifier is tested on 200 loans. It flags 50 loans as defaults, of which 40 actually defaulted. In total 60 loans …
- A neural network with many parameters achieves very low error on the training set but substantially higher error on a validation set. Which …
- A model predicts whether a trade is fraudulent. Of 1,000 trades, 20 are actually fraudulent. The model flags 25 trades, of which 15 are actu…
- A bank combines 100 decision trees by majority vote to classify counterparties as default or non-default. Compared with a single deep tree, …
- Three models are evaluated by 5-fold cross-validation. Mean squared errors on the five held-out folds are: Model A: 4, 6, 5, 7, 8; Model B: …