FRM Part I · FRM Exam Part I · Machine-Learning Methods
Which statement best distinguishes a random forest from simple bagging of decision trees?
A random forest randomly selects a subset of features at each split, in addition to bootstrapping observations. This decorrelates the trees, so averaging reduces variance more than in plain bagging. Sequential error correction is boosting, not a random forest.
- AA random forest uses only a single bootstrap sample for all trees
- BA random forest considers only a random subset of features at each split, which lowers correlation among treesCorrect
- CA random forest builds trees sequentially, each correcting the previous tree's errors
- DA random forest requires all features to be considered at every split
Explanation
Both methods train trees on bootstrap samples, but a random forest also restricts each split to a random subset of predictors. This decorrelates the trees and makes averaging more effective at reducing variance. Sequential error correction describes boosting.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- A bank's data set includes a categorical feature, 'region', with four values: North, South, East and West. The analyst wants to use it in a …
- A risk analyst has a dataset of 20,000 past retail loans, each labelled with whether the borrower defaulted within 12 months, and wants a mo…
- A node in a classification tree contains 40 loans, of which 20 defaulted. A candidate split sends 20 loans to the left child with 2 defaults…
- An analyst has 10,000 observations to forecast corporate bond downgrades. She first standardizes all features using the mean and standard de…
- In the logistic regression ln(p/(1-p)) = b0 + b1 x with b1 = 0.693, how does a one-unit increase in x affect the odds of the positive class,…
- A risk analyst fits a single unpruned decision tree to predict loan default and finds that it classifies the training data almost perfectly …