FRM Exam Part I · Machine Learning and Prediction
Decision Trees, Ensembles and Random Forests for FRM Part I
Updated 11 October 2026 · Fact-checked
A decision tree splits data into groups using yes/no rules on features, then predicts a class or an average value in each final group. Ensembles combine many trees. Bagging and random forests average independent trees to cut variance. Boosting builds trees in sequence to cut bias. Pruning controls tree size to limit overfitting.
Understand Decision Trees, Ensembles and Random Forests
A decision tree predicts an outcome by asking a series of simple questions. Each question splits the data on one feature, for example 'debt-to-income above 40%?'. You follow the answers down to a leaf, and the leaf gives the prediction. Trees are easy to explain, which matters in credit and risk work where decisions must be justified.
A classification tree predicts a category, such as default or no default. The leaf predicts the majority class, or the share of defaults in that leaf. A regression tree predicts a number, such as loss given default. The leaf predicts the average outcome of the training observations in it. At each split the algorithm picks the feature and cut-off that most reduce impurity (for classification, measured by Gini impurity or entropy) or squared error (for regression). It is a greedy method: it picks the best split now and does not look ahead.
A tree left to grow fully can fit the training data almost perfectly. It has low bias but high variance and overfits. Pruning fixes this. You can stop growth early (limit depth, require a minimum number of observations per leaf, require a minimum improvement per split). Or you can grow a large tree and then cut back branches that add little, using a penalty on the number of leaves and choosing the size by validation or cross-validation.
A single tree is unstable: small changes in the data can change the tree a lot. Ensemble methods combine many models. Bagging (bootstrap aggregating) draws many bootstrap samples, fits a tree to each, and averages the predictions (regression) or takes a majority vote (classification). This lowers variance. A random forest adds one more idea: at each split only a random subset of features is considered. That makes the trees less correlated, so averaging works better. Boosting fits trees one after another, each focusing on the errors of the earlier ones, with small steps controlled by a learning rate. It mainly reduces bias and can overfit if run too long.
Two related methods may appear alongside trees. K-nearest neighbours predicts from the k closest training points by vote or average, so it needs feature scaling and a sensible k. A support vector machine finds the boundary that separates classes with the widest margin. Know these at the level of the idea, not the algebra.
Key formulas to remember
- Gini impurity of a node
- G = 1 − Σ pᵢ²
- pᵢ is the share of class i in the node. For two classes, G = 2p(1 − p). Zero means a pure node.
- Entropy of a node
- H = − Σ pᵢ × ln(pᵢ)
- Another impurity measure. Zero for a pure node. Log base only scales the value.
- Weighted impurity after a split
- Weighted G = (n_left ÷ n) × G_left + (n_right ÷ n) × G_right
- Choose the split with the lowest weighted impurity, which is the largest impurity reduction.
- Regression tree split criterion
- Minimise SSE = Σ(y − ȳ_left)² + Σ(y − ȳ_right)²
- Leaf prediction is the mean of y in that leaf.
- Cost-complexity pruning
- Cost = Error(tree) + α × (number of leaves)
- Larger α gives a smaller tree. Choose α by validation or cross-validation.
- Bagging prediction
- Prediction = (1 ÷ B) × Σ f_b(x), b = 1 to B
- For classification use majority vote. Averaging reduces variance, not bias.
- Variance of an average of B equally correlated models
- Var = ρσ² + (1 − ρ)σ² ÷ B
- ρ is the correlation between models, σ² each model's variance. Lower ρ (random forest) helps more than just raising B.
How to solve Decision Trees, Ensembles and Random Forests questions
Use this order for any question on trees and ensembles.
- 1Identify the task: classification (category) or regression (number). This sets the leaf rule and the split criterion.
- 2If a split must be chosen, compute the impurity of the parent and the weighted impurity of each child. Pick the lowest weighted impurity.
- 3Write the leaf prediction: majority class or class proportion for classification, mean of y for regression.
- 4For size questions, link depth and number of leaves to the bias-variance tradeoff: bigger tree, lower bias, higher variance. Pruning moves the other way.
- 5For ensembles, name the main effect: bagging and random forests cut variance, boosting cuts bias.
- 6Check the mechanism: bootstrap samples, random feature subsets at each split, or sequential fitting on errors with a learning rate.
- 7Check the answer against the data hygiene issues: out-of-sample validation, scaling for k-NN, and interpretability trade-offs.
Quickest way: Keyword-to-concept shortcut
When to use it: Use for conceptual multiple-choice questions where you must tell methods apart quickly.
- Parallel trees on bootstrap samples, averaged: bagging. Add random feature subsets: random forest.
- Trees built in sequence on previous errors: boosting.
- Single tree that overfits and is unstable: high variance, so prune or ensemble.
- Need a split calculation: for two classes use G = 2p(1 − p), then weight by node size.
- Method needs distances and scaling: k-NN. Maximum margin boundary: SVM.
- Eliminate options that say ensembling removes bias in bagging or that deeper trees reduce variance.
Common mistakes in Decision Trees, Ensembles and Random Forests
Saying bagging reduces bias.
Averaging sounds like it improves everything.
Fix: Averaging models with similar bias keeps the bias and lowers variance. Boosting is the method aimed at bias.
Forgetting to weight child impurities by node size.
Students add or average the two Gini values directly.
Fix: Multiply each child's impurity by its share of observations before adding.
Thinking a random forest uses all features at every split.
It is confused with bagging.
Fix: A random forest considers only a random subset of features at each split, which decorrelates the trees.
Believing a fully grown tree is best because it fits the training data perfectly.
Training accuracy is mistaken for real performance.
Fix: Judge by validation or test error. A perfect training fit signals overfitting.
Treating a regression tree leaf as a majority vote.
The classification rule is applied to every tree.
Fix: Regression leaves predict the mean of the outcomes in the leaf.
Assuming more trees in a random forest causes overfitting in the same way as boosting.
Both are called ensembles.
Fix: Adding trees to a bagged forest does not by itself raise variance. Boosting for too many rounds can overfit.
Worked examples
Example 1
A node holds 100 loans: 40 defaulted and 60 did not. A split on 'income below ₹6,00,000' gives a left node of 50 loans (35 defaults, 15 non-defaults) and a right node of 50 loans (5 defaults, 45 non-defaults). Compute the Gini impurity reduction from the split.
Show the solution
- Parent Gini: p = 0.40, so G = 2 × 0.40 × 0.60 = 0.48.
- Left node: p = 35 ÷ 50 = 0.70, so G = 2 × 0.70 × 0.30 = 0.42.
- Right node: p = 5 ÷ 50 = 0.10, so G = 2 × 0.10 × 0.90 = 0.18.
- Weighted child Gini = 0.5 × 0.42 + 0.5 × 0.18 = 0.21 + 0.09 = 0.30.
- Reduction = 0.48 − 0.30 = 0.18.
Answer: The split reduces Gini impurity by 0.18 (from 0.48 to 0.30).
Example 2
B = 100 trees each have prediction variance σ² = 4 and pairwise correlation ρ = 0.25. What is the variance of the averaged prediction? What would it be if ρ were reduced to 0.10?
Show the solution
- Use Var = ρσ² + (1 − ρ)σ² ÷ B.
- With ρ = 0.25: 0.25 × 4 = 1.00. Second term: 0.75 × 4 ÷ 100 = 0.03. Total = 1.03.
- With ρ = 0.10: 0.10 × 4 = 0.40. Second term: 0.90 × 4 ÷ 100 = 0.036. Total = 0.436.
- Compare: lowering correlation cut variance from 1.03 to 0.436, far more than extra trees could, since the first term does not shrink as B grows.
Answer: Variance is 1.03 at ρ = 0.25 and 0.436 at ρ = 0.10. This is why random forests restrict features at each split.
Exam tips
- Expect conceptual contrasts: bagging versus boosting, random forest versus bagging, pruning versus growing. Learn the one-line difference for each.
- If a question gives class counts, compute Gini or weighted impurity by hand. Keep four decimals only if needed.
- Link every method to bias-variance: deep tree is high variance, bagging cuts variance, boosting cuts bias, pruning cuts variance.
- Remember interpretability: one tree is easy to explain, forests and boosted models are not. Questions on model governance often use this.
- For k-NN, link small k to high variance and large k to high bias, and remember to scale features.
Practice questions from Machine Learning and Prediction
- A node in a classification tree holds 100 observations: 50 defaults and 50 non-defaults. A candidate split sends 40 observations left (35 de…
- Which statement correctly distinguishes a random forest from plain bagging of decision trees?
- A neural network hidden node receives inputs x1 = 2 and x2 = -1. The weights are w1 = 0.5 and w2 = 1.5, and the bias is 0.25. The node uses …
- In k-fold cross-validation with k = 5 applied to a sample of 500 observations, how many observations are used to train the model in each ite…
- A risk analyst applies principal components analysis (PCA) to a set of 10 highly correlated yield-curve variables. Which statement best desc…
Decision Trees, Ensembles and Random Forests in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Decision Trees, Ensembles and Random Forests: frequently asked questions
What is the difference between bagging and boosting?
Bagging fits models independently on bootstrap samples and averages them, which mainly reduces variance. Boosting fits models one after another, each correcting earlier errors, which mainly reduces bias. Boosting can overfit if you run too many rounds.
How does a random forest work?
It grows many trees, each on a bootstrap sample of the data. At every split only a random subset of features is considered. Predictions are averaged for regression or voted for classification. The feature restriction makes trees less correlated.
How do you prune a decision tree?
You either stop growth early with limits on depth, leaf size or split gain, or grow a large tree and cut back branches using a penalty on the number of leaves. You choose the final size by validation or cross-validation error.
Do I need to know k-nearest neighbours and support vector machines for FRM Part I?
Know the core idea of each and their strengths and weaknesses. k-NN predicts from the closest points and needs scaling. An SVM finds the widest-margin boundary. Check the current GARP Study Guide for the exact scope this year.