CFA Level II Exam · Machine Learning
Supervised Learning: CART and Ensemble Learning for CFA Level 2
Updated 7 October 2026 · Fact-checked
CART is a supervised method that splits data into branches using feature cutoffs to predict a category (classification) or a number (regression). Deep trees overfit, so you prune them. Ensemble learning combines many models: bagging and random forests average trees built on resampled data, while boosting builds trees in sequence to fix earlier errors.
Understand Supervised Learning: CART and Ensemble Learning
A classification and regression tree (CART) is a supervised learning model. It asks a series of yes/no questions about features, such as "Is debt-to-equity above 1.5?" Each question is a split at a decision node. The final boxes are terminal nodes (leaves) and give the prediction.
If the target is a category (default or no default), the tree is a classification tree. The prediction at a leaf is the most common class there. If the target is continuous (next-year return), it is a regression tree. The prediction is the average target value at the leaf. At each node the algorithm picks the feature and cutoff that make the resulting groups as pure as possible, meaning the lowest error or impurity (for example, the lowest Gini or entropy for classification).
Trees are easy to read, need little data preparation and capture non-linear relationships and interactions. Their weakness is overfitting. If you let a tree grow until every leaf is pure, it memorizes noise in the training data and does poorly on new data. To control this, you set stopping rules (maximum depth, minimum observations per leaf, minimum gain per split) or you prune: grow the tree, then remove sections that add little predictive power on validation data. Pruning gives a simpler tree that generalizes better.
A single tree is also unstable. Small changes in data can change the structure a lot. Ensemble learning fixes this by combining several models, called weak learners or base models, into one prediction. The idea is that many different models make errors that partly cancel out. Ensembles can be built from different algorithms, or from the same algorithm trained on different data.
Bagging (bootstrap aggregating) draws many random samples from the training set with replacement, trains one model on each, then combines them. Classification uses majority vote and regression uses the average. Bagging mainly reduces variance (overfitting). A random forest is bagged trees with an extra step: at each split only a random subset of features is considered. This makes the trees less correlated, so averaging works better. Boosting trains models one after another. Each new model focuses on the observations the earlier ones got wrong, either by giving them more weight (AdaBoost) or by fitting the remaining errors (gradient boosting). Boosting reduces bias but can overfit if run too long. A forest loses the clear readability of one tree.
Key formulas to remember
- Classification tree leaf prediction
- Predicted class = most common class among training observations in the leaf
- Used for categorical targets such as default or no default.
- Regression tree leaf prediction
- Predicted value = average of the target values in the leaf
- Used for continuous targets such as returns.
- Bagging combination rule
- Classification: majority vote of the models. Regression: average of the model predictions
- Each model is trained on a bootstrap sample drawn with replacement.
- Random forest feature rule
- At each split, consider only a random subset of the features
- Lowers correlation between trees. The usual aim is lower variance.
- Boosting idea
- Models are trained in sequence; each focuses on prior errors
- Aimed at reducing bias. AdaBoost reweights misclassified observations; gradient boosting fits the residual errors.
- Overfitting check
- Training error low and validation error much higher means overfitting
- Pruning or stopping rules reduce tree complexity.
How to solve Supervised Learning: CART and Ensemble Learning questions
Most item-set questions on this topic ask you to pick the right model, read a tree, or choose a fix for overfitting. Use the same sequence each time.
- 1Identify the target in the vignette. A category means classification tree; a number means regression tree.
- 2If a tree is shown, trace the observation from the root through each decision node, applying the cutoff at each one, until you reach a leaf.
- 3Read the leaf prediction: majority class for classification, average value for regression.
- 4If the vignette shows high training accuracy but poor test accuracy, call it overfitting. Look for the fix: pruning, a maximum depth, a minimum leaf size, or an ensemble.
- 5If the question is about combining models, decide the type. Parallel models on bootstrap samples means bagging; with random feature subsets means random forest; sequential models focusing on errors means boosting.
- 6Match the goal. Reducing variance points to bagging or random forest. Reducing bias by learning from errors points to boosting.
- 7Check the trade-off the question asks about, such as interpretability. A single tree is easy to explain; an ensemble is not.
Quickest way: Keyword matching for model type
When to use it: Use when the question describes a method in words and asks you to name it or state its effect.
- Spot "bootstrap", "with replacement" or "majority vote": it is bagging.
- Spot "random subset of features at each split": it is a random forest.
- Spot "sequential", "misclassified observations get more weight" or "fits residuals": it is boosting.
- Spot "remove branches" or "reduce depth": it is pruning, which fights overfitting.
- Eliminate any option that claims one tree is stable or that a fully grown tree generalizes best.
Common mistakes in Supervised Learning: CART and Ensemble Learning
Saying a fully grown tree is the best model because it has zero training error.
Students link low error with a good model.
Fix: Training error is not the test. A tree grown until every leaf is pure usually overfits. Judge by performance on validation or test data.
Mixing up bagging and boosting.
Both combine many models, so the descriptions blur.
Fix: Bagging trains models independently on bootstrap samples and averages them. Boosting trains models in sequence, each correcting earlier errors.
Thinking a random forest uses every feature at every split.
Students assume it is just bagged trees.
Fix: A random forest considers only a random subset of features at each split. This is what makes the trees less correlated.
Using the average of a leaf for a classification tree, or the majority class for a regression tree.
The two tree types look the same on a diagram.
Fix: Classification leaf gives the majority class. Regression leaf gives the mean target value.
Claiming ensembles are easier to interpret than a single tree.
Students assume more models mean more insight.
Fix: A single tree is readable as a set of rules. A forest or boosted ensemble usually gains accuracy but loses clear interpretability.
Believing boosting can never overfit.
Students remember that boosting learns from errors.
Fix: Boosting can overfit if too many rounds are used, including fitting noise. Tune the number of rounds on validation data.
Worked examples
Example 1
A credit analyst builds a classification tree to predict loan default (1) or no default (0). The root node splits on debt-to-income (DTI): DTI ≤ 40% goes left, DTI > 40% goes right. The left node splits on years of employment: ≤ 2 years leads to a leaf with 30 defaults and 70 non-defaults; > 2 years leads to a leaf with 5 defaults and 95 non-defaults. The right leaf has 60 defaults and 40 non-defaults. Predicting the majority class in each leaf gives 75% accuracy on the 300 training loans, and the tree scores 74% on test data. The analyst then grows a much deeper tree with purer leaves. It scores 97% on training data but 62% on test data. Q1: What does a borrower with DTI 35% and 1 year of employment get predicted by the first tree? Q2: What is the likely problem with the deeper tree? Q3: Which action best addresses it?
Show the solution
- Q1: DTI 35% is ≤ 40%, so go left. Employment 1 year is ≤ 2, so reach the leaf with 30 defaults and 70 non-defaults.
- The majority class in that leaf is no default (70 of 100), so the prediction is no default.
- Q2: The deeper tree has 97% training accuracy but only 62% test accuracy. Such a large gap shows it has memorized noise in the training data, which is overfitting. (The first tree's training accuracy of 75% and test accuracy of 74% are almost the same, so it does not show this gap.)
- Q3: Prune the tree or impose stopping rules such as a maximum depth or a minimum number of observations per leaf. Using an ensemble such as a random forest would also help.
Answer: Q1: No default (the leaf is 70% non-default). Q2: Overfitting in the deeper tree (97% training versus 62% test). Q3: Prune or restrict the tree (or use an ensemble) to improve out-of-sample performance.
Example 2
An analyst forecasts quarterly stock returns. Model A averages 200 trees, each trained on a bootstrap sample, with each split limited to a random subset of the features. Model B trains 200 shallow trees one after another, giving more weight to observations that earlier trees predicted poorly. Q1: Name each model. Q2: Which one mainly aims to reduce variance? Q3: Given the target is a return, how does Model A form its final prediction?
Show the solution
- Q1: Model A uses bootstrap samples and random feature subsets at each split, so it is a random forest. Model B builds trees in sequence and reweights errors, so it is boosting.
- Q2: Averaging many de-correlated trees mainly reduces variance, so Model A.
- Q3: Return is continuous, so this is regression. Model A averages the predictions of the individual trees.
Answer: Q1: A is a random forest; B is boosting. Q2: Model A. Q3: It takes the average of the trees' predicted returns.
Exam tips
- Read the target first. Category or number decides the leaf rule and the tree type.
- Link each ensemble with its keyword: bootstrap means bagging, random feature subsets means random forest, sequential error correction means boosting.
- When a vignette gives training and test results, compare them before choosing an answer. A big gap means overfitting.
- Expect conceptual contrasts such as interpretability versus accuracy. A single tree wins on interpretability; ensembles usually win on predictive stability.
- There is no penalty for wrong answers, so answer every question even if you must guess between two options.
Supervised Learning: CART and Ensemble Learning in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Supervised Learning: CART and Ensemble Learning: frequently asked questions
What is the difference between bagging and boosting in CFA Level II?
Bagging trains models independently on bootstrap samples and combines them by vote or average, mainly to cut variance. Boosting trains models one after another, each focusing on the errors of earlier ones, mainly to cut bias.
How does a random forest differ from bagging?
A random forest is bagging applied to trees with one extra rule: each split considers only a random subset of the features. This makes the trees less alike, so averaging them reduces variance more.
How do you prevent overfitting in a decision tree?
Limit growth with stopping rules such as a maximum depth or a minimum number of observations per leaf. You can also prune a grown tree by removing branches that add little value on validation data. Using an ensemble also helps.
Why are trees called CART?
CART stands for classification and regression trees. The same splitting idea predicts a category when the target is categorical and a numeric value when the target is continuous.