Skip to content

FRM Exam Part I · Machine-Learning Methods

Decision Trees, Ensembles and K-Nearest Neighbors for FRM

Updated 11 October 2026 · Fact-checked

A decision tree splits data into groups using yes/no rules on features. Ensembles combine many models: bagging and random forests average trees built in parallel to cut variance, while boosting builds trees one after another to cut bias. K-nearest neighbors classifies a point by the majority vote of its k closest observations.

Understand Decision Trees, Ensembles and K-Nearest Neighbors

A decision tree works like a flowchart. At each node it asks a question about one feature, such as "Is debt-to-equity above 2?" The answer sends the observation down a branch. The final nodes are called leaves. A classification tree predicts a category (default or no default). A regression tree predicts a number, usually the average of the training values in that leaf.

The tree picks each split to make the resulting groups as pure as possible. For classification, purity is measured with Gini impurity or entropy. For regression, the split that most reduces the sum of squared errors is chosen. Trees are easy to read and handle nonlinear patterns. But a deep tree fits noise. It has low bias and high variance, so it overfits. You control this by limiting depth, requiring a minimum leaf size, or pruning.

A single tree is unstable, so we combine many. Bagging (bootstrap aggregating) draws many bootstrap samples, fits a tree on each, and averages the predictions (regression) or takes a majority vote (classification). This lowers variance. A random forest adds one more step: at each split, only a random subset of features is considered. That makes the trees less correlated, so averaging works better. Boosting is different. Trees are built in sequence, and each new tree focuses on the errors of the earlier ones. Boosting mainly reduces bias, but it can overfit if you run too many rounds.

A support vector machine (SVM) finds the boundary (hyperplane) that separates two classes with the widest possible margin. The observations closest to the boundary are the support vectors; only they define it. If classes overlap, a soft margin allows some misclassification, controlled by a penalty parameter. Kernels let an SVM draw nonlinear boundaries.

K-nearest neighbors (KNN) has no training step in the usual sense. To classify a new point, it finds the k closest training points by distance (usually Euclidean) and takes a majority vote. For regression it averages their values. Small k gives a flexible, high-variance fit. Large k gives a smoother, high-bias fit. Because it uses distance, you must scale features first.

Key formulas to remember

Gini impurity
G = 1 − Σ pᵢ²
pᵢ is the share of class i in the node. For two classes, the maximum is 0.5 and a pure node gives 0.
Entropy
H = − Σ pᵢ × ln(pᵢ)
Zero for a pure node. Higher means more mixed. Log base (2 or e) only rescales it.
Regression tree leaf prediction
ŷ = average of y in the leaf
Splits are chosen to minimise the sum of squared errors across the two child nodes.
Bagging prediction
ŷ = (1 ÷ B) × Σ ŷᵦ
B bootstrap trees. For classification use a majority vote.
Euclidean distance
d = √[Σ (xᵢ − zᵢ)²]
Used by KNN. Scale features first or large-valued ones dominate.
Weighted impurity after a split
G_split = (n_L ÷ n) × G_L + (n_R ÷ n) × G_R
Pick the split with the lowest value, which is the largest impurity reduction.

How to solve Decision Trees, Ensembles and K-Nearest Neighbors questions

Most questions ask you to identify the right method, explain a trade-off, or compute a simple impurity or distance. Use this order.

  1. 1Identify the task: classification (category) or regression (number).
  2. 2Identify the method named or described: single tree, bagging, random forest, boosting, SVM or KNN.
  3. 3Link the method to its bias-variance effect: deep tree = high variance; bagging and forests cut variance; boosting cuts bias; large k raises bias.
  4. 4If a calculation is needed, write the formula first (Gini, entropy, distance, or average).
  5. 5Plug in the numbers carefully, using class proportions rather than counts for impurity.
  6. 6For KNN, compute all distances, rank them, take the k smallest, then vote or average.
  7. 7Check the answer against the description: lower impurity means a better split; a pure node has impurity 0.

Quickest way: Match keyword to method

When to use it: Use for conceptual multiple-choice questions where wording points to one method.

  1. "Parallel, bootstrap samples, reduces variance" means bagging.
  2. "Random subset of features at each split" means random forest.
  3. "Sequential, focuses on errors" means boosting.
  4. "Maximum margin, support vectors" means SVM.
  5. "Majority vote of closest points, scale features" means KNN.
  6. "Interpretable, prone to overfit, prune" means single tree.
  7. Eliminate options that swap bias and variance.

Common mistakes in Decision Trees, Ensembles and K-Nearest Neighbors

  • Saying bagging and boosting both mainly reduce variance in the same way.

    Both are called ensembles, so they get lumped together.

    Fix: Bagging trains independently in parallel and lowers variance. Boosting trains sequentially on errors and mainly lowers bias.

  • Thinking a random forest uses all features at every split.

    Confusing it with plain bagging.

    Fix: A forest considers only a random subset at each split, which decorrelates the trees.

  • Not scaling features before KNN.

    Students forget KNN depends on distance.

    Fix: Standardise features so each contributes comparably to distance.

  • Choosing a small k to reduce overfitting.

    Small numbers feel simpler.

    Fix: Small k is more flexible and overfits. Larger k smooths the fit and raises bias.

  • Computing Gini with counts instead of proportions.

    Rushing through the formula.

    Fix: Divide each class count by the node total first, then square.

  • Believing all training points define an SVM boundary.

    Mixing SVM with regression.

    Fix: Only the support vectors, the points on or inside the margin, determine the boundary.

Worked examples

Example 1

A node has 40 loans: 30 performing and 10 defaulted. A split sends 20 loans left (18 performing, 2 defaulted) and 20 loans right (12 performing, 8 defaulted). Compute the weighted Gini impurity after the split and the reduction from the parent.

Show the solution
  1. Parent Gini = 1 − (30/40)² − (10/40)² = 1 − 0.5625 − 0.0625 = 0.3750.
  2. Left Gini = 1 − (18/20)² − (2/20)² = 1 − 0.81 − 0.01 = 0.18.
  3. Right Gini = 1 − (12/20)² − (8/20)² = 1 − 0.36 − 0.16 = 0.48.
  4. Weighted = (20/40) × 0.18 + (20/40) × 0.48 = 0.09 + 0.24 = 0.33.
  5. Reduction = 0.375 − 0.33 = 0.045.

Answer: Weighted Gini after the split is 0.33, a reduction of 0.045 from the parent's 0.375.

Example 2

A KNN classifier uses k = 3 with two scaled features. A new borrower is at (2, 3). Training points: A (2, 5) default; B (3, 3) no default; C (5, 7) default; D (1, 2) no default. Classify the borrower.

Show the solution
  1. Distance to A = √[(2−2)² + (3−5)²] = √4 = 2.000.
  2. Distance to B = √[(2−3)² + (3−3)²] = √1 = 1.000.
  3. Distance to C = √[(2−5)² + (3−7)²] = √(9+16) = 5.000.
  4. Distance to D = √[(2−1)² + (3−2)²] = √2 ≈ 1.414.
  5. Ranked: B (1.000), D (1.414), A (2.000), C (5.000). The three nearest are B, D and A.
  6. Votes: B no default, D no default, A default. Majority is no default (2 to 1).

Answer: The borrower is classified as no default.

Exam tips

  • Expect conceptual questions on bias versus variance for each ensemble method; learn one line per method.
  • Calculations are usually small: one Gini or entropy value, or a few distances. Show the formula and keep proportions exact.
  • Watch for options that swap bagging and boosting, or add "all features" to a random forest.
  • Remember trees are interpretable but unstable, while ensembles gain accuracy and lose interpretability.
  • For KNN questions, check whether the features need scaling and how k changes flexibility.

Practice questions from Machine-Learning Methods

Decision Trees, Ensembles and K-Nearest Neighbors in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Decision Trees, Ensembles and K-Nearest Neighbors: frequently asked questions

What is the difference between bagging and boosting?

Bagging fits many models independently on bootstrap samples and averages them, which mainly lowers variance. Boosting fits models one after another, each correcting the previous errors, which mainly lowers bias. Boosting can overfit if run too long.

How does a random forest differ from bagging?

Both average trees fit on bootstrap samples. A random forest also restricts each split to a random subset of features. This makes the trees less correlated, so averaging reduces variance more.

How does k-nearest neighbors work?

It measures the distance from a new point to every training point and selects the k closest. For classification it takes a majority vote; for regression it averages their values. Features should be scaled first.

What do I need to know about support vector machines for FRM?

Know that an SVM finds the hyperplane with the maximum margin between classes. Only the support vectors define it. A soft margin tolerates some errors, and kernels allow nonlinear boundaries.