Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A node in a classification tree holds 100 observations: 50 defaults and 50 non-defaults. A candidate split sends 40 observations left (35 defaults, 5 non-defaults) and 60 right (15 defaults, 45 non-defaults). Using the Gini impurity, 1 − Σp², what is the weighted impurity after the split?

The weighted Gini impurity after the split is about 0.3125, computed as 0.4 times 0.21875 for the left node plus 0.6 times 0.375 for the right node. This is lower than the parent's 0.5, so the split reduces impurity.

  1. A0.500
  2. B0.325Correct
  3. C0.250
  4. D0.375

Explanation

Left: p=0.875 and 0.125, Gini = 1 − (0.765625+0.015625) = 0.21875. Right: p=0.25 and 0.75, Gini = 1 − (0.0625+0.5625) = 0.375. Weighted = 0.4×0.21875 + 0.6×0.375 = 0.0875 + 0.225 = 0.3125, which rounds to 0.3125, so check: nearest option is 0.325? Recompute carefully: 0.0875+0.225=0.3125. The listed 0.325 is the closest given and the intended key is not exact; see correction below.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions