Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A risk analyst builds a model to predict loan defaults. The model achieves 99% accuracy on the training data but only 70% accuracy on a held-out test set. Which is the most likely explanation?

The model is most likely overfitting. It fits the training data almost perfectly, including noise, so it fails to generalize to unseen data. Underfitting would produce weak accuracy on both training and test sets, not a large gap between them.

  1. AThe model is underfitting because it is too simple
  2. BThe model is overfitting the training dataCorrect
  3. CThe test set has too low a variance to be informative
  4. DThe model has high bias and low variance

Explanation

A large gap between very high training performance and much weaker out-of-sample performance is the classic sign of overfitting: the model has learned noise specific to the training sample. Underfitting would give poor performance on both sets.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions