Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

An analyst compares two default models using 5-fold cross-validation on 1,000 observations. Model A has training error of 2% and mean validation error of 15%. Model B has training error of 9% and mean validation error of 10%. Which conclusion is best supported?

Model A is overfit and Model B should generalize better. Model A's validation error of 15% far exceeds its 2% training error, while Model B's gap is small and its validation error is lower. Model choice should rely on out-of-sample error, not training error.

  1. AModel A is overfit and Model B is likely to generalize betterCorrect
  2. BModel B is overfit because its training error is higher
  3. CModel A should be preferred because its training error is lower
  4. DBoth models generalize equally because errors are positive

Explanation

Model A has a large gap (13 points) between training and validation error, indicating overfitting. Model B's gap is only 1 point and its validation error is lower (10% vs 15%), so it should generalize better. Selection should be based on out-of-sample (validation) error, not training error.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions