Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

Three models are evaluated by 5-fold cross-validation. Mean squared errors on the five held-out folds are: Model A: 4, 6, 5, 7, 8; Model B: 3, 4, 5, 4, 4; Model C: 5, 5, 6, 5, 4. Model B's training MSE averages 0.5, Model A's averages 4.5 and Model C's averages 5.0. Which conclusion is correct?

Choose Model B. Its average cross-validation MSE is 4.0, versus 6.0 for A and 5.0 for C. Selection should be based on out-of-sample error, even though B's large gap from its 0.5 training MSE signals some overfitting.

  1. AChoose Model B because its training MSE is lowest
  2. BChoose Model C because its validation and training MSE are closest
  3. CChoose Model A because its training MSE is below its validation MSE
  4. DChoose Model B, which has the lowest average validation MSE of 4.0, although it shows the largest overfitting gapCorrect

Explanation

Average CV MSE: A = 30/5 = 6.0; B = 20/5 = 4.0; C = 25/5 = 5.0. Model selection should rely on out-of-sample (validation) error, so B is best at 4.0. Its gap to training MSE (4.0 − 0.5 = 3.5) shows overfitting, but this does not change that it generalizes best. Choosing by training MSE or smallest gap would be wrong.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions