FRM Part I · FRM Exam Part I · Machine Learning and Prediction
A quantitative analyst fits a very flexible model to 500 observations. It achieves a training mean squared error of 0.2 but a test mean squared error of 1.9 on held-out data. A simpler model has training MSE of 0.9 and test MSE of 1.0. Which conclusion is best supported?
The flexible model is overfitting: its training error is very low but its test error is much higher. Since prediction quality is judged out of sample, the simpler model with test MSE of 1.0 beats 1.9 and should be preferred.
- AThe flexible model is underfitting and has high bias
- BThe flexible model is overfitting and the simpler model should be preferred for predictionCorrect
- CBoth models are equally good because training errors are low
- DThe simpler model is overfitting because its training error is higher
Explanation
A large gap between low training error and much higher test error signals overfitting (high variance). Out-of-sample performance is the relevant criterion, and the simpler model's test MSE of 1.0 is lower than 1.9. Higher training error alone does not imply overfitting.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- In k-fold cross-validation with k = 5 applied to a sample of 500 observations, how many observations are used to train the model in each ite…
- A risk analyst applies principal components analysis (PCA) to a set of 10 highly correlated yield-curve variables. Which statement best desc…
- A risk analyst fits a single, very deep classification tree to predict loan default. It classifies the training data almost perfectly but pe…
- A risk analyst builds a feed-forward neural network to predict loan default. The network has an input layer, two hidden layers and an output…
- An analyst compares two default models using 5-fold cross-validation on 1,000 observations. Model A has training error of 2% and mean valida…
- A model predicts loan default using 40 predictors and achieves near-zero error on the training sample but much higher error on a separate va…