FRM Part I · FRM Exam Part I · Machine-Learning Methods
A model predicting loan losses achieves a mean squared error of 0.5 on the training sample but 4.0 on a held-out validation sample. Which interpretation and remedy is most appropriate?
The model is probably overfitted, since training error is very low but validation error is eight times larger. The sensible response is to reduce complexity or add regularization. Underfitting would show high error in both samples, so making the model more flexible would be the wrong remedy.
- AThe model is likely overfitted; apply regularization or reduce complexityCorrect
- BThe model is underfitted; add more flexible features and remove regularization
- CThe model has low variance; increase the number of parameters
- DThe validation sample is biased by definition; rely on the training error
Explanation
A large gap between low training error and much higher validation error signals overfitting: the model fits noise in the training data and generalizes poorly. Remedies include regularization, simpler models, or more data. Underfitting would show high error on both samples, so adding flexibility would worsen the problem.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- A risk team evaluates a fraud classifier where only 1% of transactions are fraudulent. A model that labels every transaction as non-fraud ac…
- An analyst has a feature with values 2, 4, 6, 8 and 20 in a training sample. The analyst applies min-max scaling to the range [0, 1] using t…
- A data scientist estimates a credit-scoring model with 200 candidate predictors, believing only about 15 truly matter, and wants the fitted …
- A random forest differs from simply bagging many fully grown decision trees because, at each split, a random forest:
- A bank builds a credit-scoring model. Preprocessing steps include standardizing features using the mean and standard deviation of the full d…
- In a KNN classifier, an analyst moves from K = 1 to K = 25 on a noisy credit dataset. Which is the expected effect?