FRM Part I · FRM Exam Part I · Machine Learning and Prediction
A risk analyst builds a model to predict loan defaults. The model achieves 99% accuracy on the training data but only 70% accuracy on a held-out test set. Which is the most likely explanation?
The model is most likely overfitting. It fits the training data almost perfectly, including noise, so it fails to generalize to unseen data. Underfitting would produce weak accuracy on both training and test sets, not a large gap between them.
- AThe model is underfitting because it is too simple
- BThe model is overfitting the training dataCorrect
- CThe test set has too low a variance to be informative
- DThe model has high bias and low variance
Explanation
A large gap between very high training performance and much weaker out-of-sample performance is the classic sign of overfitting: the model has learned noise specific to the training sample. Underfitting would give poor performance on both sets.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- Which statement correctly contrasts agglomerative hierarchical clustering with K-means?
- A data scientist has 5,000 observations and wants to select among several models and then report an unbiased estimate of final performance. …
- A risk manager clusters customers on two variables: annual transaction volume (in USD, ranging 0 to 2,000,000) and number of late payments (…
- A data scientist increases the complexity of a prediction model by adding many more predictors and higher-order terms. Holding the training …
- A risk analyst wants to group a bank's corporate borrowers into clusters based on financial ratios, with no predefined default labels availa…
- A PCA on six standardized predictors yields eigenvalues of 3.0, 1.5, 0.6, 0.45, 0.3 and 0.15. What is the smallest number of components need…