IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning
A model fitted to a training set gives a very low training error but a much higher error on a separate test set. Which is the most likely explanation?
The model is most likely overfitted. It has captured noise in the training data, so training error is low, but it generalises poorly to unseen data, giving a much higher test error. Underfitting would show high error on both sets.
- AThe model is overfitted to the training dataCorrect
- BThe model is underfitted because it is too simple
- CThe test set has been used to fit the parameters
- DThe training set is too large for the model
- The model has high bias and low variance
Explanation
A large gap between low training error and high test error is the classic sign of overfitting: the model has learned noise specific to the training sample and does not generalise. Underfitting would give high error on both sets.
Did you get it right without looking?
One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.
More Elementary principles of machine learning questions
- A general insurer wants to predict the claim amount (in rupees) for each motor policy from rating factors such as vehicle age, engine capaci…
- A lasso (L1-penalised) regression is fitted to predict lapse rates and the penalty parameter lambda is increased from a small value to a ver…
- A lapse-prediction model is fitted to 1,000 policies. Of the 100 policies that actually lapsed, the model flags 70 as lapses. Of the 900 pol…
- A data scientist at a Mumbai insurer standardises all predictors using the mean and standard deviation of the full dataset, then splits the …
- A regression model is evaluated on a test set of 4 observations with actual values 10, 12, 14, 20 and predicted values 11, 10, 14, 16. What …
- A model fitted to training data gives a very low training error but a much higher error on a separate validation set. What is the most likel…