IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning
A data scientist fits a very flexible model to a claims dataset. It achieves a very low error on the training data but a much higher error on unseen test data. Which statement best describes this outcome?
This is overfitting. The model is so flexible that it fits random noise in the training data, giving very low training error but poor performance on unseen data. Underfitting would instead produce high error on both training and test sets.
- AUnderfitting, because the model has too few parameters
- BOverfitting, because the model has learned noise specific to the training dataCorrect
- CData leakage, because the test set was used to select the parameters
- DBias, because the model ignores the relationship between the variables
- Regularisation, because the model penalises complexity
Explanation
A large gap between low training error and high test error is the classic sign of overfitting. The flexible model has fitted random noise in the training sample, so it does not generalise. Underfitting would show high error on both the training and test sets.
Did you get it right without looking?
One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.
More Elementary principles of machine learning questions
- A binary classifier for predicting policy lapse is tested on 200 policies. It gives 40 true positives, 20 false positives, 10 false negative…
- In principal components analysis (PCA) of a standardised data set with 5 variables, the eigenvalues of the correlation matrix are 2.5, 1.3, …
- A claims-classification model is trained with early stopping. Training error keeps decreasing each iteration, while validation error decreas…
- A health insurer in Pune has records of 50,000 past policyholders, each with age, BMI, smoking status and a recorded label showing whether a…
- A pricing team fits a very flexible model to claim data and finds that its error on the training data is very small but its error on a separ…
- A data scientist at an Indian insurer tunes the regularisation parameter of a lasso regression. She tries many values and picks the one with…