IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning
A pricing team fits a claim-severity model and finds that its error on the data used for fitting is very small, but its error on a separate hold-out set is much larger. Which is the most likely explanation?
The model is most likely overfitted. It has captured noise peculiar to the training data, so it fits that data very closely but generalises poorly to unseen data, producing a much larger error on the hold-out set.
- AThe model is overfitted to the training dataCorrect
- BThe model is underfitted and too simple
- CThe hold-out set has been used to train the model
- DThe model has too little flexibility to capture the signal
- The training data contains no random noise
Explanation
A large gap between low training error and high test error is the classic sign of overfitting: the model has learned noise specific to the training sample. Underfitting would show high error on both sets. Leakage of the hold-out set into training would make hold-out error low, not high.
Did you get it right without looking?
One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.
More Elementary principles of machine learning questions
- A general insurer wants to predict the claim amount (in rupees) for each motor policy from rating factors such as vehicle age, engine capaci…
- A binary classifier for insurance claim fraud is tested on 200 claims. It flags 40 as fraud, of which 30 are truly fraud. In total 50 of the…
- A lasso (L1-penalised) regression is fitted to predict lapse rates and the penalty parameter lambda is increased from a small value to a ver…
- A lapse-prediction model is fitted to 1,000 policies. Of the 100 policies that actually lapsed, the model flags 70 as lapses. Of the 900 pol…
- A data scientist at a Mumbai insurer standardises all predictors using the mean and standard deviation of the full dataset, then splits the …
- When building a predictive model for insurance data, which action best reduces the risk of data leakage?