Skip to content

IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning

A pricing team fits a claim-severity model and finds that its error on the data used for fitting is very small, but its error on a separate hold-out set is much larger. Which is the most likely explanation?

The model is most likely overfitted. It has captured noise peculiar to the training data, so it fits that data very closely but generalises poorly to unseen data, producing a much larger error on the hold-out set.

  1. AThe model is overfitted to the training dataCorrect
  2. BThe model is underfitted and too simple
  3. CThe hold-out set has been used to train the model
  4. DThe model has too little flexibility to capture the signal
  5. The training data contains no random noise

Explanation

A large gap between low training error and high test error is the classic sign of overfitting: the model has learned noise specific to the training sample. Underfitting would show high error on both sets. Leakage of the hold-out set into training would make hold-out error low, not high.

Did you get it right without looking?

One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.

More Elementary principles of machine learning questions