Skip to content

IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning

A data scientist fits a very flexible model to a claims dataset. It achieves a very low error on the training data but a much higher error on unseen test data. Which statement best describes this outcome?

This is overfitting. The model is so flexible that it fits random noise in the training data, giving very low training error but poor performance on unseen data. Underfitting would instead produce high error on both training and test sets.

  1. AUnderfitting, because the model has too few parameters
  2. BOverfitting, because the model has learned noise specific to the training dataCorrect
  3. CData leakage, because the test set was used to select the parameters
  4. DBias, because the model ignores the relationship between the variables
  5. Regularisation, because the model penalises complexity

Explanation

A large gap between low training error and high test error is the classic sign of overfitting. The flexible model has fitted random noise in the training sample, so it does not generalise. Underfitting would show high error on both the training and test sets.

Did you get it right without looking?

One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.

More Elementary principles of machine learning questions