Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A data scientist has 5,000 observations and wants to select among several models and then report an unbiased estimate of final performance. Which data-splitting practice is most appropriate?

Train on the training set, tune and select using the validation set, and evaluate the chosen model once on an untouched test set. Keeping the test set separate gives an unbiased out-of-sample estimate, whereas reusing selection data makes performance look optimistic.

  1. ASelect the model using the test set, then retrain and report performance on the training set
  2. BTrain on the training set, tune and select using the validation set, and evaluate the chosen model once on an untouched test setCorrect
  3. CUse the same validation set to both select the model and report its final performance
  4. DTrain and evaluate on the full sample to maximize the data used

Explanation

The validation set guides hyperparameter tuning and model choice. The test set must stay untouched until the end so it provides an unbiased estimate of out-of-sample performance. Using the same data for selection and reporting makes the estimate optimistic.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions