Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A data scientist splits data into training, validation and test sets to build a default-prediction model. She tries several regularization strengths and picks the one with the lowest error. Which procedure best avoids biased estimates of future performance?

Tune the regularization strength on the validation set and report performance on the untouched test set. Using the test set for tuning leaks information and gives optimistic error estimates, while tuning on training data favors overly complex models.

  1. AChoose the regularization strength on the test set, then report test error
  2. BChoose the regularization strength on the training set, then report validation error
  3. CChoose the regularization strength on the validation set, then report error on the untouched test setCorrect
  4. DChoose the regularization strength on the test set, then retrain on the validation set

Explanation

Hyperparameters should be tuned using validation data, and the test set must remain unused until the final evaluation. Tuning on the test set leaks information and makes its error optimistic. Tuning on training data favors overfit settings because training error falls with complexity.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions