Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

A data scientist splits data into training, validation and test sets to build a default-prediction model. Which use of the three sets is correct?

Parameters are fitted on the training set, hyperparameters are tuned on the validation set, and the test set is used once at the end to estimate out-of-sample performance. Tuning on the test set leaks information and gives an overly optimistic performance estimate.

  1. AFit parameters on the training set, tune hyperparameters on the validation set, and assess final performance once on the test setCorrect
  2. BFit parameters on the test set, tune hyperparameters on the training set, and assess performance on the validation set
  3. CTune hyperparameters repeatedly on the test set so the final error estimate is as low as possible
  4. DFit parameters and tune hyperparameters on the training set, then use validation and test sets interchangeably

Explanation

The training set estimates model parameters, the validation set selects hyperparameters such as regularization strength, and the untouched test set gives an unbiased estimate of out-of-sample performance. Tuning on the test set leaks information and makes the performance estimate optimistic.

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions