Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

An analyst splits 1,000 observations as follows: 600 for training, 200 for validation and 200 for testing. She trains five candidate models on the training set, selects the one with the lowest validation error, and then reports that model's error on the test set. What is the main purpose of the test set in this procedure?

The test set gives an unbiased estimate of how the chosen model will perform on new data. Because it was not used for fitting, tuning or model selection, it avoids the optimistic bias that affects the validation error.

  1. ATo tune the hyperparameters of the selected model
  2. BTo provide an unbiased estimate of out-of-sample performance of the chosen modelCorrect
  3. CTo increase the amount of data used for fitting the parameters
  4. DTo choose among the five candidate models

Explanation

Validation data were already used to choose the model, so its error is optimistically biased. The test set is untouched by training and selection, so it gives an unbiased estimate of performance on new data. Tuning and model selection belong to the validation stage.

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions