Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A dataset of 1,000 observations is split into training, validation, and test sets in a 60/20/20 ratio. A analyst tunes a hyperparameter by trying 5 values, picking the one with the lowest error on one of the sets, then reports the final error on another. Which assignment is correct, and how many observations are in the set used for the final unbiased error estimate?

Tune hyperparameters on the validation set and report final performance on the held-out test set, which contains 200 observations (20% of 1,000). Using the test set for tuning or reporting training error would give a biased, optimistic estimate.

  1. ATune on test set; report on validation set; 200 observations
  2. BTune on training set; report on validation set; 600 observations
  3. CTune on validation set; report on test set; 200 observationsCorrect
  4. DTune on validation set; report on training set; 600 observations

Explanation

Hyperparameters are chosen using the validation set (20% of 1,000 = 200). The test set (also 200) is held out untouched for the final unbiased performance estimate. Reporting on training or tuning on test biases the result.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions