Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

An analyst has 10,000 observations to forecast corporate bond downgrades. She first standardizes all features using the mean and standard deviation of the full dataset, and then splits the data into training and validation sets to tune hyperparameters. Which statement BEST describes the problem with this approach?

Data leakage occurs, so validation performance will tend to be overstated. Using the mean and standard deviation of the full dataset lets validation observations influence the transformation. Scaling parameters should be estimated from the training set only and then applied unchanged to validation and test data.

  1. AValidation information leaks into the training process, so validation performance will tend to be overstatedCorrect
  2. BStandardization is invalid for any dataset larger than 5,000 observations
  3. CThe training set will have zero variance after standardization
  4. DThe approach lowers training error but has no effect on validation error estimation

Explanation

Computing scaling parameters on the full dataset lets the validation data influence the transformation applied in training, which is data leakage. The correct process is to split first, compute parameters on the training set only, then apply them to validation and test sets. This makes validation estimates optimistic. The other statements are false.

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions