FRM Part I · FRM Exam Part I · Machine-Learning Methods
An analyst has 10,000 observations to forecast corporate bond downgrades. She first standardizes all features using the mean and standard deviation of the full dataset, and then splits the data into training and validation sets to tune hyperparameters. Which statement BEST describes the problem with this approach?
Data leakage occurs, so validation performance will tend to be overstated. Using the mean and standard deviation of the full dataset lets validation observations influence the transformation. Scaling parameters should be estimated from the training set only and then applied unchanged to validation and test data.
- AValidation information leaks into the training process, so validation performance will tend to be overstatedCorrect
- BStandardization is invalid for any dataset larger than 5,000 observations
- CThe training set will have zero variance after standardization
- DThe approach lowers training error but has no effect on validation error estimation
Explanation
Computing scaling parameters on the full dataset lets the validation data influence the transformation applied in training, which is data leakage. The correct process is to split first, compute parameters on the training set only, then apply them to validation and test sets. This makes validation estimates optimistic. The other statements are false.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- A node in a classification tree holds 40 observations: 30 non-defaults and 10 defaults. Using the Gini impurity, 1 minus the sum of squared …
- A ridge regression with one standardized predictor and no intercept has the OLS slope estimate of 1.20, where the predictor's sum of squares…
- A data scientist standardizes a feature using the training set, which has a mean of 40 and a standard deviation of 8. A test observation has…
- A model predicting loan losses achieves a mean squared error of 0.5 on the training sample but 4.0 on a held-out validation sample. Which in…
- A model predicting credit losses achieves a training mean squared error of 0.5 and a validation mean squared error of 4.8 on held-out data. …
- A deep neural network for credit scoring achieves very low training error but much higher validation error. Which single action is most dire…