FRM Part I · FRM Exam Part I · Machine-Learning Methods
A risk team has a training sample of 1,000 observations for a credit scoring model. A feature 'missing_income' is present for 10% of records. The team fills missing income with the mean of income computed over the ENTIRE data set (training plus the held-out test set of 250 records) before splitting the data. Which statement best describes the problem?
This is data leakage. Using a mean computed over both training and test records lets test-set information influence preprocessing, making out-of-sample performance look better than it truly is. Imputation values should be estimated from the training data only and then applied unchanged to the test set.
- AIt causes data leakage, because information from the test set influences the training data and inflates the apparent out-of-sample performanceCorrect
- BIt introduces no problem because the mean is a simple summary statistic and cannot leak information
- CIt causes underfitting because mean imputation always increases model complexity
- DIt violates the requirement that imputation values must be the median rather than the mean
Explanation
Computing the imputation mean on the full data set lets test-set information enter the preprocessing used for training, which is data leakage and makes out-of-sample evaluation overly optimistic. The mean should be estimated on the training set only and then applied to the test set. Mean imputation does not increase complexity, and no rule mandates the median.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- A model predicting loan losses achieves a mean squared error of 0.5 on the training sample but 4.0 on a held-out validation sample. Which in…
- In a KNN classifier, an analyst moves from K = 1 to K = 25 on a noisy credit dataset. Which is the expected effect?
- A bank has transaction records for 200,000 corporate clients with no predefined categories. The risk team applies k-means to group clients w…
- A ridge regression with one standardized predictor and no intercept has the OLS slope estimate of 1.20, where the predictor's sum of squares…
- A default classifier is tested on 200 loans. Results: 30 true positives, 10 false positives, 20 false negatives, and 140 true negatives. Wha…
- A data scientist standardizes a feature using the training set, which has a mean of 40 and a standard deviation of 8. A test observation has…