Skip to content

FRM Part I · FRM Exam Part I · Machine-Learning Methods

A risk team has a training sample of 1,000 observations for a credit scoring model. A feature 'missing_income' is present for 10% of records. The team fills missing income with the mean of income computed over the ENTIRE data set (training plus the held-out test set of 250 records) before splitting the data. Which statement best describes the problem?

This is data leakage. Using a mean computed over both training and test records lets test-set information influence preprocessing, making out-of-sample performance look better than it truly is. Imputation values should be estimated from the training data only and then applied unchanged to the test set.

  1. AIt causes data leakage, because information from the test set influences the training data and inflates the apparent out-of-sample performanceCorrect
  2. BIt introduces no problem because the mean is a simple summary statistic and cannot leak information
  3. CIt causes underfitting because mean imputation always increases model complexity
  4. DIt violates the requirement that imputation values must be the median rather than the mean

Explanation

Computing the imputation mean on the full data set lets test-set information enter the preprocessing used for training, which is data leakage and makes out-of-sample evaluation overly optimistic. The mean should be estimated on the training set only and then applied to the test set. Mean imputation does not increase complexity, and no rule mandates the median.

Did you get it right without looking?

One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.

More Machine-Learning Methods questions