Skip to content

FRM Part I · FRM Exam Part I · Machine Learning and Prediction

A risk team has 40 yield-curve and macro predictors and is predicting credit spread changes. They run PCA on the full dataset, keep the first 5 components, then use 10-fold cross-validation to estimate out-of-sample error of a regression on those components. Which is the most important concern with this procedure?

The main concern is data leakage: PCA was fitted on the full dataset, including observations later used as validation folds, so cross-validated error is likely understated. PCA should be estimated within each training fold only. Components are uncorrelated, and PCA suits correlated predictors.

  1. APCA was fitted using data that later serve as validation folds, so the cross-validation error may be understatedCorrect
  2. BPrincipal components are correlated, which invalidates the regression coefficients
  3. CKeeping only 5 components guarantees the model underfits
  4. DPCA cannot be used when predictors are correlated

Explanation

The PCA loadings were estimated using all observations, including those in each validation fold, so information leaks from validation into training and the error estimate is optimistically biased. The correct approach is to fit PCA within each training fold only. Principal components are in fact uncorrelated, and PCA works especially well with correlated data.

Did you get it right without looking?

One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.

More Machine Learning and Prediction questions