FRM Part I · FRM Exam Part I · Machine Learning and Prediction
A risk team has 40 yield-curve and macro predictors and is predicting credit spread changes. They run PCA on the full dataset, keep the first 5 components, then use 10-fold cross-validation to estimate out-of-sample error of a regression on those components. Which is the most important concern with this procedure?
The main concern is data leakage: PCA was fitted on the full dataset, including observations later used as validation folds, so cross-validated error is likely understated. PCA should be estimated within each training fold only. Components are uncorrelated, and PCA suits correlated predictors.
- APCA was fitted using data that later serve as validation folds, so the cross-validation error may be understatedCorrect
- BPrincipal components are correlated, which invalidates the regression coefficients
- CKeeping only 5 components guarantees the model underfits
- DPCA cannot be used when predictors are correlated
Explanation
The PCA loadings were estimated using all observations, including those in each validation fold, so information leaks from validation into training and the error estimate is optimistically biased. The correct approach is to fit PCA within each training fold only. Principal components are in fact uncorrelated, and PCA works especially well with correlated data.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- Compared with a regression using all original predictors, a principal components regression (PCR) that uses the first few components has whi…
- A risk team has transaction data for 50,000 corporate clients with no labels. They run k-means to group clients with similar trading behavio…
- A risk manager clusters customers on two variables: annual transaction volume (in USD, ranging 0 to 2,000,000) and number of late payments (…
- A risk analyst fits a linear model to predict loan losses using 60 correlated explanatory variables and only 120 observations. The analyst w…
- A single neuron receives inputs x1 = 2 and x2 = -1 with weights w1 = 0.5 and w2 = 1.5 and bias b = 0.5. It uses a ReLU activation, f(z) = ma…
- A node in a classification tree holds 100 observations: 50 defaults and 50 non-defaults. A candidate split sends 40 observations left (35 de…