FRM Part I · FRM Exam Part I · Machine Learning and Prediction
In k-fold cross-validation with k = 5 applied to a sample of 500 observations, how many observations are used to train the model in each iteration, and how many times is each observation used for validation?
Each iteration trains on 400 observations and validates on the remaining 100. Because every observation belongs to exactly one fold, it is used for validation exactly once and for training in the other four iterations.
- A400 observations for training; each observation validated exactly onceCorrect
- B100 observations for training; each observation validated exactly once
- C400 observations for training; each observation validated 5 times
- D100 observations for training; each observation validated 4 times
Explanation
With 5 folds of 100 observations each, every iteration trains on the other 4 folds, which is 400 observations, and validates on the held-out fold of 100. Each observation sits in exactly one fold, so it is used for validation exactly once and for training in the other four iterations.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- A fraud model is evaluated on 1,000 transactions. The confusion matrix shows: true positives 30, false positives 20, false negatives 10, tru…
- An analyst uses ridge regression with one predictor and no intercept, where the predictor has been scaled so that the sum of x squared equal…
- A bank uses the first three principal components of 12 equity factor returns as regressors in a model to predict portfolio losses (principal…
- Compared with a logistic regression using the same inputs, a deep neural network used for a bank's default prediction is most likely to pres…
- A risk team has 40 yield-curve and macro predictors and is predicting credit spread changes. They run PCA on the full dataset, keep the firs…
- A risk team runs agglomerative hierarchical clustering on five funds using Euclidean distance. The first merge joins Funds A and B, which ar…