FRM Part I · FRM Exam Part I · Machine Learning and Prediction
A data scientist splits data into training, validation and test sets to build a default-prediction model. She tries several regularization strengths and picks the one with the lowest error. Which procedure best avoids biased estimates of future performance?
Tune the regularization strength on the validation set and report performance on the untouched test set. Using the test set for tuning leaks information and gives optimistic error estimates, while tuning on training data favors overly complex models.
- AChoose the regularization strength on the test set, then report test error
- BChoose the regularization strength on the training set, then report validation error
- CChoose the regularization strength on the validation set, then report error on the untouched test setCorrect
- DChoose the regularization strength on the test set, then retrain on the validation set
Explanation
Hyperparameters should be tuned using validation data, and the test set must remain unused until the final evaluation. Tuning on the test set leaks information and makes its error optimistic. Tuning on training data favors overfit settings because training error falls with complexity.
Did you get it right without looking?
One question tells you little. A timed set on Machine Learning and Prediction shows your real accuracy, how long you take and where you lose marks.
More Machine Learning and Prediction questions
- A bank's credit-default classifier is tested on 200 loans. It flags 50 loans as defaults, of which 40 actually defaulted. In total 60 loans …
- A neural network with many parameters achieves very low error on the training set but substantially higher error on a validation set. Which …
- A model predicts whether a trade is fraudulent. Of 1,000 trades, 20 are actually fraudulent. The model flags 25 trades, of which 15 are actu…
- A risk analyst fits a single, very deep classification tree to predict loan default. It classifies the training data almost perfectly but pe…
- A bank combines 100 decision trees by majority vote to classify counterparties as default or non-default. Compared with a single deep tree, …
- Three models are evaluated by 5-fold cross-validation. Mean squared errors on the five held-out folds are: Model A: 4, 6, 5, 7, 8; Model B: …