FRM Part I · FRM Exam Part I · Machine-Learning Methods
A analyst compares penalized regressions on a data set with many correlated predictors, some of which are believed to be irrelevant. She wants a method that can set some coefficients exactly to zero, performing variable selection. Which statement is correct?
LASSO can set coefficients exactly to zero because its L1 penalty on the sum of absolute coefficients creates corner solutions, giving variable selection. Ridge uses a squared (L2) penalty that shrinks coefficients toward zero but generally leaves all predictors in the model.
- ARidge regression sets coefficients exactly to zero because it penalizes squared coefficients
- BLASSO can set coefficients exactly to zero because it penalizes the sum of absolute coefficient valuesCorrect
- CBoth ridge and LASSO always keep every predictor in the model
- DNeither method can reduce the number of predictors
Explanation
LASSO uses an L1 penalty (sum of absolute values), whose geometry produces corner solutions with some coefficients exactly zero. Ridge uses an L2 penalty, which shrinks coefficients toward zero but typically not exactly to zero, so option A is wrong.
Did you get it right without looking?
One question tells you little. A timed set on Machine-Learning Methods shows your real accuracy, how long you take and where you lose marks.
More Machine-Learning Methods questions
- A data scientist standardizes a feature using the training set, which has a mean of 40 and a standard deviation of 8. A test observation has…
- A model predicting credit losses achieves a training mean squared error of 0.5 and a validation mean squared error of 4.8 on held-out data. …
- A deep neural network for credit scoring achieves very low training error but much higher validation error. Which single action is most dire…
- In k-means clustering, an analyst increases k from 3 to 4 on the same dataset and re-runs the algorithm to convergence. Which outcome is mos…
- A single weight w in a network is updated by gradient descent. The current weight is 0.80, the learning rate is 0.10, and the partial deriva…
- A data scientist uses 5-fold cross-validation on a dataset of 1,000 observations to select a model. Fold validation mean squared errors are …