IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning
In ridge regression, a penalty equal to a positive constant λ times the sum of squared coefficients is added to the residual sum of squares. Compared with ordinary least squares, what is the effect of increasing λ?
Increasing λ shrinks the coefficients towards zero, which usually reduces variance at the cost of some bias. Ridge uses a squared penalty, so it rarely sets coefficients exactly to zero, which is a lasso feature, and its training error cannot beat ordinary least squares.
- ACoefficients are shrunk towards zero, which typically reduces variance at the cost of some added biasCorrect
- BCoefficients are set exactly to zero, giving automatic variable selection
- CCoefficients are inflated, which reduces bias and increases variance
- DThe model is guaranteed to have lower training error than OLS
- The penalty only affects the intercept
Explanation
The L2 penalty shrinks coefficients towards zero, trading a little bias for lower variance. Exact zeros are characteristic of the lasso (L1), not ridge. OLS minimises training error, so ridge cannot have lower training error. The intercept is normally not penalised.
Did you get it right without looking?
One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.
More Elementary principles of machine learning questions
- A health insurer groups 50,000 customers into segments using only age, income and claim frequency, with no predefined segment labels. Which …
- A pricing team fits a claim-severity model and finds that its error on the data used for fitting is very small, but its error on a separate …
- A model fitted to training data gives a very low training error but a much higher error on a separate validation set. What is the most likel…
- A modeller fits polynomial regressions of increasing degree to the same training data. The training error falls steadily with degree, but th…
- A dataset is split into training, validation and test sets to choose the number of neighbours k in a k-nearest-neighbours model. What is the…
- An insurer has a dataset of past motor policies, each labelled with whether a claim was made, and wants a model to predict claims for new po…