Skip to content

IAI Actuarial Core Principles · Risk Modelling and Survival Analysis · Elementary principles of machine learning

In ridge regression, a penalty equal to a positive constant λ times the sum of squared coefficients is added to the residual sum of squares. Compared with ordinary least squares, what is the effect of increasing λ?

Increasing λ shrinks the coefficients towards zero, which usually reduces variance at the cost of some bias. Ridge uses a squared penalty, so it rarely sets coefficients exactly to zero, which is a lasso feature, and its training error cannot beat ordinary least squares.

  1. ACoefficients are shrunk towards zero, which typically reduces variance at the cost of some added biasCorrect
  2. BCoefficients are set exactly to zero, giving automatic variable selection
  3. CCoefficients are inflated, which reduces bias and increases variance
  4. DThe model is guaranteed to have lower training error than OLS
  5. The penalty only affects the intercept

Explanation

The L2 penalty shrinks coefficients towards zero, trading a little bias for lower variance. Exact zeros are characteristic of the lasso (L1), not ridge. OLS minimises training error, so ridge cannot have lower training error. The intercept is normally not penalised.

Did you get it right without looking?

One question tells you little. A timed set on Elementary principles of machine learning shows your real accuracy, how long you take and where you lose marks.

More Elementary principles of machine learning questions