FRM Exam Part I · Machine-Learning Methods
Ridge, LASSO and Elastic Net Regression Explained
Updated 11 October 2026 · Fact-checked
Ridge, LASSO and Elastic Net are penalized regressions. They add a penalty on coefficient size to the squared-error loss. Ridge shrinks coefficients toward zero but keeps all variables. LASSO can set coefficients exactly to zero, so it selects variables. Elastic Net mixes both. Lambda sets penalty strength and is chosen by cross-validation.
Understand Regularization: Ridge, LASSO and Elastic Net
Ordinary least squares picks coefficients that minimize the sum of squared errors. With many predictors, or predictors that are highly correlated, OLS fits noise in the training data. Coefficients become large and unstable, and out-of-sample performance is poor. This is overfitting.
Regularization fixes this by adding a penalty for large coefficients. The model now minimizes squared errors plus a penalty term. A bit of bias is accepted in exchange for a larger drop in variance. This is the bias-variance tradeoff at work.
Ridge (L2) penalizes the sum of squared coefficients. It shrinks all coefficients smoothly toward zero, but none reach exactly zero. It works well when many predictors each carry some signal and when predictors are correlated. LASSO (L1) penalizes the sum of absolute coefficients. The shape of this penalty lets some coefficients land exactly at zero, so LASSO does variable selection and gives a sparser, easier-to-explain model.
Elastic Net combines the L1 and L2 penalties. It selects variables like LASSO but handles groups of correlated predictors better. LASSO tends to pick one from a correlated group and drop the others; Elastic Net tends to keep them together.
The penalty parameter λ controls strength. λ = 0 gives plain OLS. As λ grows, coefficients shrink more, and the model becomes simpler. Choose λ by k-fold cross-validation: try many values, and pick the one with the lowest validation error. Standardize predictors first, because the penalty depends on coefficient size, which depends on units. The intercept is normally not penalized.
Key formulas to remember
- OLS objective
- minimize Σ(yᵢ − ŷᵢ)²
- Baseline with no penalty. Equals ridge or LASSO when λ = 0.
- Ridge objective
- minimize Σ(yᵢ − ŷᵢ)² + λ Σ βⱼ²
- L2 penalty. Shrinks coefficients, never exactly to zero (for finite λ). Sum runs over slopes, not the intercept.
- LASSO objective
- minimize Σ(yᵢ − ŷᵢ)² + λ Σ |βⱼ|
- L1 penalty. Can set coefficients exactly to zero, so it selects variables.
- Elastic Net objective
- minimize Σ(yᵢ − ŷᵢ)² + λ₁ Σ |βⱼ| + λ₂ Σ βⱼ²
- Mix of L1 and L2 penalties. Both λ values are chosen by cross-validation.
- Effect of λ
- λ = 0 → OLS; λ → ∞ → all slopes → 0
- Larger λ means more bias, less variance, a simpler model.
How to solve Regularization: Ridge, LASSO and Elastic Net questions
Use this routine for conceptual and calculation questions on penalized regression.
- 1Identify the penalty: squares of coefficients (ridge), absolute values (LASSO), or both (Elastic Net).
- 2Write the objective as squared errors plus λ times the penalty. Check whether the intercept is excluded.
- 3If asked to compute, find the sum of squared errors first, then the penalty from the given coefficients, then add them.
- 4If asked about behavior, link it to λ: small λ is close to OLS and high variance; large λ means more shrinkage, more bias, lower variance.
- 5If asked about variable selection, only L1-based methods (LASSO, Elastic Net) give exact zeros.
- 6If asked how to choose λ, answer k-fold cross-validation using out-of-sample error, not training error.
- 7Check for correlated predictors or standardization cues, which point to ridge or Elastic Net and to scaling first.
Quickest way: Penalty-shape shortcut
When to use it: Use for multiple-choice questions comparing the three methods or describing the effect of λ.
- Squared coefficients: ridge. Shrinks, keeps all variables.
- Absolute coefficients: LASSO. Sparse, selects variables.
- Both penalties: Elastic Net. Sparse and stable with correlated features.
- Raising λ: more bias, less variance. Pick λ by cross-validation.
- For calculations, compute SSE + λ × penalty and compare, with no heavy algebra.
Common mistakes in Regularization: Ridge, LASSO and Elastic Net
Saying ridge performs variable selection.
Both methods shrink coefficients, so they seem alike.
Fix: Ridge shrinks toward zero but does not set coefficients exactly to zero. Only LASSO and Elastic Net do.
Thinking a larger λ always improves the model.
Students link more regularization with less overfitting.
Fix: Too large a λ underfits: high bias. The best λ minimizes validation error, found by cross-validation.
Choosing λ by minimizing training error.
Training error is easy to compute.
Fix: Training error is lowest at λ = 0. Use held-out or cross-validated error.
Penalizing the intercept or forgetting to standardize predictors.
The penalty formula looks like it covers every parameter.
Fix: Penalty applies to slopes only. Standardize predictors so units do not drive the shrinkage.
Saying regularization makes estimates unbiased.
Confusing lower variance with better estimates in every sense.
Fix: Penalized estimates are biased on purpose. The gain is lower variance and better out-of-sample prediction.
Worked examples
Example 1
A ridge regression has two slope coefficients, β₁ = 2 and β₂ = −3, with λ = 0.5. The sum of squared errors is 40. Compute the ridge objective value.
Show the solution
- Ridge objective = SSE + λ Σ βⱼ².
- Penalty sum: 2² + (−3)² = 4 + 9 = 13.
- Penalty term: 0.5 × 13 = 6.5.
- Objective = 40 + 6.5 = 46.5.
Answer: 46.5
Example 2
A LASSO model uses β₁ = 1.5, β₂ = −0.5 and β₃ = 0 with λ = 4, and has SSE of 12. Compute the LASSO objective, and state what β₃ = 0 implies.
Show the solution
- LASSO objective = SSE + λ Σ |βⱼ|.
- Sum of absolute values: 1.5 + 0.5 + 0 = 2.0.
- Penalty term: 4 × 2.0 = 8.
- Objective = 12 + 8 = 20.
- β₃ = 0 means the third variable has been dropped from the model, which is LASSO's variable selection.
Answer: Objective = 20; the third variable is excluded by LASSO.
Exam tips
- Expect conceptual questions: which method selects variables, which handles correlated predictors, what λ does.
- Memorize the penalty forms: squares for ridge, absolute values for LASSO. Many calculation questions need only these.
- Tie regularization to overfitting and the bias-variance tradeoff; questions often combine them.
- When the question mentions choosing λ, the answer is cross-validation on out-of-sample error.
- Watch wording: ridge shrinks but does not eliminate; LASSO can eliminate.
Practice questions from Machine-Learning Methods
- A risk analyst fits a single unpruned decision tree to predict loan default and finds that it classifies the training data almost perfectly …
- Which statement best distinguishes a random forest from simple bagging of decision trees?
- A node in a classification tree holds 40 observations: 30 non-defaults and 10 defaults. Using the Gini impurity, 1 minus the sum of squared …
- A neuron has two inputs x1 = 2 and x2 = -1, weights w1 = 0.5 and w2 = 1.5, and bias b = 0.25. The neuron uses a ReLU activation, f(z) = max(…
- A risk team chooses the penalty parameter lambda for an elastic net model used to predict loan defaults. Which procedure is most appropriate…
Regularization: Ridge, LASSO and Elastic Net: frequently asked questions
What is the difference between ridge and LASSO regression?
Ridge adds a penalty on squared coefficients and shrinks them without making them exactly zero. LASSO penalizes absolute values and can set coefficients to exactly zero, so it selects variables. Ridge suits many small effects; LASSO suits sparse models.
How is the penalty parameter lambda chosen?
Use k-fold cross-validation. Fit the model across a range of λ values, measure error on the held-out folds, and pick the λ with the lowest validation error. Training error cannot be used, since it always favors λ = 0.
How does regularization reduce overfitting?
It penalizes large coefficients, which constrains model complexity. This adds some bias but cuts variance, so predictions on new data are usually more stable and accurate.
When should I use Elastic Net?
Use it when predictors are highly correlated and you also want variable selection. LASSO may pick one variable from a correlated group at random, while Elastic Net tends to keep the group together.