CFA Level II Exam · Machine Learning
LASSO, Penalized Regression, SVM and KNN for CFA Level II
Updated 7 October 2026 · Fact-checked
Penalized regression (such as LASSO) adds a penalty on coefficient size to reduce overfitting and drop weak features. A support vector machine (SVM) finds the widest-margin boundary between two classes. KNN classifies a point by the majority class of its k nearest neighbors. Match the algorithm to the vignette's clues.
Understand Supervised Learning: Penalized Regression and SVM
These three are supervised learning methods. You give the algorithm labeled data (features X and a known target Y), and it learns to predict Y for new data. The main risk is overfitting: the model fits noise in the training data and does badly on new data.
Penalized regression fights overfitting in a regression with many features. Ordinary least squares (OLS) minimizes the sum of squared errors. Penalized regression minimizes that sum plus a penalty that grows with the size of the coefficients. This forces the model to keep only features that earn their place. LASSO (least absolute shrinkage and selection operator) uses a penalty on the sum of absolute values of the coefficients, multiplied by a hyperparameter λ. A larger λ means a stronger penalty. Because of the absolute-value penalty, LASSO can shrink some coefficients exactly to zero, so it also does feature selection. Ridge uses squared coefficients, which shrinks them but does not usually set them to zero. Elastic net mixes both penalties. Penalized regression can also be used for classification (penalized logistic regression). Standardize features first, because the penalty depends on coefficient size, and size depends on units.
Support vector machine (SVM) is a classifier. For two classes, it finds the hyperplane (a linear boundary) that separates them with the widest possible margin. The points closest to the boundary are the support vectors; only they define the boundary. If classes overlap, a soft margin classification allows some misclassified points and adds a penalty for them. SVM works well when classes are not perfectly clean, uses few observations to fix the boundary, and can handle nonlinear boundaries with kernel functions. It is mainly used for classification, though it can be used for regression.
K-nearest neighbor (KNN) classifies a new observation by looking at the k closest training observations and taking the majority class. It can also predict a number by averaging neighbors. It needs a distance measure, and you choose k. A very small k tends to overfit (noisy boundary). A very large k tends to underfit (too smooth). KNN assumes similar observations are near each other, and it does not make assumptions about the data distribution. Outliers and irrelevant features hurt it, and features should be on similar scales.
Key formulas to remember
- LASSO objective
- Minimize Σ(Yᵢ − Ŷᵢ)² + λ × Σ|bₖ|
- Penalty is the sum of absolute coefficients. Larger λ means more shrinkage and more coefficients at zero. λ = 0 gives OLS. The intercept is normally not penalized.
- Ridge objective
- Minimize Σ(Yᵢ − Ŷᵢ)² + λ × Σbₖ²
- Squared penalty. Shrinks coefficients toward zero but rarely to exactly zero.
- Elastic net
- Minimize Σ(Yᵢ − Ŷᵢ)² + λ₁Σ|bₖ| + λ₂Σbₖ²
- Combines LASSO and ridge penalties.
- KNN rule
- Class of new point = majority class among its k nearest training points
- Small k risks overfitting; large k risks underfitting. Use an odd k with two classes to avoid ties.
- SVM boundary
- Choose the hyperplane that maximizes the margin between classes
- Support vectors are the points closest to the boundary. Soft margin allows some misclassification.
How to solve Supervised Learning: Penalized Regression and SVM questions
Use this method on any item-set question about penalized regression, SVM or KNN.
- 1Find the task in the vignette: predicting a number (regression) or assigning a category (classification).
- 2Find the clues: many features, correlated features, need to drop variables, overlapping classes, or small datasets.
- 3Match the algorithm: penalized regression for many features and overfitting; LASSO for automatic feature selection; SVM for a classification boundary with a margin; KNN for classification by similarity to nearby points.
- 4Read the exhibit: coefficients at zero show features LASSO removed; the number of support vectors or the neighbors' classes show how SVM or KNN decided.
- 5Apply the direction rule: higher λ means more penalty, simpler model, less variance and more bias; changing k works the opposite way (higher k is smoother).
- 6Check the answer against overfitting versus underfitting, and pick the option that fits the vignette's stated problem.
Quickest way: Keyword matching for the three algorithms
When to use it: When time is short and the question asks which method fits or what a parameter change does.
- Coefficients shrunk to zero or feature selection means LASSO.
- Many correlated features and a penalty on size means penalized regression (ridge or elastic net).
- Margin, hyperplane or support vectors means SVM.
- Majority vote of neighbors or distance means KNN.
- λ up or k up means simpler model. Expect less overfitting and more bias.
Common mistakes in Supervised Learning: Penalized Regression and SVM
Saying ridge sets coefficients to exactly zero.
Ridge and LASSO are both called penalized regression and get blended.
Fix: Only LASSO (absolute-value penalty) performs feature selection. Ridge shrinks without usually eliminating.
Thinking a larger λ gives a more complex model.
Students link a bigger number with more flexibility.
Fix: Larger λ means a stronger penalty, so a simpler model with fewer or smaller coefficients.
Believing small k reduces overfitting in KNN.
Students assume more local means more accurate.
Fix: k = 1 follows every noisy point and overfits. Raising k smooths the boundary but too high underfits.
Saying all training points define the SVM boundary.
Regression logic uses all observations.
Fix: Only the support vectors, the points closest to the boundary, determine it.
Ignoring feature scaling.
Students focus on the algorithm and not the data.
Fix: Penalties and distances depend on scale. Standardize features before LASSO or KNN.
Worked examples
Example 1
An analyst builds a model to forecast bond fund returns using 60 correlated macro features and 80 observations. OLS fits the training data almost perfectly but performs poorly on new data. She switches to LASSO and finds that 42 of the 60 coefficients are zero. (1) What problem did OLS show? (2) Why are 42 coefficients zero? (3) What happens if she raises λ?
Show the solution
- Part 1: near-perfect training fit but poor new-data results, with many features and few observations, is overfitting.
- Part 2: the LASSO penalty uses absolute values of coefficients, so it can shrink weak features exactly to zero. That is feature selection, leaving 18 features.
- Part 3: a higher λ strengthens the penalty, so more coefficients move to zero and the model becomes simpler, with lower variance and more bias.
Answer: (1) Overfitting. (2) The absolute-value penalty in LASSO shrinks weak coefficients to exactly zero. (3) More coefficients go to zero; the model gets simpler and may start to underfit if λ is too high.
Example 2
A lender uses KNN with k = 5 to classify an applicant as Default or No Default. The five nearest borrowers in the training set are: Default, No Default, Default, Default, No Default. (1) What class is predicted? (2) The lender sets k = 1 and the nearest borrower is No Default. What is predicted and what is the risk? (3) Another analyst proposes an SVM. What does it find?
Show the solution
- Part 1: count the votes. Default has 3, No Default has 2. Majority is Default.
- Part 2: with k = 1 only the single nearest point counts, so the prediction is No Default. A single noisy or unusual neighbor can drive the result, which is overfitting.
- Part 3: an SVM finds the hyperplane that separates Default from No Default with the maximum margin, defined by the support vectors, rather than voting among neighbors.
Answer: (1) Default (3 votes to 2). (2) No Default, with a high risk of overfitting to noise. (3) SVM finds the widest-margin separating boundary determined by the support vectors.
Exam tips
- Expect the vignette to describe a problem (overfitting, too many features, overlapping classes) and ask you to pick or justify an algorithm. Match the clue to the method.
- Memorize direction rules: λ up means simpler; k up means smoother. Most conceptual questions test only the direction.
- LASSO versus ridge is a common trap. Only LASSO zeroes out coefficients.
- In KNN questions, count the neighbors carefully and check you are using the stated k.
- Remember that SVM is defined by support vectors, not by all the data.
Supervised Learning: Penalized Regression and SVM in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Supervised Learning: Penalized Regression and SVM: frequently asked questions
How does LASSO reduce overfitting?
LASSO adds a penalty equal to λ times the sum of absolute coefficients to the error being minimized. This discourages large coefficients and pushes weak ones to zero, giving a simpler model that generalizes better.
What is the difference between LASSO and ridge regression?
LASSO penalizes absolute coefficient values and can set coefficients to exactly zero, so it selects features. Ridge penalizes squared coefficients and shrinks them without usually removing any. Elastic net combines both.
What is a support vector in SVM?
A support vector is a training point that lies closest to the separating hyperplane, on or inside the margin. These points alone define the boundary. Moving other points far from the margin does not change it.
How do you choose k in KNN?
k is a hyperparameter you set before training, usually by testing values on validation data. A small k tends to overfit and a large k tends to underfit. With two classes, an odd k avoids tied votes.