Skip to content

CFA Level II · CFA Level II Exam

Machine Learning: formula sheet

Full chapter guide

Key formulas

Generalization error decomposition
Out-of-sample error = bias error + variance error + base error
Bias is from a model too simple. Variance is from a model too sensitive to the training sample. Base error is random noise that no model removes.
Overfitting signature
Low training error + high validation/test error
Underfitting shows high error in both training and validation samples.
k-fold cross-validation
Each of k folds is the validation set once; training uses k − 1 folds; average the k results
Used when data is limited. Each observation is used for validation exactly once.
Learning type rule
Labeled target → supervised; no target → unsupervised
Continuous target = regression; categorical target = classification.
LASSO objective
Minimize Σ(Yᵢ − Ŷᵢ)² + λ × Σ|bₖ|
Penalty is the sum of absolute coefficients. Larger λ means more shrinkage and more coefficients at zero. λ = 0 gives OLS. The intercept is normally not penalized.
Ridge objective
Minimize Σ(Yᵢ − Ŷᵢ)² + λ × Σbₖ²
Squared penalty. Shrinks coefficients toward zero but rarely to exactly zero.
Elastic net
Minimize Σ(Yᵢ − Ŷᵢ)² + λ₁Σ|bₖ| + λ₂Σbₖ²
Combines LASSO and ridge penalties.
KNN rule
Class of new point = majority class among its k nearest training points
Small k risks overfitting; large k risks underfitting. Use an odd k with two classes to avoid ties.
SVM boundary
Choose the hyperplane that maximizes the margin between classes
Support vectors are the points closest to the boundary. Soft margin allows some misclassification.
Classification tree leaf prediction
Predicted class = most common class among training observations in the leaf
Used for categorical targets such as default or no default.
Regression tree leaf prediction
Predicted value = average of the target values in the leaf
Used for continuous targets such as returns.
Bagging combination rule
Classification: majority vote of the models. Regression: average of the model predictions
Each model is trained on a bootstrap sample drawn with replacement.
Random forest feature rule
At each split, consider only a random subset of the features
Lowers correlation between trees. The usual aim is lower variance.
Boosting idea
Models are trained in sequence; each focuses on prior errors
Aimed at reducing bias. AdaBoost reweights misclassified observations; gradient boosting fits the residual errors.
Overfitting check
Training error low and validation error much higher means overfitting
Pruning or stopping rules reduce tree complexity.
Proportion of variance explained by a component
Proportion for PCj = eigenvalue of PCj ÷ Σ of all eigenvalues
Eigenvalues sum to total variance. The proportions across all components sum to 100%.
Cumulative variance explained
Cumulative proportion for first m components = Σ (proportions of PC1 to PCm)
Use this to decide how many components to keep for a target level such as 90%.
Euclidean distance between two points
d = √[(x1 − y1)² + (x2 − y2)² + … + (xn − yn)²]
The usual similarity measure in clustering. Smaller distance means more similar.
K-means centroid update
New centroid = average of each feature across all points assigned to the cluster
Repeat assignment and update until no point changes cluster.
PCA properties
Components are uncorrelated; PC1 variance ≥ PC2 variance ≥ PC3 variance …
Eigenvalues are ordered from largest to smallest.
Node calculation
Node output = f(Σ wᵢxᵢ + b)
wᵢ are weights, xᵢ inputs, b the bias, f the activation function. The curriculum describes this as a summation operator followed by an activation function.
Sigmoid activation
f(x) = 1 ÷ (1 + e^(-x))
Output lies between 0 and 1. Useful when the output is read as a probability.
ReLU activation
f(x) = max(0, x)
Zero for negative inputs, equal to x for positive inputs. Simple and widely used in hidden layers.
Layer structure rule
Input nodes = number of features; hidden layers ≥ 1; DNN = many hidden layers
The output layer matches the task (one node for a single continuous prediction, for example).
RL objective
Maximise expected cumulative (discounted) reward
The agent learns a policy for choosing actions given the state.

Quick revision

  • Supervised learning uses labelled data; unsupervised learning finds structure without labels; reinforcement learning learns from rewards.
  • Overfitting means a model fits training data too closely and performs poorly on new data.
  • High bias means underfitting; high variance means overfitting.
  • Data is split into training, validation and test sets; the test set is used only for final evaluation.
  • Penalized regression adds a penalty on coefficient size; LASSO can shrink coefficients to exactly zero, so it selects features.
  • SVM finds the boundary that gives the widest margin between classes.
  • CART builds a tree of splits; a deep tree overfits, and pruning reduces this.
  • Bagging averages many models trained on resampled data; boosting builds models in sequence to fix earlier errors.
  • A random forest is a bagged ensemble of decision trees, each trained on a bootstrap sample and using a random subset of features at each split.
  • PCA reduces many correlated features into fewer uncorrelated components; clustering groups similar observations.
  • K-means needs you to choose the number of clusters; hierarchical clustering does not.
  • Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN); accuracy = (TP + TN) ÷ all observations.

Common mistakes

  • Calling clustering a supervised method because groups are the output. Fix: Supervised needs labels in the training data. Clustering creates groups without labels, so it is unsupervised.
  • Saying an overfit model has high bias. Fix: Overfit means low bias, high variance. Underfit means high bias, low variance.
  • Saying ridge sets coefficients to exactly zero. Fix: Only LASSO (absolute-value penalty) performs feature selection. Ridge shrinks without usually eliminating.
  • Thinking a larger λ gives a more complex model. Fix: Larger λ means a stronger penalty, so a simpler model with fewer or smaller coefficients.
  • Saying a fully grown tree is the best model because it has zero training error. Fix: Training error is not the test. A tree grown until every leaf is pure usually overfits. Judge by performance on validation or test data.
  • Mixing up bagging and boosting. Fix: Bagging trains models independently on bootstrap samples and averages them. Boosting trains models in sequence, each correcting earlier errors.
  • Treating PCA as a method that selects the best original features. Fix: PCA creates new variables that are combinations of all original features. It does not drop original features one by one.
  • Forgetting that principal components are uncorrelated. Fix: Inputs are correlated; components are orthogonal and uncorrelated with each other.
  • Calling reinforcement learning a type of supervised learning because it uses feedback. Fix: A reward is a delayed signal from the environment, not a correct answer supplied for each input. RL has no labelled dataset.
  • Thinking the activation function is optional for a nonlinear model. Fix: Without a nonlinear activation, stacked layers collapse to a linear model. The activation provides the nonlinearity.

Exam tips

  • Start every item by asking whether the data has a labeled target. This settles most classification questions.
  • Quote the training-validation gap in your reasoning. Exam options often differ only in bias vs variance wording.
  • Remember the sample roles: train, validate, test. A trap option lets the test sample guide tuning.
  • Do not assume deep learning is always supervised or always unsupervised. Use the vignette detail.
  • In k-fold questions, count carefully: k runs, k − 1 training folds each.
  • Expect the vignette to describe a problem (overfitting, too many features, overlapping classes) and ask you to pick or justify an algorithm. Match the clue to the method.
  • Memorize direction rules: λ up means simpler; k up means smoother. Most conceptual questions test only the direction.
  • LASSO versus ridge is a common trap. Only LASSO zeroes out coefficients.