CFA Level II · CFA Level II Exam
Machine Learning: formula sheet
Key formulas
- Generalization error decomposition
- Out-of-sample error = bias error + variance error + base error
- Bias is from a model too simple. Variance is from a model too sensitive to the training sample. Base error is random noise that no model removes.
- Overfitting signature
- Low training error + high validation/test error
- Underfitting shows high error in both training and validation samples.
- k-fold cross-validation
- Each of k folds is the validation set once; training uses k − 1 folds; average the k results
- Used when data is limited. Each observation is used for validation exactly once.
- Learning type rule
- Labeled target → supervised; no target → unsupervised
- Continuous target = regression; categorical target = classification.
- LASSO objective
- Minimize Σ(Yᵢ − Ŷᵢ)² + λ × Σ|bₖ|
- Penalty is the sum of absolute coefficients. Larger λ means more shrinkage and more coefficients at zero. λ = 0 gives OLS. The intercept is normally not penalized.
- Ridge objective
- Minimize Σ(Yᵢ − Ŷᵢ)² + λ × Σbₖ²
- Squared penalty. Shrinks coefficients toward zero but rarely to exactly zero.
- Elastic net
- Minimize Σ(Yᵢ − Ŷᵢ)² + λ₁Σ|bₖ| + λ₂Σbₖ²
- Combines LASSO and ridge penalties.
- KNN rule
- Class of new point = majority class among its k nearest training points
- Small k risks overfitting; large k risks underfitting. Use an odd k with two classes to avoid ties.
- SVM boundary
- Choose the hyperplane that maximizes the margin between classes
- Support vectors are the points closest to the boundary. Soft margin allows some misclassification.
- Classification tree leaf prediction
- Predicted class = most common class among training observations in the leaf
- Used for categorical targets such as default or no default.
- Regression tree leaf prediction
- Predicted value = average of the target values in the leaf
- Used for continuous targets such as returns.
- Bagging combination rule
- Classification: majority vote of the models. Regression: average of the model predictions
- Each model is trained on a bootstrap sample drawn with replacement.
- Random forest feature rule
- At each split, consider only a random subset of the features
- Lowers correlation between trees. The usual aim is lower variance.
- Boosting idea
- Models are trained in sequence; each focuses on prior errors
- Aimed at reducing bias. AdaBoost reweights misclassified observations; gradient boosting fits the residual errors.
- Overfitting check
- Training error low and validation error much higher means overfitting
- Pruning or stopping rules reduce tree complexity.
- Proportion of variance explained by a component
- Proportion for PCj = eigenvalue of PCj ÷ Σ of all eigenvalues
- Eigenvalues sum to total variance. The proportions across all components sum to 100%.
- Cumulative variance explained
- Cumulative proportion for first m components = Σ (proportions of PC1 to PCm)
- Use this to decide how many components to keep for a target level such as 90%.
- Euclidean distance between two points
- d = √[(x1 − y1)² + (x2 − y2)² + … + (xn − yn)²]
- The usual similarity measure in clustering. Smaller distance means more similar.
- K-means centroid update
- New centroid = average of each feature across all points assigned to the cluster
- Repeat assignment and update until no point changes cluster.
- PCA properties
- Components are uncorrelated; PC1 variance ≥ PC2 variance ≥ PC3 variance …
- Eigenvalues are ordered from largest to smallest.
- Node calculation
- Node output = f(Σ wᵢxᵢ + b)
- wᵢ are weights, xᵢ inputs, b the bias, f the activation function. The curriculum describes this as a summation operator followed by an activation function.
- Sigmoid activation
- f(x) = 1 ÷ (1 + e^(-x))
- Output lies between 0 and 1. Useful when the output is read as a probability.
- ReLU activation
- f(x) = max(0, x)
- Zero for negative inputs, equal to x for positive inputs. Simple and widely used in hidden layers.
- Layer structure rule
- Input nodes = number of features; hidden layers ≥ 1; DNN = many hidden layers
- The output layer matches the task (one node for a single continuous prediction, for example).
- RL objective
- Maximise expected cumulative (discounted) reward
- The agent learns a policy for choosing actions given the state.
Quick revision
- Supervised learning uses labelled data; unsupervised learning finds structure without labels; reinforcement learning learns from rewards.
- Overfitting means a model fits training data too closely and performs poorly on new data.
- High bias means underfitting; high variance means overfitting.
- Data is split into training, validation and test sets; the test set is used only for final evaluation.
- Penalized regression adds a penalty on coefficient size; LASSO can shrink coefficients to exactly zero, so it selects features.
- SVM finds the boundary that gives the widest margin between classes.
- CART builds a tree of splits; a deep tree overfits, and pruning reduces this.
- Bagging averages many models trained on resampled data; boosting builds models in sequence to fix earlier errors.
- A random forest is a bagged ensemble of decision trees, each trained on a bootstrap sample and using a random subset of features at each split.
- PCA reduces many correlated features into fewer uncorrelated components; clustering groups similar observations.
- K-means needs you to choose the number of clusters; hierarchical clustering does not.
- Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN); accuracy = (TP + TN) ÷ all observations.
Common mistakes
- Calling clustering a supervised method because groups are the output. Fix: Supervised needs labels in the training data. Clustering creates groups without labels, so it is unsupervised.
- Saying an overfit model has high bias. Fix: Overfit means low bias, high variance. Underfit means high bias, low variance.
- Saying ridge sets coefficients to exactly zero. Fix: Only LASSO (absolute-value penalty) performs feature selection. Ridge shrinks without usually eliminating.
- Thinking a larger λ gives a more complex model. Fix: Larger λ means a stronger penalty, so a simpler model with fewer or smaller coefficients.
- Saying a fully grown tree is the best model because it has zero training error. Fix: Training error is not the test. A tree grown until every leaf is pure usually overfits. Judge by performance on validation or test data.
- Mixing up bagging and boosting. Fix: Bagging trains models independently on bootstrap samples and averages them. Boosting trains models in sequence, each correcting earlier errors.
- Treating PCA as a method that selects the best original features. Fix: PCA creates new variables that are combinations of all original features. It does not drop original features one by one.
- Forgetting that principal components are uncorrelated. Fix: Inputs are correlated; components are orthogonal and uncorrelated with each other.
- Calling reinforcement learning a type of supervised learning because it uses feedback. Fix: A reward is a delayed signal from the environment, not a correct answer supplied for each input. RL has no labelled dataset.
- Thinking the activation function is optional for a nonlinear model. Fix: Without a nonlinear activation, stacked layers collapse to a linear model. The activation provides the nonlinearity.
Exam tips
- Start every item by asking whether the data has a labeled target. This settles most classification questions.
- Quote the training-validation gap in your reasoning. Exam options often differ only in bias vs variance wording.
- Remember the sample roles: train, validate, test. A trap option lets the test sample guide tuning.
- Do not assume deep learning is always supervised or always unsupervised. Use the vignette detail.
- In k-fold questions, count carefully: k runs, k − 1 training folds each.
- Expect the vignette to describe a problem (overfitting, too many features, overlapping classes) and ask you to pick or justify an algorithm. Match the clue to the method.
- Memorize direction rules: λ up means simpler; k up means smoother. Most conceptual questions test only the direction.
- LASSO versus ridge is a common trap. Only LASSO zeroes out coefficients.