FRM Part I · FRM Exam Part I
Machine-Learning Methods: formula sheet
Key formulas
- Supervised learning
- Inputs (features) + known labels → learn a mapping to predict the label
- Numeric label = regression. Categorical label = classification.
- Unsupervised learning
- Inputs (features) only, no labels → find clusters or lower-dimensional structure
- Tasks: clustering and dimension reduction.
- Reinforcement learning
- Agent takes action in a state → receives reward → updates policy to maximize cumulative reward
- Sequential decisions; feedback is a reward, not a correct label.
- Prediction versus inference
- ML: emphasis on out-of-sample prediction. Traditional statistics: emphasis on estimating and testing parameters
- A guide to the usual emphasis, not a strict rule. Many methods serve both.
- Expected prediction error decomposition
- Expected error = Bias² + Variance + Irreducible error
- Irreducible error is noise no model can remove. Only bias and variance change with model choice.
- Cross-validation error
- CV error = (1 ÷ k) × Σ (error on fold i), i = 1 to k
- Each fold is the validation set once. This is a simple average when folds are equal in size.
- Observations per fold
- Fold size = n ÷ k
- Each round trains on n × (k − 1) ÷ k observations.
- Complexity pattern
- Complexity ↑ → Bias ↓, Variance ↑, training error ↓
- Validation error is U-shaped: it falls, then rises once overfitting starts.
- Data roles
- Training = fit; Validation = tune/select; Test = final unbiased estimate
- Use the test set once. Using it to tune makes it part of training.
- Standardization (z-score)
- z = (x − μ) ÷ σ
- μ and σ come from the training set. Result has mean 0 and standard deviation 1. It does not make the data normal.
- Min-max normalization
- x' = (x − min) ÷ (max − min)
- Maps training data to 0 to 1. Very sensitive to outliers because min and max are extreme values.
- Mean imputation
- x_missing = (Σ x_observed) ÷ n_observed
- Keeps the mean unchanged but shrinks variance. Use the median for skewed data.
- Interquartile range outlier rule
- Outlier if x < Q1 − 1.5 × IQR or x > Q3 + 1.5 × IQR, where IQR = Q3 − Q1
- A common rule of thumb, not a law. Flag for review rather than delete automatically.
- OLS objective
- minimize Σ(yᵢ − ŷᵢ)²
- Baseline with no penalty. Equals ridge or LASSO when λ = 0.
- Ridge objective
- minimize Σ(yᵢ − ŷᵢ)² + λ Σ βⱼ²
- L2 penalty. Shrinks coefficients, never exactly to zero (for finite λ). Sum runs over slopes, not the intercept.
- LASSO objective
- minimize Σ(yᵢ − ŷᵢ)² + λ Σ |βⱼ|
- L1 penalty. Can set coefficients exactly to zero, so it selects variables.
- Elastic Net objective
- minimize Σ(yᵢ − ŷᵢ)² + λ₁ Σ |βⱼ| + λ₂ Σ βⱼ²
- Mix of L1 and L2 penalties. Both λ values are chosen by cross-validation.
- Effect of λ
- λ = 0 → OLS; λ → ∞ → all slopes → 0
- Larger λ means more bias, less variance, a simpler model.
- Logistic function
- P(y = 1) = 1 ÷ (1 + e^(−z)), where z = b0 + b1x1 + ... + bkxk
- Output always lies between 0 and 1.
- Log-odds (logit)
- ln[p ÷ (1 − p)] = b0 + b1x1 + ... + bkxk
- Coefficients are linear in log-odds, not in probability. Odds = p ÷ (1 − p).
- Accuracy
- (TP + TN) ÷ (TP + TN + FP + FN)
- Misleading when one class is rare.
- Precision
- TP ÷ (TP + FP)
- Denominator is everything predicted positive.
- Recall (sensitivity, TPR)
- TP ÷ (TP + FN)
- Denominator is everything actually positive.
- Specificity
- TN ÷ (TN + FP)
- True negative rate.
- False positive rate
- FP ÷ (FP + TN) = 1 − specificity
- The x-axis of the ROC curve.
- F1 score
- 2 × Precision × Recall ÷ (Precision + Recall) = 2TP ÷ (2TP + FP + FN)
- Harmonic mean; stays low if either precision or recall is low.
- AUC interpretation
- AUC = 0.5: random; AUC = 1: perfect
- Higher AUC means better ranking of positives above negatives across thresholds.
- Gini impurity
- G = 1 − Σ pᵢ²
- pᵢ is the share of class i in the node. For two classes, the maximum is 0.5 and a pure node gives 0.
- Entropy
- H = − Σ pᵢ × ln(pᵢ)
- Zero for a pure node. Higher means more mixed. Log base (2 or e) only rescales it.
- Regression tree leaf prediction
- ŷ = average of y in the leaf
- Splits are chosen to minimise the sum of squared errors across the two child nodes.
- Bagging prediction
- ŷ = (1 ÷ B) × Σ ŷᵦ
- B bootstrap trees. For classification use a majority vote.
- Euclidean distance
- d = √[Σ (xᵢ − zᵢ)²]
- Used by KNN. Scale features first or large-valued ones dominate.
- Weighted impurity after a split
- G_split = (n_L ÷ n) × G_L + (n_R ÷ n) × G_R
- Pick the split with the lowest value, which is the largest impurity reduction.
- Euclidean distance
- d(x, y) = √[Σ (xᵢ − yᵢ)²]
- Standard distance for k-means. Scale features first.
- Centroid
- centroid = average of each feature across the points in the cluster
- Recomputed after every assignment step in k-means.
- Within-cluster sum of squares (inertia)
- WCSS = Σ over clusters Σ over points ‖x − centroid‖²
- K-means minimizes this. It falls as k rises, so use the elbow, not the minimum.
- Variance explained by a component
- share of component j = λⱼ ÷ Σ λ
- λ are eigenvalues of the covariance (or correlation) matrix.
- Cumulative variance explained
- (λ₁ + … + λₘ) ÷ Σ λ
- Use to decide how many components to keep.
- Principal component
- PCⱼ = w₁ⱼX₁ + w₂ⱼX₂ + … + wₙⱼXₙ
- Weights (loadings) form a unit-length eigenvector. Components are uncorrelated.
- Node output
- a = f(z), where z = w₁x₁ + w₂x₂ + … + wₙxₙ + b
- w are weights, b is the bias, f is the activation function.
- Sigmoid activation
- σ(z) = 1 ÷ (1 + e^(−z))
- Output lies between 0 and 1. Often used for probabilities. Can cause vanishing gradients for large |z|.
- ReLU activation
- ReLU(z) = max(0, z)
- Simple and widely used in hidden layers. Output is zero for negative z.
- Tanh activation
- tanh(z) = (e^z − e^(−z)) ÷ (e^z + e^(−z))
- Output lies between −1 and 1 and is centred on zero.
- Mean squared error loss
- MSE = (1 ÷ n) × Σ (yᵢ − ŷᵢ)²
- Typical loss for numeric targets.
- Gradient descent update
- w_new = w_old − η × ∂L/∂w
- η is the learning rate. L is the loss.
- Chain rule in backpropagation
- ∂L/∂w = ∂L/∂a × ∂a/∂z × ∂z/∂w
- Gradients are passed backward from output layer to earlier layers.
Quick revision
- Supervised learning uses labelled data; unsupervised learning finds structure without labels; reinforcement learning learns from rewards.
- Overfitting means low training error but high error on new data; underfitting means the model is too simple to capture the pattern.
- Higher model complexity usually lowers bias and raises variance.
- Use separate training, validation and test data; cross-validation rotates the validation fold.
- Scale features before using penalized regression, KNN or PCA, because these depend on the size of variables.
- Ridge shrinks coefficients but does not usually set them to exactly zero; LASSO can set some to exactly zero, so it selects features.
- Elastic Net mixes the ridge and LASSO penalties.
- Logistic regression outputs a probability between 0 and 1 using p = 1 ÷ (1 + e^(−z)).
- Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN); accuracy = (TP + TN) ÷ total.
- Single decision trees overfit easily; ensembles such as bagging and random forests reduce variance by combining many trees.
- K-means needs the number of clusters set in advance; PCA builds uncorrelated components ordered by variance explained.
- Neural networks learn through layers, weights and activation functions, and need regularization or early stopping to limit overfitting.
Common mistakes
- Calling clustering a supervised method. Fix: In clustering the groups are discovered, not given. If no labels were supplied in training, it is unsupervised.
- Treating default prediction as unsupervised because it is about risk. Fix: If past loans are tagged default or not default, it is supervised classification.
- Choosing the model with the lowest training error. Fix: Select on validation or cross-validated error. Training error says nothing about generalization.
- Mixing up bias and variance for overfitting. Fix: Overfitting is high variance and low bias. Underfitting is high bias and low variance.
- Scaling the whole dataset before splitting into training and test sets Fix: Compute μ and σ (or min and max) on training data only, then apply them to the test data.
- Believing standardization makes data normally distributed Fix: Standardization only shifts and rescales. The shape of the distribution, including skewness, is unchanged.
- Saying ridge performs variable selection. Fix: Ridge shrinks toward zero but does not set coefficients exactly to zero. Only LASSO and Elastic Net do.
- Thinking a larger λ always improves the model. Fix: Too large a λ underfits: high bias. The best λ minimizes validation error, found by cross-validation.
- Mixing up precision and recall. Fix: Precision divides by predicted positives (TP + FP). Recall divides by actual positives (TP + FN). Ask: 'of those flagged' versus 'of all real cases'.
- Trusting accuracy on imbalanced data. Fix: Check recall and precision for the rare class. Use F1 or AUC when classes are imbalanced.
Exam tips
- Decide on labels first. Most questions on this topic are answered by that single check.
- Watch for key words: 'labeled', 'target' mean supervised; 'segment', 'group', 'compress' mean unsupervised; 'agent', 'reward', 'policy' mean reinforcement.
- Expect finance links: credit scoring and fraud as supervised, customer segmentation as unsupervised, execution and hedging as reinforcement.
- In contrast questions, link ML to prediction and flexibility, and traditional statistics to inference and interpretability, without saying one always wins.
- Questions usually give a table of training and validation errors. Decide on validation error first, then diagnose.
- Learn the one-line mapping: overfit = high variance, underfit = high bias. Many MCQs test only this.
- When an option says 'evaluate on the test set repeatedly to tune', it is almost always the wrong choice.
- For k-fold, do the arithmetic: fold size n ÷ k, training size n × (k − 1) ÷ k, then average the errors.