Skip to content

FRM Part I · FRM Exam Part I

Machine-Learning Methods: formula sheet

Full chapter guide

Key formulas

Supervised learning
Inputs (features) + known labels → learn a mapping to predict the label
Numeric label = regression. Categorical label = classification.
Unsupervised learning
Inputs (features) only, no labels → find clusters or lower-dimensional structure
Tasks: clustering and dimension reduction.
Reinforcement learning
Agent takes action in a state → receives reward → updates policy to maximize cumulative reward
Sequential decisions; feedback is a reward, not a correct label.
Prediction versus inference
ML: emphasis on out-of-sample prediction. Traditional statistics: emphasis on estimating and testing parameters
A guide to the usual emphasis, not a strict rule. Many methods serve both.
Expected prediction error decomposition
Expected error = Bias² + Variance + Irreducible error
Irreducible error is noise no model can remove. Only bias and variance change with model choice.
Cross-validation error
CV error = (1 ÷ k) × Σ (error on fold i), i = 1 to k
Each fold is the validation set once. This is a simple average when folds are equal in size.
Observations per fold
Fold size = n ÷ k
Each round trains on n × (k − 1) ÷ k observations.
Complexity pattern
Complexity ↑ → Bias ↓, Variance ↑, training error ↓
Validation error is U-shaped: it falls, then rises once overfitting starts.
Data roles
Training = fit; Validation = tune/select; Test = final unbiased estimate
Use the test set once. Using it to tune makes it part of training.
Standardization (z-score)
z = (x − μ) ÷ σ
μ and σ come from the training set. Result has mean 0 and standard deviation 1. It does not make the data normal.
Min-max normalization
x' = (x − min) ÷ (max − min)
Maps training data to 0 to 1. Very sensitive to outliers because min and max are extreme values.
Mean imputation
x_missing = (Σ x_observed) ÷ n_observed
Keeps the mean unchanged but shrinks variance. Use the median for skewed data.
Interquartile range outlier rule
Outlier if x < Q1 − 1.5 × IQR or x > Q3 + 1.5 × IQR, where IQR = Q3 − Q1
A common rule of thumb, not a law. Flag for review rather than delete automatically.
OLS objective
minimize Σ(yᵢ − ŷᵢ)²
Baseline with no penalty. Equals ridge or LASSO when λ = 0.
Ridge objective
minimize Σ(yᵢ − ŷᵢ)² + λ Σ βⱼ²
L2 penalty. Shrinks coefficients, never exactly to zero (for finite λ). Sum runs over slopes, not the intercept.
LASSO objective
minimize Σ(yᵢ − ŷᵢ)² + λ Σ |βⱼ|
L1 penalty. Can set coefficients exactly to zero, so it selects variables.
Elastic Net objective
minimize Σ(yᵢ − ŷᵢ)² + λ₁ Σ |βⱼ| + λ₂ Σ βⱼ²
Mix of L1 and L2 penalties. Both λ values are chosen by cross-validation.
Effect of λ
λ = 0 → OLS; λ → ∞ → all slopes → 0
Larger λ means more bias, less variance, a simpler model.
Logistic function
P(y = 1) = 1 ÷ (1 + e^(−z)), where z = b0 + b1x1 + ... + bkxk
Output always lies between 0 and 1.
Log-odds (logit)
ln[p ÷ (1 − p)] = b0 + b1x1 + ... + bkxk
Coefficients are linear in log-odds, not in probability. Odds = p ÷ (1 − p).
Accuracy
(TP + TN) ÷ (TP + TN + FP + FN)
Misleading when one class is rare.
Precision
TP ÷ (TP + FP)
Denominator is everything predicted positive.
Recall (sensitivity, TPR)
TP ÷ (TP + FN)
Denominator is everything actually positive.
Specificity
TN ÷ (TN + FP)
True negative rate.
False positive rate
FP ÷ (FP + TN) = 1 − specificity
The x-axis of the ROC curve.
F1 score
2 × Precision × Recall ÷ (Precision + Recall) = 2TP ÷ (2TP + FP + FN)
Harmonic mean; stays low if either precision or recall is low.
AUC interpretation
AUC = 0.5: random; AUC = 1: perfect
Higher AUC means better ranking of positives above negatives across thresholds.
Gini impurity
G = 1 − Σ pᵢ²
pᵢ is the share of class i in the node. For two classes, the maximum is 0.5 and a pure node gives 0.
Entropy
H = − Σ pᵢ × ln(pᵢ)
Zero for a pure node. Higher means more mixed. Log base (2 or e) only rescales it.
Regression tree leaf prediction
ŷ = average of y in the leaf
Splits are chosen to minimise the sum of squared errors across the two child nodes.
Bagging prediction
ŷ = (1 ÷ B) × Σ ŷᵦ
B bootstrap trees. For classification use a majority vote.
Euclidean distance
d = √[Σ (xᵢ − zᵢ)²]
Used by KNN. Scale features first or large-valued ones dominate.
Weighted impurity after a split
G_split = (n_L ÷ n) × G_L + (n_R ÷ n) × G_R
Pick the split with the lowest value, which is the largest impurity reduction.
Euclidean distance
d(x, y) = √[Σ (xᵢ − yᵢ)²]
Standard distance for k-means. Scale features first.
Centroid
centroid = average of each feature across the points in the cluster
Recomputed after every assignment step in k-means.
Within-cluster sum of squares (inertia)
WCSS = Σ over clusters Σ over points ‖x − centroid‖²
K-means minimizes this. It falls as k rises, so use the elbow, not the minimum.
Variance explained by a component
share of component j = λⱼ ÷ Σ λ
λ are eigenvalues of the covariance (or correlation) matrix.
Cumulative variance explained
(λ₁ + … + λₘ) ÷ Σ λ
Use to decide how many components to keep.
Principal component
PCⱼ = w₁ⱼX₁ + w₂ⱼX₂ + … + wₙⱼXₙ
Weights (loadings) form a unit-length eigenvector. Components are uncorrelated.
Node output
a = f(z), where z = w₁x₁ + w₂x₂ + … + wₙxₙ + b
w are weights, b is the bias, f is the activation function.
Sigmoid activation
σ(z) = 1 ÷ (1 + e^(−z))
Output lies between 0 and 1. Often used for probabilities. Can cause vanishing gradients for large |z|.
ReLU activation
ReLU(z) = max(0, z)
Simple and widely used in hidden layers. Output is zero for negative z.
Tanh activation
tanh(z) = (e^z − e^(−z)) ÷ (e^z + e^(−z))
Output lies between −1 and 1 and is centred on zero.
Mean squared error loss
MSE = (1 ÷ n) × Σ (yᵢ − ŷᵢ)²
Typical loss for numeric targets.
Gradient descent update
w_new = w_old − η × ∂L/∂w
η is the learning rate. L is the loss.
Chain rule in backpropagation
∂L/∂w = ∂L/∂a × ∂a/∂z × ∂z/∂w
Gradients are passed backward from output layer to earlier layers.

Quick revision

  • Supervised learning uses labelled data; unsupervised learning finds structure without labels; reinforcement learning learns from rewards.
  • Overfitting means low training error but high error on new data; underfitting means the model is too simple to capture the pattern.
  • Higher model complexity usually lowers bias and raises variance.
  • Use separate training, validation and test data; cross-validation rotates the validation fold.
  • Scale features before using penalized regression, KNN or PCA, because these depend on the size of variables.
  • Ridge shrinks coefficients but does not usually set them to exactly zero; LASSO can set some to exactly zero, so it selects features.
  • Elastic Net mixes the ridge and LASSO penalties.
  • Logistic regression outputs a probability between 0 and 1 using p = 1 ÷ (1 + e^(−z)).
  • Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN); accuracy = (TP + TN) ÷ total.
  • Single decision trees overfit easily; ensembles such as bagging and random forests reduce variance by combining many trees.
  • K-means needs the number of clusters set in advance; PCA builds uncorrelated components ordered by variance explained.
  • Neural networks learn through layers, weights and activation functions, and need regularization or early stopping to limit overfitting.

Common mistakes

  • Calling clustering a supervised method. Fix: In clustering the groups are discovered, not given. If no labels were supplied in training, it is unsupervised.
  • Treating default prediction as unsupervised because it is about risk. Fix: If past loans are tagged default or not default, it is supervised classification.
  • Choosing the model with the lowest training error. Fix: Select on validation or cross-validated error. Training error says nothing about generalization.
  • Mixing up bias and variance for overfitting. Fix: Overfitting is high variance and low bias. Underfitting is high bias and low variance.
  • Scaling the whole dataset before splitting into training and test sets Fix: Compute μ and σ (or min and max) on training data only, then apply them to the test data.
  • Believing standardization makes data normally distributed Fix: Standardization only shifts and rescales. The shape of the distribution, including skewness, is unchanged.
  • Saying ridge performs variable selection. Fix: Ridge shrinks toward zero but does not set coefficients exactly to zero. Only LASSO and Elastic Net do.
  • Thinking a larger λ always improves the model. Fix: Too large a λ underfits: high bias. The best λ minimizes validation error, found by cross-validation.
  • Mixing up precision and recall. Fix: Precision divides by predicted positives (TP + FP). Recall divides by actual positives (TP + FN). Ask: 'of those flagged' versus 'of all real cases'.
  • Trusting accuracy on imbalanced data. Fix: Check recall and precision for the rare class. Use F1 or AUC when classes are imbalanced.

Exam tips

  • Decide on labels first. Most questions on this topic are answered by that single check.
  • Watch for key words: 'labeled', 'target' mean supervised; 'segment', 'group', 'compress' mean unsupervised; 'agent', 'reward', 'policy' mean reinforcement.
  • Expect finance links: credit scoring and fraud as supervised, customer segmentation as unsupervised, execution and hedging as reinforcement.
  • In contrast questions, link ML to prediction and flexibility, and traditional statistics to inference and interpretability, without saying one always wins.
  • Questions usually give a table of training and validation errors. Decide on validation error first, then diagnose.
  • Learn the one-line mapping: overfit = high variance, underfit = high bias. Many MCQs test only this.
  • When an option says 'evaluate on the test set repeatedly to tune', it is almost always the wrong choice.
  • For k-fold, do the arithmetic: fold size n ÷ k, training size n × (k − 1) ÷ k, then average the errors.