Skip to content

FRM Part I · FRM Exam Part I

Machine Learning and Prediction: formula sheet

Full chapter guide

Key formulas

Supervised learning
Inputs (features) + known target → learn mapping → predict target
Numeric target = regression. Categorical target = classification.
Unsupervised learning
Inputs only, no target → find structure (clusters, components)
Main tools: clustering and PCA. No right answer is supplied.
Reinforcement learning
Agent → action → environment → reward → update policy
Goal is to maximize cumulative reward, not to match labels.
Data split
Training set (fit) | Validation set (tune) | Test set (final check)
Judge a model on data it did not train on.
Error decomposition
Expected out-of-sample squared error = Bias² + Variance + Irreducible error
Irreducible error is noise you cannot remove by any model. Bias and variance are the parts you can trade off.
Complexity effect
More complexity → Bias ↓, Variance ↑ (and the reverse)
Training error always tends to fall with complexity. Validation error is U-shaped.
k-fold cross-validation error
CV error = (1 ÷ k) × Σ (error on fold i), i = 1 to k
Each observation is used for validation exactly once and for training k − 1 times.
Training size in k-fold
Training observations per round = N × (k − 1) ÷ k; validation = N ÷ k
Leave-one-out is the case k = N.
OLS loss
Σ(yᵢ − ŷᵢ)²
Sum of squared residuals. Regularized models add a penalty to this.
Ridge objective
Σ(yᵢ − ŷᵢ)² + λ Σβⱼ²
L2 penalty. Shrinks coefficients but rarely makes them exactly zero. Intercept not penalized.
LASSO objective
Σ(yᵢ − ŷᵢ)² + λ Σ|βⱼ|
L1 penalty. Can set coefficients exactly to zero, so it selects variables.
Elastic net objective
Σ(yᵢ − ŷᵢ)² + λ₁ Σ|βⱼ| + λ₂ Σβⱼ²
Mix of L1 and L2. Some texts use one λ and a mixing weight instead.
Effect of λ
λ = 0 → OLS; λ ↑ → bias ↑, variance ↓
Choose λ by cross-validation on held-out data.
Proportion of variance explained
PVE(k) = λk ÷ (λ1 + λ2 + ... + λp)
λk is the eigenvalue of component k, and p is the number of original features. The denominator is total variance.
Cumulative variance explained
Cumulative PVE(m) = (λ1 + ... + λm) ÷ (λ1 + ... + λp)
Use it to choose m, the number of components kept, to reach a target such as 90%.
Principal component score
PCk = w1k·x1 + w2k·x2 + ... + wpk·xp
The weights w are the loadings. Features are usually standardised first. Loadings of a component, taken as a vector, have unit length.
Total variance with standardised features
Σλ = p
This holds when you use the correlation matrix. An eigenvalue above 1 then explains more than an average single feature (the Kaiser rule of thumb).
Orthogonality
Correlation(PCi, PCj) = 0 for i ≠ j
Components are uncorrelated by construction.
Euclidean distance
d(x, y) = √[Σ (xᵢ − yᵢ)²]
Sum over all features i. Standardise features first.
Manhattan distance
d(x, y) = Σ |xᵢ − yᵢ|
Less sensitive to a single large difference than Euclidean.
Centroid
centroid = (1 ÷ n) × Σ of the points in the cluster
The mean of each feature across the cluster's members.
Within-cluster sum of squares
WCSS = Σ over clusters Σ over points ‖x − centroid‖²
K-means minimises this. It always falls as K increases.
Silhouette coefficient
s = (b − a) ÷ max(a, b)
a = average distance to own cluster; b = average distance to nearest other cluster. Range −1 to 1.
Gini impurity of a node
G = 1 − Σ pᵢ²
pᵢ is the share of class i in the node. For two classes, G = 2p(1 − p). Zero means a pure node.
Entropy of a node
H = − Σ pᵢ × ln(pᵢ)
Another impurity measure. Zero for a pure node. Log base only scales the value.
Weighted impurity after a split
Weighted G = (n_left ÷ n) × G_left + (n_right ÷ n) × G_right
Choose the split with the lowest weighted impurity, which is the largest impurity reduction.
Regression tree split criterion
Minimise SSE = Σ(y − ȳ_left)² + Σ(y − ȳ_right)²
Leaf prediction is the mean of y in that leaf.
Cost-complexity pruning
Cost = Error(tree) + α × (number of leaves)
Larger α gives a smaller tree. Choose α by validation or cross-validation.
Bagging prediction
Prediction = (1 ÷ B) × Σ f_b(x), b = 1 to B
For classification use majority vote. Averaging reduces variance, not bias.
Variance of an average of B equally correlated models
Var = ρσ² + (1 − ρ)σ² ÷ B
ρ is the correlation between models, σ² each model's variance. Lower ρ (random forest) helps more than just raising B.
Node pre-activation
z = Σ (wᵢ × xᵢ) + b
Weighted sum of inputs plus a bias. The activation function is applied to z.
Node output
a = f(z)
f is the activation function. If f is linear, stacked layers collapse to one linear model.
Sigmoid
σ(z) = 1 ÷ (1 + e^(−z))
Output lies between 0 and 1. Used for probabilities, for example default probability. Gradient is small for large positive or negative z.
Sigmoid derivative
σ'(z) = σ(z) × (1 − σ(z))
Maximum value is 0.25 at z = 0. Explains the vanishing gradient problem.
Tanh
tanh(z) = (e^z − e^(−z)) ÷ (e^z + e^(−z))
Output between −1 and 1 and centred on zero.
ReLU
ReLU(z) = max(0, z)
Output is zero for negative z and equal to z otherwise. Cheap to compute and eases vanishing gradients.
Mean squared error loss
L = (1 ÷ n) × Σ (yᵢ − ŷᵢ)²
Common loss for regression. For a single observation with a ½ factor, L = ½(y − ŷ)².
Chain rule for a weight
∂L/∂w = ∂L/∂a × ∂a/∂z × ∂z/∂w
This is backpropagation. For a node, ∂z/∂w equals the input x.
Gradient descent update
w(new) = w(old) − η × ∂L/∂w
η is the learning rate. Too large can overshoot; too small makes training slow.
Accuracy
Accuracy = (TP + TN) ÷ (TP + TN + FP + FN)
Share of all predictions that are correct. Misleading with imbalanced classes.
Precision
Precision = TP ÷ (TP + FP)
Denominator is everything predicted positive.
Recall (sensitivity, TPR)
Recall = TP ÷ (TP + FN)
Denominator is everything actually positive.
False positive rate
FPR = FP ÷ (FP + TN)
Equals 1 − specificity. This is the x-axis of the ROC curve.
Specificity
Specificity = TN ÷ (TN + FP)
Share of actual negatives correctly identified.
F1 score
F1 = 2 × Precision × Recall ÷ (Precision + Recall)
Harmonic mean of precision and recall.
Logistic function
p = 1 ÷ (1 + e^−z), where z = b0 + b1x1 + ... + bkxk
Log-odds ln(p ÷ (1 − p)) = z is linear in the inputs.
RMSE
RMSE = √[ Σ(yᵢ − ŷᵢ)² ÷ n ]
Same units as y. Larger errors are penalised more.

Quick revision

  • Supervised learning uses labelled outcomes; unsupervised finds structure without labels; reinforcement learns from rewards.
  • Overfitting means low training error but high error on new data; the model has learned noise.
  • Higher model complexity usually lowers bias and raises variance.
  • Use separate training, validation and test sets; the test set is for final evaluation only.
  • Ridge shrinks coefficients toward zero but does not usually set them exactly to zero.
  • LASSO can set coefficients exactly to zero, so it performs variable selection.
  • Elastic Net blends the Ridge and LASSO penalties.
  • PCA components are uncorrelated; the first explains the most variance, and the shares of variance across all components sum to 100%.
  • K-Means needs you to choose the number of clusters K in advance; hierarchical clustering does not.
  • Random forests average many trees built on random samples and features to reduce variance.
  • Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN); accuracy = (TP + TN) ÷ total.
  • Accuracy can mislead when classes are imbalanced.

Common mistakes

  • Calling clustering a supervised method Fix: Clusters are discovered by the algorithm. Classification uses labels that were given in advance.
  • Treating a default prediction model as regression because it outputs a probability Fix: Look at the target. Default or not is a category, so it is classification.
  • Choosing the model with the lowest training error. Fix: Select on validation or cross-validation error. Training error is not a measure of predictive power.
  • Using the test set to tune hyperparameters. Fix: Tune on validation data. Touch the test set once at the end, otherwise its estimate is optimistic.
  • Saying ridge performs variable selection. Fix: Ridge makes coefficients small but almost never exactly zero. Only LASSO and elastic net give exact zeros.
  • Choosing λ to minimize training error. Fix: Pick λ using cross-validation or a validation set, which measures out-of-sample error.
  • Dividing an eigenvalue by the number of features instead of the sum of eigenvalues. Fix: Always divide by the sum of all eigenvalues unless the question says the features are standardised.
  • Saying components are correlated with each other. Fix: Remember PCA produces uncorrelated, orthogonal components by construction.
  • Treating clustering as supervised learning. Fix: Clustering finds groups from the data alone. If the question mentions known labels, it is classification.
  • Choosing the K that gives the lowest WCSS. Fix: WCSS reaches zero when K equals the number of points. Use the elbow or the silhouette score.

Exam tips

  • Most questions are scenario matching. Find the label and the goal first.
  • Expect contrasts: inference versus prediction, labeled versus unlabeled, supervised versus reinforcement.
  • Watch for options that wrongly say unsupervised learning uses labeled data.
  • Remember the finance examples: credit scoring and fraud (supervised), clustering and PCA (unsupervised), trade execution and dynamic hedging (reinforcement).
  • Expect conceptual questions that give a pair of training and validation errors and ask you to diagnose overfitting or underfitting.
  • Memorize the role of each sample: training fits, validation tunes, test confirms once.
  • Be ready for simple k-fold arithmetic: fold sizes and the average of fold errors.
  • Remember that regularization, simpler models and more data reduce variance, and a note on time series splits is a likely trap.