FRM Part I · FRM Exam Part I
Machine Learning and Prediction: formula sheet
Key formulas
- Supervised learning
- Inputs (features) + known target → learn mapping → predict target
- Numeric target = regression. Categorical target = classification.
- Unsupervised learning
- Inputs only, no target → find structure (clusters, components)
- Main tools: clustering and PCA. No right answer is supplied.
- Reinforcement learning
- Agent → action → environment → reward → update policy
- Goal is to maximize cumulative reward, not to match labels.
- Data split
- Training set (fit) | Validation set (tune) | Test set (final check)
- Judge a model on data it did not train on.
- Error decomposition
- Expected out-of-sample squared error = Bias² + Variance + Irreducible error
- Irreducible error is noise you cannot remove by any model. Bias and variance are the parts you can trade off.
- Complexity effect
- More complexity → Bias ↓, Variance ↑ (and the reverse)
- Training error always tends to fall with complexity. Validation error is U-shaped.
- k-fold cross-validation error
- CV error = (1 ÷ k) × Σ (error on fold i), i = 1 to k
- Each observation is used for validation exactly once and for training k − 1 times.
- Training size in k-fold
- Training observations per round = N × (k − 1) ÷ k; validation = N ÷ k
- Leave-one-out is the case k = N.
- OLS loss
- Σ(yᵢ − ŷᵢ)²
- Sum of squared residuals. Regularized models add a penalty to this.
- Ridge objective
- Σ(yᵢ − ŷᵢ)² + λ Σβⱼ²
- L2 penalty. Shrinks coefficients but rarely makes them exactly zero. Intercept not penalized.
- LASSO objective
- Σ(yᵢ − ŷᵢ)² + λ Σ|βⱼ|
- L1 penalty. Can set coefficients exactly to zero, so it selects variables.
- Elastic net objective
- Σ(yᵢ − ŷᵢ)² + λ₁ Σ|βⱼ| + λ₂ Σβⱼ²
- Mix of L1 and L2. Some texts use one λ and a mixing weight instead.
- Effect of λ
- λ = 0 → OLS; λ ↑ → bias ↑, variance ↓
- Choose λ by cross-validation on held-out data.
- Proportion of variance explained
- PVE(k) = λk ÷ (λ1 + λ2 + ... + λp)
- λk is the eigenvalue of component k, and p is the number of original features. The denominator is total variance.
- Cumulative variance explained
- Cumulative PVE(m) = (λ1 + ... + λm) ÷ (λ1 + ... + λp)
- Use it to choose m, the number of components kept, to reach a target such as 90%.
- Principal component score
- PCk = w1k·x1 + w2k·x2 + ... + wpk·xp
- The weights w are the loadings. Features are usually standardised first. Loadings of a component, taken as a vector, have unit length.
- Total variance with standardised features
- Σλ = p
- This holds when you use the correlation matrix. An eigenvalue above 1 then explains more than an average single feature (the Kaiser rule of thumb).
- Orthogonality
- Correlation(PCi, PCj) = 0 for i ≠ j
- Components are uncorrelated by construction.
- Euclidean distance
- d(x, y) = √[Σ (xᵢ − yᵢ)²]
- Sum over all features i. Standardise features first.
- Manhattan distance
- d(x, y) = Σ |xᵢ − yᵢ|
- Less sensitive to a single large difference than Euclidean.
- Centroid
- centroid = (1 ÷ n) × Σ of the points in the cluster
- The mean of each feature across the cluster's members.
- Within-cluster sum of squares
- WCSS = Σ over clusters Σ over points ‖x − centroid‖²
- K-means minimises this. It always falls as K increases.
- Silhouette coefficient
- s = (b − a) ÷ max(a, b)
- a = average distance to own cluster; b = average distance to nearest other cluster. Range −1 to 1.
- Gini impurity of a node
- G = 1 − Σ pᵢ²
- pᵢ is the share of class i in the node. For two classes, G = 2p(1 − p). Zero means a pure node.
- Entropy of a node
- H = − Σ pᵢ × ln(pᵢ)
- Another impurity measure. Zero for a pure node. Log base only scales the value.
- Weighted impurity after a split
- Weighted G = (n_left ÷ n) × G_left + (n_right ÷ n) × G_right
- Choose the split with the lowest weighted impurity, which is the largest impurity reduction.
- Regression tree split criterion
- Minimise SSE = Σ(y − ȳ_left)² + Σ(y − ȳ_right)²
- Leaf prediction is the mean of y in that leaf.
- Cost-complexity pruning
- Cost = Error(tree) + α × (number of leaves)
- Larger α gives a smaller tree. Choose α by validation or cross-validation.
- Bagging prediction
- Prediction = (1 ÷ B) × Σ f_b(x), b = 1 to B
- For classification use majority vote. Averaging reduces variance, not bias.
- Variance of an average of B equally correlated models
- Var = ρσ² + (1 − ρ)σ² ÷ B
- ρ is the correlation between models, σ² each model's variance. Lower ρ (random forest) helps more than just raising B.
- Node pre-activation
- z = Σ (wᵢ × xᵢ) + b
- Weighted sum of inputs plus a bias. The activation function is applied to z.
- Node output
- a = f(z)
- f is the activation function. If f is linear, stacked layers collapse to one linear model.
- Sigmoid
- σ(z) = 1 ÷ (1 + e^(−z))
- Output lies between 0 and 1. Used for probabilities, for example default probability. Gradient is small for large positive or negative z.
- Sigmoid derivative
- σ'(z) = σ(z) × (1 − σ(z))
- Maximum value is 0.25 at z = 0. Explains the vanishing gradient problem.
- Tanh
- tanh(z) = (e^z − e^(−z)) ÷ (e^z + e^(−z))
- Output between −1 and 1 and centred on zero.
- ReLU
- ReLU(z) = max(0, z)
- Output is zero for negative z and equal to z otherwise. Cheap to compute and eases vanishing gradients.
- Mean squared error loss
- L = (1 ÷ n) × Σ (yᵢ − ŷᵢ)²
- Common loss for regression. For a single observation with a ½ factor, L = ½(y − ŷ)².
- Chain rule for a weight
- ∂L/∂w = ∂L/∂a × ∂a/∂z × ∂z/∂w
- This is backpropagation. For a node, ∂z/∂w equals the input x.
- Gradient descent update
- w(new) = w(old) − η × ∂L/∂w
- η is the learning rate. Too large can overshoot; too small makes training slow.
- Accuracy
- Accuracy = (TP + TN) ÷ (TP + TN + FP + FN)
- Share of all predictions that are correct. Misleading with imbalanced classes.
- Precision
- Precision = TP ÷ (TP + FP)
- Denominator is everything predicted positive.
- Recall (sensitivity, TPR)
- Recall = TP ÷ (TP + FN)
- Denominator is everything actually positive.
- False positive rate
- FPR = FP ÷ (FP + TN)
- Equals 1 − specificity. This is the x-axis of the ROC curve.
- Specificity
- Specificity = TN ÷ (TN + FP)
- Share of actual negatives correctly identified.
- F1 score
- F1 = 2 × Precision × Recall ÷ (Precision + Recall)
- Harmonic mean of precision and recall.
- Logistic function
- p = 1 ÷ (1 + e^−z), where z = b0 + b1x1 + ... + bkxk
- Log-odds ln(p ÷ (1 − p)) = z is linear in the inputs.
- RMSE
- RMSE = √[ Σ(yᵢ − ŷᵢ)² ÷ n ]
- Same units as y. Larger errors are penalised more.
Quick revision
- Supervised learning uses labelled outcomes; unsupervised finds structure without labels; reinforcement learns from rewards.
- Overfitting means low training error but high error on new data; the model has learned noise.
- Higher model complexity usually lowers bias and raises variance.
- Use separate training, validation and test sets; the test set is for final evaluation only.
- Ridge shrinks coefficients toward zero but does not usually set them exactly to zero.
- LASSO can set coefficients exactly to zero, so it performs variable selection.
- Elastic Net blends the Ridge and LASSO penalties.
- PCA components are uncorrelated; the first explains the most variance, and the shares of variance across all components sum to 100%.
- K-Means needs you to choose the number of clusters K in advance; hierarchical clustering does not.
- Random forests average many trees built on random samples and features to reduce variance.
- Precision = TP ÷ (TP + FP); recall = TP ÷ (TP + FN); accuracy = (TP + TN) ÷ total.
- Accuracy can mislead when classes are imbalanced.
Common mistakes
- Calling clustering a supervised method Fix: Clusters are discovered by the algorithm. Classification uses labels that were given in advance.
- Treating a default prediction model as regression because it outputs a probability Fix: Look at the target. Default or not is a category, so it is classification.
- Choosing the model with the lowest training error. Fix: Select on validation or cross-validation error. Training error is not a measure of predictive power.
- Using the test set to tune hyperparameters. Fix: Tune on validation data. Touch the test set once at the end, otherwise its estimate is optimistic.
- Saying ridge performs variable selection. Fix: Ridge makes coefficients small but almost never exactly zero. Only LASSO and elastic net give exact zeros.
- Choosing λ to minimize training error. Fix: Pick λ using cross-validation or a validation set, which measures out-of-sample error.
- Dividing an eigenvalue by the number of features instead of the sum of eigenvalues. Fix: Always divide by the sum of all eigenvalues unless the question says the features are standardised.
- Saying components are correlated with each other. Fix: Remember PCA produces uncorrelated, orthogonal components by construction.
- Treating clustering as supervised learning. Fix: Clustering finds groups from the data alone. If the question mentions known labels, it is classification.
- Choosing the K that gives the lowest WCSS. Fix: WCSS reaches zero when K equals the number of points. Use the elbow or the silhouette score.
Exam tips
- Most questions are scenario matching. Find the label and the goal first.
- Expect contrasts: inference versus prediction, labeled versus unlabeled, supervised versus reinforcement.
- Watch for options that wrongly say unsupervised learning uses labeled data.
- Remember the finance examples: credit scoring and fraud (supervised), clustering and PCA (unsupervised), trade execution and dynamic hedging (reinforcement).
- Expect conceptual questions that give a pair of training and validation errors and ask you to diagnose overfitting or underfitting.
- Memorize the role of each sample: training fits, validation tunes, test confirms once.
- Be ready for simple k-fold arithmetic: fold sizes and the average of fold errors.
- Remember that regularization, simpler models and more data reduce variance, and a note on time series splits is a likely trap.