IAI Actuarial Core Principles · Risk Modelling and Survival Analysis
Elementary principles of machine learning: formula sheet
Key formulas
- Supervised learning set-up
- y = f(x) + ε
- x are the features, y the target, f the unknown relationship to be learned, ε random error. Supervised learning estimates f.
- Choosing the type of learning
- Labelled target? Yes → supervised. No → unsupervised. Reward-driven actions? → reinforcement.
- A decision rule, not a formula. Use it to classify any scenario.
- Supervised sub-types
- Numeric target → regression. Categorical target → classification.
- A binary outcome such as lapse or no lapse is classification, even if the model outputs a probability.
- Linear regression model
- y = β0 + β1x1 + β2x2 + ... + βpxp + ε
- Used when the target is continuous. Coefficients are usually fitted by least squares, minimising Σ(yi − ŷi)².
- Euclidean distance (k-NN)
- d(a, b) = √[Σ (aj − bj)²]
- Sum over all features j. Scale features first, or large-valued features dominate.
- k-NN prediction
- Classification: majority class among the k nearest. Regression: ŷ = (1/k) Σ yi over the k nearest
- Small k gives a flexible, noisy fit. Large k gives a smoother fit. Use an odd k for two classes to avoid ties.
- Gini impurity
- G = 1 − Σ pk²
- pk is the proportion of class k in the node. G = 0 means a pure node.
- Entropy
- H = − Σ pk log2(pk)
- Take 0 × log 0 as 0. Information gain is the parent's impurity minus the weighted average impurity of the child nodes.
- Naive Bayes
- P(C | x1,...,xp) ∝ P(C) × Π P(xj | C)
- Assumes features are independent given the class. Compare the product for each class and choose the largest. Divide by their sum to get probabilities.
- Regression tree split criterion
- Choose the split that minimises Σ (yi − ȳleft)² + Σ (yi − ȳright)²
- The leaf prediction is the mean of the target in that leaf.
- Euclidean distance
- d(x, y) = √( Σ (xᵢ − yᵢ)² )
- Sum over all variables i. Standardise variables first if scales differ.
- Cluster centre (centroid)
- centroid = (1 ÷ n) Σ of the points in the cluster
- Take the mean of each variable separately. k-means recomputes this after every assignment step.
- Within-cluster sum of squares
- WCSS = Σ over clusters Σ over points in cluster ‖x − centroid‖²
- k-means tries to minimise this. It always falls as k rises, so do not choose k by minimum WCSS alone.
- Standardisation
- z = (x − mean) ÷ standard deviation
- Use before clustering or PCA when variables are in different units.
- Principal component
- PC₁ = φ₁₁X₁ + φ₁₂X₂ + … + φ₁ₚXₚ, with Σ φ₁ⱼ² = 1
- The loadings φ are the entries of the eigenvector. They are scaled to unit length.
- Proportion of variance explained
- PVE of component m = λₘ ÷ Σ λⱼ
- λ are eigenvalues of the covariance or correlation matrix. For standardised data, Σ λⱼ = p, the number of variables.
- Expected test error decomposition
- Expected squared prediction error = Bias² + Variance + Irreducible error
- For squared-error loss. Irreducible error is the noise variance and cannot be removed by any model.
- Ridge regression
- Minimise Σ (yᵢ − ŷᵢ)² + λ Σ βⱼ²
- The sum is over the slope coefficients. The intercept is normally not penalised. λ ≥ 0. When λ = 0 you get ordinary least squares.
- LASSO regression
- Minimise Σ (yᵢ − ŷᵢ)² + λ Σ |βⱼ|
- Can set coefficients exactly to zero, so it performs variable selection.
- k-fold cross-validation error
- CV = (1 ÷ k) Σ Eᵢ, for i = 1 to k
- Eᵢ is the error measured on fold i when the model is trained on the other k − 1 folds.
- Effect of λ
- λ ↑ ⇒ bias ↑, variance ↓
- Choose λ that minimises cross-validated error.
- Accuracy
- Accuracy = (TP + TN) ÷ (TP + TN + FP + FN)
- Share of all cases classified correctly. Misleading when classes are very unbalanced.
- Precision
- Precision = TP ÷ (TP + FP)
- Of those predicted positive, the share that really are positive.
- Recall (sensitivity, true positive rate)
- Recall = TP ÷ (TP + FN)
- Of all actual positives, the share found.
- Specificity
- Specificity = TN ÷ (TN + FP)
- Of all actual negatives, the share correctly identified.
- False positive rate
- FPR = FP ÷ (FP + TN) = 1 − Specificity
- This is the x-axis of the ROC curve. The y-axis is recall.
- F1 score
- F1 = 2 × Precision × Recall ÷ (Precision + Recall)
- Harmonic mean of precision and recall.
- Mean squared error
- MSE = (1/n) Σ (yᵢ − ŷᵢ)²
- RMSE = √MSE, which is in the same units as y.
- Mean absolute error
- MAE = (1/n) Σ |yᵢ − ŷᵢ|
- Less sensitive to large errors than MSE.
- Standardisation
- z = (x − mean) ÷ standard deviation
- Min-max scaling uses (x − min) ÷ (max − min), which gives values from 0 to 1.
Quick revision
- Supervised learning uses labelled data with a target; unsupervised learning has no target.
- Regression predicts a numerical value; classification predicts a category.
- Clustering groups similar observations; dimension reduction cuts the number of variables while keeping the main information.
- Data is split into training data to fit the model and held-back data to test it on unseen cases.
- A validation set is used to choose between models or tune settings; a test set gives a final unbiased check.
- Overfitting means the model fits noise in the training data and performs poorly on new data.
- Signs of overfitting: very good training results but clearly worse validation results.
- Simpler models, more data and regularisation are common ways to reduce overfitting.
- In classification, a confusion matrix counts correct and incorrect predictions by class.
- Accuracy alone can mislead when one class is rare.
- Always state the assumptions and the purpose of the model before choosing a measure.
- Data quality, such as missing values and scaling, affects every method.
Common mistakes
- Saying unsupervised learning has no data to learn from, or has no goal. Fix: Say it has no labelled target. It still has features and aims to find structure such as groups.
- Calling a lapse or default prediction regression because it gives a probability. Fix: Look at the target. If it is a category such as lapse or not, it is classification.
- Calling a problem regression because it uses the word 'probability'. Fix: Look at the target, not the output. If the target is a category, it is classification, even when the model outputs probabilities.
- Using k-NN without scaling the features. Fix: Rescale first, for example to a common range or to standard scores. Otherwise a feature such as income in rupees swamps one such as age in years.
- Not standardising variables with different units before clustering or PCA. Fix: Always say whether you standardise and why. If scales differ, standardise, or use the correlation matrix for PCA.
- Stopping k-means after one assignment step. Fix: After updating centroids, reassign every point. Stop only when no point changes cluster.
- Choosing the model with the lowest training error. Fix: Compare models on validation or cross-validated error, never on training error alone.
- Using the test set repeatedly to tune the model. Fix: Tune on validation data or by cross-validation. Use the test set once at the end.
- Swapping precision and recall Fix: Look at the denominator. Precision divides by predicted positives (row or column of predictions). Recall divides by actual positives.
- Quoting accuracy as proof of a good model on unbalanced data Fix: Compare with the naive model that always predicts the majority class. Report recall and precision for the rare class.
Exam tips
- Start every classification answer by stating whether a labelled target exists. Examiners look for this first.
- Give a short insurance example for each type. Claim size, lapse, segmentation and pricing policy are safe choices.
- When asked to compare ML with traditional modelling, give one point on each side: interpretability and assumptions versus flexibility and prediction.
- Always mention overfitting and out-of-sample testing when discussing performance. It is a standard mark.
- In multiple-choice questions, read the target variable carefully before choosing regression or classification.
- Start every answer by naming the target and stating regression or classification. This earns easy marks.
- Show the formula before the numbers. Markers give method marks even if arithmetic slips.
- State assumptions explicitly: conditional independence for naive Bayes, scaling for k-NN, stopping rule for trees.