Skip to content

IAI Actuarial Core Principles · Risk Modelling and Survival Analysis

Elementary principles of machine learning: formula sheet

Full chapter guide

Key formulas

Supervised learning set-up
y = f(x) + ε
x are the features, y the target, f the unknown relationship to be learned, ε random error. Supervised learning estimates f.
Choosing the type of learning
Labelled target? Yes → supervised. No → unsupervised. Reward-driven actions? → reinforcement.
A decision rule, not a formula. Use it to classify any scenario.
Supervised sub-types
Numeric target → regression. Categorical target → classification.
A binary outcome such as lapse or no lapse is classification, even if the model outputs a probability.
Linear regression model
y = β0 + β1x1 + β2x2 + ... + βpxp + ε
Used when the target is continuous. Coefficients are usually fitted by least squares, minimising Σ(yi − ŷi)².
Euclidean distance (k-NN)
d(a, b) = √[Σ (aj − bj)²]
Sum over all features j. Scale features first, or large-valued features dominate.
k-NN prediction
Classification: majority class among the k nearest. Regression: ŷ = (1/k) Σ yi over the k nearest
Small k gives a flexible, noisy fit. Large k gives a smoother fit. Use an odd k for two classes to avoid ties.
Gini impurity
G = 1 − Σ pk²
pk is the proportion of class k in the node. G = 0 means a pure node.
Entropy
H = − Σ pk log2(pk)
Take 0 × log 0 as 0. Information gain is the parent's impurity minus the weighted average impurity of the child nodes.
Naive Bayes
P(C | x1,...,xp) ∝ P(C) × Π P(xj | C)
Assumes features are independent given the class. Compare the product for each class and choose the largest. Divide by their sum to get probabilities.
Regression tree split criterion
Choose the split that minimises Σ (yi − ȳleft)² + Σ (yi − ȳright)²
The leaf prediction is the mean of the target in that leaf.
Euclidean distance
d(x, y) = √( Σ (xᵢ − yᵢ)² )
Sum over all variables i. Standardise variables first if scales differ.
Cluster centre (centroid)
centroid = (1 ÷ n) Σ of the points in the cluster
Take the mean of each variable separately. k-means recomputes this after every assignment step.
Within-cluster sum of squares
WCSS = Σ over clusters Σ over points in cluster ‖x − centroid‖²
k-means tries to minimise this. It always falls as k rises, so do not choose k by minimum WCSS alone.
Standardisation
z = (x − mean) ÷ standard deviation
Use before clustering or PCA when variables are in different units.
Principal component
PC₁ = φ₁₁X₁ + φ₁₂X₂ + … + φ₁ₚXₚ, with Σ φ₁ⱼ² = 1
The loadings φ are the entries of the eigenvector. They are scaled to unit length.
Proportion of variance explained
PVE of component m = λₘ ÷ Σ λⱼ
λ are eigenvalues of the covariance or correlation matrix. For standardised data, Σ λⱼ = p, the number of variables.
Expected test error decomposition
Expected squared prediction error = Bias² + Variance + Irreducible error
For squared-error loss. Irreducible error is the noise variance and cannot be removed by any model.
Ridge regression
Minimise Σ (yᵢ − ŷᵢ)² + λ Σ βⱼ²
The sum is over the slope coefficients. The intercept is normally not penalised. λ ≥ 0. When λ = 0 you get ordinary least squares.
LASSO regression
Minimise Σ (yᵢ − ŷᵢ)² + λ Σ |βⱼ|
Can set coefficients exactly to zero, so it performs variable selection.
k-fold cross-validation error
CV = (1 ÷ k) Σ Eᵢ, for i = 1 to k
Eᵢ is the error measured on fold i when the model is trained on the other k − 1 folds.
Effect of λ
λ ↑ ⇒ bias ↑, variance ↓
Choose λ that minimises cross-validated error.
Accuracy
Accuracy = (TP + TN) ÷ (TP + TN + FP + FN)
Share of all cases classified correctly. Misleading when classes are very unbalanced.
Precision
Precision = TP ÷ (TP + FP)
Of those predicted positive, the share that really are positive.
Recall (sensitivity, true positive rate)
Recall = TP ÷ (TP + FN)
Of all actual positives, the share found.
Specificity
Specificity = TN ÷ (TN + FP)
Of all actual negatives, the share correctly identified.
False positive rate
FPR = FP ÷ (FP + TN) = 1 − Specificity
This is the x-axis of the ROC curve. The y-axis is recall.
F1 score
F1 = 2 × Precision × Recall ÷ (Precision + Recall)
Harmonic mean of precision and recall.
Mean squared error
MSE = (1/n) Σ (yᵢ − ŷᵢ)²
RMSE = √MSE, which is in the same units as y.
Mean absolute error
MAE = (1/n) Σ |yᵢ − ŷᵢ|
Less sensitive to large errors than MSE.
Standardisation
z = (x − mean) ÷ standard deviation
Min-max scaling uses (x − min) ÷ (max − min), which gives values from 0 to 1.

Quick revision

  • Supervised learning uses labelled data with a target; unsupervised learning has no target.
  • Regression predicts a numerical value; classification predicts a category.
  • Clustering groups similar observations; dimension reduction cuts the number of variables while keeping the main information.
  • Data is split into training data to fit the model and held-back data to test it on unseen cases.
  • A validation set is used to choose between models or tune settings; a test set gives a final unbiased check.
  • Overfitting means the model fits noise in the training data and performs poorly on new data.
  • Signs of overfitting: very good training results but clearly worse validation results.
  • Simpler models, more data and regularisation are common ways to reduce overfitting.
  • In classification, a confusion matrix counts correct and incorrect predictions by class.
  • Accuracy alone can mislead when one class is rare.
  • Always state the assumptions and the purpose of the model before choosing a measure.
  • Data quality, such as missing values and scaling, affects every method.

Common mistakes

  • Saying unsupervised learning has no data to learn from, or has no goal. Fix: Say it has no labelled target. It still has features and aims to find structure such as groups.
  • Calling a lapse or default prediction regression because it gives a probability. Fix: Look at the target. If it is a category such as lapse or not, it is classification.
  • Calling a problem regression because it uses the word 'probability'. Fix: Look at the target, not the output. If the target is a category, it is classification, even when the model outputs probabilities.
  • Using k-NN without scaling the features. Fix: Rescale first, for example to a common range or to standard scores. Otherwise a feature such as income in rupees swamps one such as age in years.
  • Not standardising variables with different units before clustering or PCA. Fix: Always say whether you standardise and why. If scales differ, standardise, or use the correlation matrix for PCA.
  • Stopping k-means after one assignment step. Fix: After updating centroids, reassign every point. Stop only when no point changes cluster.
  • Choosing the model with the lowest training error. Fix: Compare models on validation or cross-validated error, never on training error alone.
  • Using the test set repeatedly to tune the model. Fix: Tune on validation data or by cross-validation. Use the test set once at the end.
  • Swapping precision and recall Fix: Look at the denominator. Precision divides by predicted positives (row or column of predictions). Recall divides by actual positives.
  • Quoting accuracy as proof of a good model on unbalanced data Fix: Compare with the naive model that always predicts the majority class. Report recall and precision for the rare class.

Exam tips

  • Start every classification answer by stating whether a labelled target exists. Examiners look for this first.
  • Give a short insurance example for each type. Claim size, lapse, segmentation and pricing policy are safe choices.
  • When asked to compare ML with traditional modelling, give one point on each side: interpretability and assumptions versus flexibility and prediction.
  • Always mention overfitting and out-of-sample testing when discussing performance. It is a standard mark.
  • In multiple-choice questions, read the target variable carefully before choosing regression or classification.
  • Start every answer by naming the target and stating regression or classification. This earns easy marks.
  • Show the formula before the numbers. Markers give method marks even if arithmetic slips.
  • State assumptions explicitly: conditional independence for naive Bayes, scaling for k-NN, stopping rule for trees.