Skip to content

Risk Modelling and Survival Analysis · Elementary principles of machine learning

Model Performance Measures and Practical Considerations in Machine Learning

Updated 11 October 2026 · Fact-checked

Model performance measures tell you how well a model predicts unseen data. For classification, build a confusion matrix and compute accuracy, precision, recall and specificity. For regression, use error metrics such as MSE. ROC curves and AUC show performance across thresholds. Good data preparation and ethical checks matter just as much.

Understand Model Performance Measures and Practical Considerations

A model is only useful if it predicts well on new data, not just the data it was built on. So we measure performance on a test set that the model has not seen during training.

For regression (predicting a number), we look at the errors, which are actual minus predicted. Common metrics are the mean squared error (MSE), the root mean squared error (RMSE) and the mean absolute error (MAE). Squaring errors punishes large misses more heavily than MAE does.

For classification (predicting a category), we use a confusion matrix. It counts four outcomes for a yes/no problem: true positives (TP), false positives (FP), true negatives (TN) and false negatives (FN). From these you get accuracy, precision, recall (also called sensitivity) and specificity. Accuracy can mislead when one class is rare, as with fraud or large claims. A model that always says "no fraud" can be 99% accurate and still useless.

Most classifiers output a probability. You pick a threshold to turn it into a yes/no. Changing the threshold trades off true positives against false positives. The ROC curve plots the true positive rate against the false positive rate for every threshold. The AUC is the area under that curve. An AUC of 0.5 is no better than random guessing and 1 is perfect separation.

Practical points matter too. Data cleaning deals with missing values, errors and outliers. Feature scaling puts variables on similar ranges, which matters for methods based on distances or gradients. Data should be split so that test data does not leak into training. Finally, models used in insurance must be checked for bias, fairness, transparency and data protection. A model can discriminate indirectly through proxy variables even if a protected characteristic is removed.

Key rules to remember

Accuracy
Accuracy = (TP + TN) ÷ (TP + TN + FP + FN)
Share of all cases classified correctly. Misleading when classes are very unbalanced.
Precision
Precision = TP ÷ (TP + FP)
Of those predicted positive, the share that really are positive.
Recall (sensitivity, true positive rate)
Recall = TP ÷ (TP + FN)
Of all actual positives, the share found.
Specificity
Specificity = TN ÷ (TN + FP)
Of all actual negatives, the share correctly identified.
False positive rate
FPR = FP ÷ (FP + TN) = 1 − Specificity
This is the x-axis of the ROC curve. The y-axis is recall.
F1 score
F1 = 2 × Precision × Recall ÷ (Precision + Recall)
Harmonic mean of precision and recall.
Mean squared error
MSE = (1/n) Σ (yᵢ − ŷᵢ)²
RMSE = √MSE, which is in the same units as y.
Mean absolute error
MAE = (1/n) Σ |yᵢ − ŷᵢ|
Less sensitive to large errors than MSE.
Standardisation
z = (x − mean) ÷ standard deviation
Min-max scaling uses (x − min) ÷ (max − min), which gives values from 0 to 1.

How to solve Model Performance Measures and Practical Considerations questions

Use this order for any question on measuring or applying a model.

  1. 1Identify the problem type: regression (numerical target) or classification (categorical target).
  2. 2For classification, define which outcome is "positive" and write out the 2 by 2 confusion matrix with TP, FP, FN and TN.
  3. 3Check that the four counts add up to the total number of cases.
  4. 4Pick the formula the question asks for. Write it down, then substitute the counts.
  5. 5Interpret the result in context. Say what a false positive or false negative costs in the actuarial setting.
  6. 6For ROC and AUC questions, compute TPR and FPR at each threshold, plot or compare, and remember that AUC 0.5 means random.
  7. 7For practical questions, name the issue (missing data, scaling, leakage, bias), explain the harm and give a remedy.

Quickest way: Confusion matrix in four lines

When to use it: Use this for MCQs asking for precision, recall, accuracy or specificity from given counts.

  1. Write TP, FP, FN, TN from the question. Check the sum.
  2. Remember the denominators: precision uses predicted positives (TP + FP), recall uses actual positives (TP + FN).
  3. Specificity uses actual negatives (TN + FP).
  4. Compute only the one ratio asked and check it lies between 0 and 1.

Common mistakes in Model Performance Measures and Practical Considerations

  • Swapping precision and recall

    Both have TP on top and the names sound alike.

    Fix: Look at the denominator. Precision divides by predicted positives (row or column of predictions). Recall divides by actual positives.

  • Quoting accuracy as proof of a good model on unbalanced data

    Accuracy looks simple and high numbers feel reassuring.

    Fix: Compare with the naive model that always predicts the majority class. Report recall and precision for the rare class.

  • Evaluating the model on the training data

    It is easy and gives impressive results.

    Fix: Always use a separate test set or cross-validation. Training error understates the true error because of overfitting.

  • Mixing up the axes of the ROC curve

    Students remember the shape but not the labels.

    Fix: The x-axis is FPR (1 − specificity). The y-axis is TPR (recall). The diagonal is random guessing.

  • Scaling using the whole dataset before splitting

    It seems efficient to scale once.

    Fix: Compute the mean and standard deviation on the training set only, then apply them to the test set. Otherwise information leaks.

  • Saying that removing a protected variable removes bias

    It sounds logical.

    Fix: Other variables can act as proxies for it. Test outcomes across groups and review the data and the use of the model.

Worked examples

Example 1

A model predicts whether a motor claim is fraudulent. On 200 test claims the results are: TP = 16, FP = 4, FN = 14, TN = 166. Calculate accuracy, precision, recall and specificity.

Show the solution
  1. Check the total: 16 + 4 + 14 + 166 = 200.
  2. Accuracy = (16 + 166) ÷ 200 = 182 ÷ 200 = 0.91.
  3. Precision = 16 ÷ (16 + 4) = 16 ÷ 20 = 0.80.
  4. Recall = 16 ÷ (16 + 14) = 16 ÷ 30 = 0.5333.
  5. Specificity = 166 ÷ (166 + 4) = 166 ÷ 170 = 0.9765.
  6. Comment: accuracy is high, but the model finds only about 53% of fraudulent claims.

Answer: Accuracy 0.91, precision 0.80, recall 0.533, specificity 0.976. The model misses nearly half of the fraud, so accuracy alone is misleading.

Example 2

A classifier gives these results at two thresholds on 100 actual positives and 400 actual negatives. Threshold A: TP = 80, FP = 60. Threshold B: TP = 60, FP = 20. Find the ROC point for each threshold and say which threshold gives the lower false positive rate.

Show the solution
  1. Recall (TPR) = TP ÷ 100, FPR = FP ÷ 400.
  2. Threshold A: TPR = 80 ÷ 100 = 0.80. FPR = 60 ÷ 400 = 0.15.
  3. Threshold B: TPR = 60 ÷ 100 = 0.60. FPR = 20 ÷ 400 = 0.05.
  4. Threshold B has the lower FPR (0.05 against 0.15) but also the lower TPR.
  5. Both points lie above the diagonal line TPR = FPR, so both beat random guessing.

Answer: Threshold A is the point (FPR, TPR) = (0.15, 0.80). Threshold B is (0.05, 0.60). B has the lower false positive rate, at the cost of lower recall.

Exam tips

  • Write the confusion matrix out before calculating. It stops denominator errors and earns method marks.
  • Always interpret the number in the insurance context, for example the cost of missing a fraud case versus investigating a genuine claim.
  • For discussion questions, give points on data quality, overfitting, scaling, bias and explainability, with one sentence of explanation for each.
  • In the computer-based paper, state how you split the data and which metric you used, as well as giving the code output.
  • Know the ROC axes and the meaning of AUC 0.5 and 1. These are frequent MCQ targets.

Practice questions from Elementary principles of machine learning

Model Performance Measures and Practical Considerations in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Model Performance Measures and Practical Considerations: frequently asked questions

How do I remember precision and recall?

Precision asks how many of the cases you flagged were right, so it divides by flagged cases (TP + FP). Recall asks how many of the real positives you caught, so it divides by real positives (TP + FN).

What does AUC tell me?

AUC is the area under the ROC curve. It measures how well the model ranks positives above negatives across all thresholds. A value of 0.5 means no better than chance and 1 means perfect ranking.

Why do we scale features?

Methods that use distances or gradient steps, such as k-means or penalised regression, are dominated by variables with large numeric ranges. Scaling gives each variable a fair influence. Tree-based methods are generally not affected.

How can machine learning be biased in insurance?

Bias can come from unrepresentative training data, past decisions that were unfair, or proxy variables that stand in for protected characteristics. Actuaries should test outcomes across groups, document the model and keep it explainable.