FRM Exam Part I · Machine Learning and Prediction
Confusion Matrix, Precision, Recall and ROC Curve
Updated 11 October 2026 · Fact-checked
Classification metrics judge how well a model sorts cases into classes, such as default or no default. Build the confusion matrix (TP, FP, TN, FN), then compute accuracy, precision, recall and the false positive rate. The ROC curve plots recall against the false positive rate; AUC summarises it. RMSE measures prediction error.
Understand Model Evaluation and Classification Metrics
A classification model predicts a category, for example whether a borrower will default. Logistic regression is the standard starting point. It models the probability of the positive class as p = 1 ÷ (1 + e^−(b0 + b1x1 + ...)). The output is always between 0 and 1. You then pick a threshold, often 0.5, and label a borrower as default if p is above it.
After predictions are made, you compare them with actual outcomes in a confusion matrix. A true positive (TP) is a predicted default that did default. A false positive (FP) is a predicted default that did not. A true negative (TN) is a predicted non-default that did not default. A false negative (FN) is a predicted non-default that did default. In credit work, an FN is usually the costly error.
Each metric answers a different question. Accuracy asks how often the model is right overall. Precision asks: of the cases I flagged, how many were really positive? Recall (sensitivity, true positive rate) asks: of the real positives, how many did I catch? Accuracy can mislead when classes are imbalanced. If only 2% of loans default, a model that never predicts default is 98% accurate and useless.
The threshold creates a trade-off. Lowering it flags more cases, so recall rises but precision usually falls. The ROC curve plots recall (true positive rate) against the false positive rate (FPR = FP ÷ (FP + TN)) across all thresholds. AUC is the area under that curve. An AUC of 0.5 is no better than random guessing, and 1.0 is perfect ranking. AUC does not depend on one threshold.
For models that predict a number, such as a loss amount, the common error measure is RMSE, the square root of the average squared error. It is in the same units as the target and penalises large errors more than small ones. Lower is better, and it should be judged on data not used to fit the model.
Key formulas to remember
- Accuracy
- Accuracy = (TP + TN) ÷ (TP + TN + FP + FN)
- Share of all predictions that are correct. Misleading with imbalanced classes.
- Precision
- Precision = TP ÷ (TP + FP)
- Denominator is everything predicted positive.
- Recall (sensitivity, TPR)
- Recall = TP ÷ (TP + FN)
- Denominator is everything actually positive.
- False positive rate
- FPR = FP ÷ (FP + TN)
- Equals 1 − specificity. This is the x-axis of the ROC curve.
- Specificity
- Specificity = TN ÷ (TN + FP)
- Share of actual negatives correctly identified.
- F1 score
- F1 = 2 × Precision × Recall ÷ (Precision + Recall)
- Harmonic mean of precision and recall.
- Logistic function
- p = 1 ÷ (1 + e^−z), where z = b0 + b1x1 + ... + bkxk
- Log-odds ln(p ÷ (1 − p)) = z is linear in the inputs.
- RMSE
- RMSE = √[ Σ(yᵢ − ŷᵢ)² ÷ n ]
- Same units as y. Larger errors are penalised more.
How to solve Model Evaluation and Classification Metrics questions
Use this routine for any confusion-matrix or evaluation question.
- 1Identify the positive class (usually default or the event of interest).
- 2Write down TP, FP, TN and FN. If only some are given, use the total and the row or column sums to find the rest.
- 3Check the question: does it ask about predicted positives (precision) or actual positives (recall)?
- 4Write the formula, then substitute the counts.
- 5Compute the answer and sanity-check that it lies between 0 and 1.
- 6For ROC and AUC questions, reason about direction: lower threshold means higher TPR and higher FPR; AUC of 0.5 is random.
- 7For RMSE, square each error, average, then take the square root.
Quickest way: Denominator check
When to use it: Any precision, recall or FPR question with a table of counts.
- Precision: divide by the predicted-positive total (TP + FP).
- Recall: divide by the actual-positive total (TP + FN).
- FPR: divide by the actual-negative total (FP + TN).
- The numerator for precision and recall is always TP.
- Eliminate options that exceed 1 or ignore the correct denominator.
Common mistakes in Model Evaluation and Classification Metrics
Swapping precision and recall.
Both use TP in the numerator and the names sound similar.
Fix: Precision: denominator is predicted positives. Recall: denominator is actual positives.
Trusting accuracy on imbalanced data.
A high percentage looks good.
Fix: Compare accuracy with the base rate and check recall and precision on the rare class.
Putting precision on the ROC axes.
Confusing ROC with the precision-recall curve.
Fix: ROC plots TPR (recall) on the y-axis against FPR on the x-axis.
Reading AUC of 0.5 as 50% accuracy or a good model.
AUC looks like a percentage.
Fix: 0.5 means no discrimination, equal to random ranking. Higher is better.
Assuming a lower threshold improves both precision and recall.
Ignoring the trade-off.
Fix: A lower threshold usually raises recall and FPR, and often lowers precision.
Evaluating RMSE on training data only.
It is the easiest number to get.
Fix: Use a validation or test set to judge out-of-sample performance and spot overfitting.
Worked examples
Example 1
A default model is tested on 1,000 loans. It gives TP = 40, FP = 20, FN = 10, TN = 930. Compute accuracy, precision and recall.
Show the solution
- Check the total: 40 + 20 + 10 + 930 = 1,000.
- Accuracy = (40 + 930) ÷ 1,000 = 970 ÷ 1,000 = 0.97.
- Precision = 40 ÷ (40 + 20) = 40 ÷ 60 = 0.6667.
- Recall = 40 ÷ (40 + 10) = 40 ÷ 50 = 0.80.
Answer: Accuracy = 97%, precision = 66.7%, recall = 80%.
Example 2
Using the same results, compute the false positive rate. Then state what happens to recall and FPR if the bank lowers the default threshold.
Show the solution
- Actual negatives = FP + TN = 20 + 930 = 950.
- FPR = 20 ÷ 950 = 0.0211.
- A lower threshold labels more loans as default.
- More actual defaulters are caught, so recall rises.
- More good loans are also flagged, so FPR rises.
Answer: FPR ≈ 2.11%. Lowering the threshold raises both recall and FPR, moving up and to the right along the ROC curve.
Exam tips
- Always identify the positive class first, then fill in the four cells before computing anything.
- Expect conceptual questions on the precision-recall trade-off and why accuracy fails with rare defaults.
- For ROC questions, remember the diagonal is random (AUC 0.5) and a curve closer to the top-left is better.
- In credit scenarios, ask which error costs more. Missing a defaulter (FN) usually favours prioritising recall.
- Do the arithmetic by hand; most questions use simple counts.
Practice questions from Machine Learning and Prediction
- An analyst uses a regularized model with a tuning parameter that controls complexity. As the penalty is increased substantially from a very …
- A K-means model is fitted to stocks for K = 1 to 5, producing total within-cluster sum of squares of 500, 220, 100, 90 and 85 respectively. …
- A risk manager clusters customers on two variables: annual transaction volume (in USD, ranging 0 to 2,000,000) and number of late payments (…
- A risk analyst groups 500 corporate borrowers into segments using only their financial ratios, with no default labels available. Which descr…
- An analyst at a risk consultancy wants to group 400 corporate borrowers into segments based on leverage, interest coverage and asset volatil…
Model Evaluation and Classification Metrics: frequently asked questions
What is the difference between precision and recall?
Precision is the share of predicted positives that are truly positive: TP ÷ (TP + FP). Recall is the share of actual positives the model finds: TP ÷ (TP + FN). Precision penalises false alarms, while recall penalises misses.
How do I interpret the ROC curve and AUC?
The ROC curve plots the true positive rate against the false positive rate as the threshold changes. A curve nearer the top-left corner is better. AUC summarises it: 0.5 is random ranking and 1.0 is perfect separation.
Why use logistic regression for default prediction?
The dependent variable is binary, and logistic regression keeps predicted probabilities between 0 and 1. Its coefficients are also interpretable through log-odds. You then choose a threshold to turn probabilities into default or non-default labels.
Is accuracy a good metric for default models?
Often not. Defaults are rare, so a model that predicts no default for everyone can score high accuracy. Use recall, precision and AUC alongside it.