Skip to content

FRM Exam Part I · Machine-Learning Methods

Logistic Regression and Classification Metrics for FRM Part I

Updated 11 October 2026 · Fact-checked

Logistic regression models the probability of a binary outcome, such as default or no default, using the S-shaped logistic function. You then pick a cutoff to classify, build a confusion matrix, and compute accuracy, precision, recall, F1 and the ROC curve. To solve questions, fill the matrix first, then apply the formula.

Understand Logistic Regression and Classification Metrics

Many risk questions have a yes-or-no answer: will the borrower default, is this transaction fraud, will the counterparty breach a limit. A linear regression is a poor fit because it can predict values below 0 or above 1. Logistic regression fixes this by passing a linear score through the logistic function, so the output always lies between 0 and 1 and can be read as a probability.

The model says: score z = b0 + b1x1 + b2x2 + ..., and P(y = 1) = 1 ÷ (1 + e^(−z)). Equivalently, the log-odds ln[p ÷ (1 − p)] equals z. This means each coefficient is the change in log-odds for a one-unit rise in that variable. A positive coefficient raises the probability of the event. The effect on probability itself is not constant; it is largest near p = 0.5 and small near 0 or 1. Coefficients are estimated by maximum likelihood, not OLS.

A probability is not yet a decision. You choose a threshold (often 0.5) and classify as positive when the predicted probability is at or above it. Comparing predictions with actual outcomes gives a confusion matrix with four cells: true positives (TP), false positives (FP), true negatives (TN) and false negatives (FN).

From these four cells come the metrics. Accuracy is the share of all cases classified correctly. Precision asks: of the cases I flagged, how many were truly positive? Recall (sensitivity, true positive rate) asks: of all the real positives, how many did I catch? The F1 score is the harmonic mean of precision and recall. Accuracy can mislead when classes are imbalanced, which is typical for defaults and fraud.

The ROC curve plots the true positive rate (recall) against the false positive rate as you move the threshold from high to low. The AUC is the area under it. An AUC of 0.5 means no better than random guessing (the diagonal line); 1.0 means perfect separation. A curve that bows toward the top-left corner is better. AUC does not depend on one threshold, so it compares models in general.

Key formulas to remember

Logistic function
P(y = 1) = 1 ÷ (1 + e^(−z)), where z = b0 + b1x1 + ... + bkxk
Output always lies between 0 and 1.
Log-odds (logit)
ln[p ÷ (1 − p)] = b0 + b1x1 + ... + bkxk
Coefficients are linear in log-odds, not in probability. Odds = p ÷ (1 − p).
Accuracy
(TP + TN) ÷ (TP + TN + FP + FN)
Misleading when one class is rare.
Precision
TP ÷ (TP + FP)
Denominator is everything predicted positive.
Recall (sensitivity, TPR)
TP ÷ (TP + FN)
Denominator is everything actually positive.
Specificity
TN ÷ (TN + FP)
True negative rate.
False positive rate
FP ÷ (FP + TN) = 1 − specificity
The x-axis of the ROC curve.
F1 score
2 × Precision × Recall ÷ (Precision + Recall) = 2TP ÷ (2TP + FP + FN)
Harmonic mean; stays low if either precision or recall is low.
AUC interpretation
AUC = 0.5: random; AUC = 1: perfect
Higher AUC means better ranking of positives above negatives across thresholds.

How to solve Logistic Regression and Classification Metrics questions

Use this routine for any logistic regression or classification metric question.

  1. 1Identify what counts as the positive class (for example, default = positive). Everything depends on this.
  2. 2Write down or build the confusion matrix: TP, FP, TN, FN. If given totals and rates, back out the missing cells.
  3. 3Check the sums: TP + FN is the actual positives; TP + FP is the predicted positives; all four cells make the total.
  4. 4Pick the metric asked for and write its formula with the right denominator.
  5. 5Substitute the numbers and compute. Keep fractions until the end to avoid rounding errors.
  6. 6For a logistic model question, compute z first, then p = 1 ÷ (1 + e^(−z)), then compare p with the threshold.
  7. 7For ROC and AUC questions, reason about direction: lowering the threshold raises both TPR and FPR.
  8. 8Sanity check: all rates lie between 0 and 1, and F1 lies between precision and recall.

Quickest way: Four-cell shortcut

When to use it: Use it for any confusion matrix question where you must find precision, recall, F1 or accuracy under time pressure.

  1. Sketch a 2 by 2 grid and label rows as actual, columns as predicted.
  2. Fill TP, FN, FP, TN from the question.
  3. Remember: precision uses the predicted column, recall uses the actual row.
  4. For F1, use 2TP ÷ (2TP + FP + FN); it avoids computing precision and recall separately.
  5. Eliminate options: F1 must sit between precision and recall, and closer to the lower one.

Common mistakes in Logistic Regression and Classification Metrics

  • Mixing up precision and recall.

    Both have TP on top and the names sound alike.

    Fix: Precision divides by predicted positives (TP + FP). Recall divides by actual positives (TP + FN). Ask: 'of those flagged' versus 'of all real cases'.

  • Trusting accuracy on imbalanced data.

    A model that predicts 'no default' for everyone scores high accuracy when defaults are rare.

    Fix: Check recall and precision for the rare class. Use F1 or AUC when classes are imbalanced.

  • Reading a logistic coefficient as a change in probability.

    Linear regression habits carry over.

    Fix: A coefficient is the change in log-odds per unit. The change in probability depends on the starting point.

  • Averaging precision and recall to get F1.

    The arithmetic mean feels natural.

    Fix: F1 is the harmonic mean: 2PR ÷ (P + R). It is always at most the arithmetic mean.

  • Thinking a lower threshold reduces false positives.

    Confusing the direction of the threshold.

    Fix: A lower threshold classifies more cases as positive, so recall rises and false positives also rise. Precision usually falls.

  • Treating AUC = 0.5 as a bad-but-useful model, or AUC below 0.5 as meaningless.

    Poor grasp of the random benchmark.

    Fix: AUC of 0.5 equals random guessing and has no discriminating power. A value well above 0.5 shows real ranking skill.

Worked examples

Example 1

A bank's default model is tested on 1,000 loans. It flags 80 loans as defaults, of which 60 actually default. In total, 100 loans actually default. Calculate accuracy, precision, recall and F1.

Show the solution
  1. Positive = default. TP = 60 (flagged and defaulted).
  2. FP = flagged − TP = 80 − 60 = 20.
  3. FN = actual defaults − TP = 100 − 60 = 40.
  4. TN = 1,000 − 60 − 20 − 40 = 880.
  5. Accuracy = (60 + 880) ÷ 1,000 = 0.94.
  6. Precision = 60 ÷ (60 + 20) = 0.75.
  7. Recall = 60 ÷ (60 + 40) = 0.60.
  8. F1 = 2 × 60 ÷ (2 × 60 + 20 + 40) = 120 ÷ 180 = 0.667.

Answer: Accuracy 94%, precision 75%, recall 60%, F1 about 0.667.

Example 2

A logistic regression for default gives z = −4 + 0.8x, where x is leverage. For a borrower with x = 5, find the default probability and the classification at a 0.5 threshold. Use e^0 = 1.

Show the solution
  1. Compute the score: z = −4 + 0.8 × 5 = −4 + 4 = 0.
  2. Apply the logistic function: p = 1 ÷ (1 + e^(−0)) = 1 ÷ (1 + 1) = 0.5.
  3. Compare with the threshold: p = 0.5 is at the threshold, so under the 'at or above' rule the borrower is classified as default.
  4. Interpretation: at z = 0 the odds are 1 to 1, so log-odds are 0.
  5. Check the effect of x = 6: z = 0.8, which is positive, so p is above 0.5. Higher leverage raises the log-odds by 0.8 per unit.

Answer: Default probability is 0.5 (odds 1:1); classified as default at a threshold of 0.5 under the at-or-above rule.

Exam tips

  • Always state the positive class before computing; if the question flips it, precision and recall change.
  • Questions often give totals and rates rather than the four cells. Rebuild the grid first.
  • Know qualitative statements: lowering the threshold raises recall and FPR; AUC of 0.5 is random; accuracy misleads with imbalanced classes.
  • In risk settings, ask which error is costlier. Missing a default (FN) usually favours recall; wrongly rejecting good clients (FP) favours precision.
  • Use the F1 shortcut 2TP ÷ (2TP + FP + FN) to save time.

Practice questions from Machine-Learning Methods

Logistic Regression and Classification Metrics: frequently asked questions

What is the difference between precision and recall?

Precision is the share of predicted positives that are truly positive: TP ÷ (TP + FP). Recall is the share of actual positives that you caught: TP ÷ (TP + FN). Precision penalises false alarms; recall penalises missed cases.

Why use logistic regression instead of linear regression for a binary outcome?

Linear regression can give fitted values outside 0 to 1 and assumes a constant effect on probability. Logistic regression keeps outputs between 0 and 1 and models log-odds as linear. It is estimated by maximum likelihood.

How do I read a ROC curve and AUC?

The ROC curve plots true positive rate against false positive rate as the threshold changes. A curve closer to the top-left corner is better. AUC summarises the curve: 0.5 is random guessing and 1.0 is perfect separation.

When is F1 better than accuracy?

F1 is better when classes are imbalanced, such as rare defaults or fraud. Accuracy can look high even if the model misses most positives. F1 balances precision and recall and ignores true negatives.