Skip to content

CFA Level II Exam · Big Data Projects

Model Training and Performance Evaluation: Precision, Recall, F1 and ROC

Updated 7 October 2026 · Fact-checked

Model evaluation tests how well a trained model predicts. For classification, build a confusion matrix and compute accuracy, precision, recall and F1. Use the ROC curve and AUC to judge discrimination across thresholds. For numeric forecasts, use RMSE. Class imbalance makes accuracy misleading, so favour F1 or AUC.

Understand Model Training and Performance Evaluation

After you build a model on big data, you must check whether it works on data it has not seen. You split the data into a training set to fit the model and a test set (or validation set) to judge it. A model that scores well on training data but poorly on new data is overfit.

For a classification model, the output is a predicted class, such as default or no default. You count the results in a confusion matrix. True positives (TP) are positives correctly predicted. False positives (FP) are negatives wrongly called positive. False negatives (FN) are positives the model missed. True negatives (TN) are negatives correctly predicted.

Class imbalance means one class is far rarer than the other, such as 2% of loans defaulting. A model that always predicts no default gets 98% accuracy yet finds no defaults. So accuracy alone is a poor guide. Precision asks how many predicted positives were right. Recall (sensitivity) asks how many actual positives were found. The F1 score is the harmonic mean of the two and is useful when classes are imbalanced.

The curve measure is the ROC curve. It plots the true positive rate (recall) on the vertical axis against the false positive rate on the horizontal axis as the cutoff threshold changes. The area under the curve (AUC) summarises it. An AUC of 0.5 is no better than random guessing. An AUC of 1.0 is perfect. A higher AUC means a better model. A more convex curve toward the top-left corner means better performance.

For models that predict a continuous number, use RMSE, the square root of the average squared error. Smaller is better. It is in the same units as the target, and large errors are penalised more because errors are squared.

The trade-off: lowering the cutoff raises recall but lowers precision. Pick the measure that matches the cost of error. If missing a default is costly, favour recall. If false alarms are costly, favour precision.

Key formulas to remember

Precision (P)
P = TP ÷ (TP + FP)
Share of predicted positives that are truly positive. Penalises false positives (Type I errors).
Recall (R)
R = TP ÷ (TP + FN)
Also the true positive rate. Penalises false negatives (Type II errors).
Accuracy
Accuracy = (TP + TN) ÷ (TP + FP + TN + FN)
Share of all predictions that are correct. Misleading when classes are imbalanced.
F1 score
F1 = (2 × P × R) ÷ (P + R)
Harmonic mean of precision and recall. Ranges from 0 to 1.
False positive rate (FPR)
FPR = FP ÷ (FP + TN)
Horizontal axis of the ROC curve.
RMSE
RMSE = √[ Σ(predicted − actual)² ÷ n ]
For numeric predictions. Lower is better. Same units as the target variable.
AUC interpretation
AUC = 0.5 means random; AUC = 1 means perfect
Higher AUC means better separation of classes; the curve is plotted of TPR against FPR.

How to solve Model Training and Performance Evaluation questions

Use this method for any question on classification or forecast evaluation in a vignette.

  1. 1Identify the model type: classification (classes) or prediction of a number (use RMSE).
  2. 2Find the data in the vignette: either confusion matrix counts or the ratios already given. Label TP, FP, FN and TN carefully by rereading what counts as the positive class.
  3. 3Check class balance. If one class is rare, do not rely on accuracy.
  4. 4Compute the requested measure from the formulas. If asked for F1, compute precision and recall first.
  5. 5For ROC or AUC questions, compare AUC with 0.5 and with other models. Higher AUC wins. Note that a point on the curve is one threshold.
  6. 6Decide which error matters. Costly misses point to recall. Costly false alarms point to precision.
  7. 7Check that your answer has the right scale (a ratio between 0 and 1) and answer the exact question asked.

Quickest way: Fill the 2×2 grid, then compute

When to use it: When the vignette gives raw counts and asks for precision, recall, accuracy or F1.

  1. Draw a quick 2×2 grid with TP, FP, FN, TN. Fill from the exhibit.
  2. Total = TP + FP + FN + TN. Check it matches the sample size.
  3. Precision = TP over all predicted positives (TP + FP). Recall = TP over all actual positives (TP + FN). Check in the exhibit whether predicted classes are in rows or columns.
  4. For F1, use the shortcut F1 = 2TP ÷ (2TP + FP + FN).
  5. Eliminate answer options that are above 1 or inconsistent with which of precision and recall is larger.

Common mistakes in Model Training and Performance Evaluation

  • Mixing up precision and recall.

    Both use TP in the numerator and the names sound alike.

    Fix: Precision divides by predicted positives (TP + FP). Recall divides by actual positives (TP + FN). Ask what the denominator counts.

  • Recommending accuracy for an imbalanced data set.

    Accuracy looks simple and high numbers seem good.

    Fix: When one class is rare, use F1, recall, precision or AUC. State that a trivial model can have high accuracy.

  • Putting FP and FN in the wrong cells.

    Students read the matrix rows and columns in the wrong order.

    Fix: Check the labels on the exhibit. A false positive is predicted positive but actually negative.

  • Averaging precision and recall with a simple mean for F1.

    The arithmetic mean is the familiar average.

    Fix: F1 is the harmonic mean: 2PR ÷ (P + R). It is pulled toward the lower of the two.

  • Reading an AUC of 0.5 as a good model, or a lower RMSE as worse.

    Confusion about the scale and direction of each measure.

    Fix: AUC of 0.5 is random guessing, and higher is better. For RMSE, lower is better.

  • Forgetting to take the square root in RMSE.

    Students stop after computing the mean squared error.

    Fix: After summing squared errors and dividing by n, take the square root.

Worked examples

Example 1

A lender tests a default-prediction model on 200 loans. Defaults are the positive class. The model gives TP = 16, FP = 4, FN = 8, TN = 172. (1) Compute accuracy. (2) Compute precision and recall. (3) Compute F1.

Show the solution
  1. Check the total: 16 + 4 + 8 + 172 = 200. Correct.
  2. Accuracy = (16 + 172) ÷ 200 = 188 ÷ 200 = 0.94.
  3. Precision = 16 ÷ (16 + 4) = 16 ÷ 20 = 0.80.
  4. Recall = 16 ÷ (16 + 8) = 16 ÷ 24 = 0.6667.
  5. F1 = 2 × 0.80 × 0.6667 ÷ (0.80 + 0.6667) = 1.0667 ÷ 1.4667 = 0.7273.
  6. Check with the shortcut: 2 × 16 ÷ (32 + 4 + 8) = 32 ÷ 44 = 0.7273.

Answer: Accuracy is 94%, precision is 80%, recall is 66.67% and F1 is about 0.727. Only 24 of 200 loans are defaults, so accuracy flatters the model; recall shows it misses one third of defaults.

Example 2

An analyst compares two models that predict quarterly sales growth (in %) for a firm. The actual values for four quarters are 4, 6, 2, 8. Model A predicts 5, 5, 3, 7. Model B predicts 4, 8, 2, 6. (1) Compute RMSE for each model. (2) Which is better? (3) A different classifier has an AUC of 0.50. What does that mean?

Show the solution
  1. Model A errors: 1, −1, 1, −1. Squared: 1, 1, 1, 1. Sum = 4. Mean = 1. RMSE = √1 = 1.00.
  2. Model B errors: 0, 2, 0, −2. Squared: 0, 4, 0, 4. Sum = 8. Mean = 2. RMSE = √2 = 1.414.
  3. Lower RMSE is better, so Model A is better even though Model B has two exact predictions.
  4. An AUC of 0.50 means the ROC curve lies on the diagonal, so the classifier separates the classes no better than random guessing.

Answer: RMSE is 1.00 for Model A and about 1.41 for Model B, so Model A is better. An AUC of 0.50 means the classifier has no predictive skill beyond chance.

Exam tips

  • Read which class is the positive class before filling the matrix. Wrong labelling flips precision and recall.
  • If the vignette stresses a rare event such as fraud or default, expect the answer to favour recall, F1 or AUC over accuracy.
  • Expect a conceptual option about the trade-off: a lower cutoff usually raises recall and lowers precision.
  • Check your arithmetic with the shortcut F1 = 2TP ÷ (2TP + FP + FN). It avoids rounding errors.
  • For RMSE, square the errors, average them, then take the square root. Lower RMSE means a better fit on the test set.

Model Training and Performance Evaluation: frequently asked questions

What is the difference between precision and recall in CFA Level II?

Precision is the share of predicted positives that are actually positive: TP ÷ (TP + FP). Recall is the share of actual positives the model found: TP ÷ (TP + FN). Precision penalises false positives, and recall penalises false negatives.

When should I use F1 instead of accuracy?

Use F1 when classes are imbalanced or when both false positives and false negatives matter. Accuracy can be high just by predicting the majority class. F1 combines precision and recall into one figure.

How do I read an ROC curve and AUC?

The ROC curve plots the true positive rate against the false positive rate at different thresholds. The closer it bends toward the top-left corner, the better. AUC of 0.5 is random and 1.0 is perfect.

What does RMSE tell me about a model?

RMSE measures the typical size of prediction errors for a numeric target, in the same units as the target. Lower values mean a better fit. Because errors are squared, large misses weigh more.