FRM Exam Part II · Credit Scoring and Rating
Validation and Performance Testing of Rating Models
Updated 11 October 2026 · Fact-checked
Validation checks whether a rating model ranks borrowers correctly (discriminatory power) and predicts default rates accurately (calibration). Plot the ROC curve, compute AUC, then Gini = 2 × AUC − 1. The CAP accuracy ratio equals the Gini coefficient. Back-testing compares predicted PDs with realised default rates by grade.
Understand Validation and Performance Testing of Rating Models
A rating model has two jobs. First, it must rank borrowers: bad ones should get worse ratings than good ones. This is discriminatory power. Second, its probability of default (PD) for each grade must match what actually happens. This is calibration. Validation tests both.
To test ranking, you need a past sample where you know who defaulted. Pick a cut-off score. Borrowers scoring below it are flagged as bad. The hit rate (true positive rate) is the share of defaulters correctly flagged. The false alarm rate (false positive rate) is the share of non-defaulters wrongly flagged.
The ROC curve plots hit rate (vertical) against false alarm rate (horizontal) as you move the cut-off across all values. A random model gives a 45-degree diagonal, with AUC (area under the curve) of 0.5. A perfect model gives AUC of 1.0. The Gini coefficient rescales AUC so random = 0 and perfect = 1.
The CAP curve (cumulative accuracy profile) plots the share of all defaulters captured against the share of all borrowers, ordered from worst to best score. The accuracy ratio (AR) is the area between the model CAP and the random line, divided by the same area for a perfect model. AR and Gini are the same number. Both are linked to AUC.
Calibration is tested by back-testing: compare the predicted PD of each grade with the realised default frequency. A binomial test is common. It assumes defaults are independent, which understates the chance of too many defaults when defaults are correlated. Always validate out-of-sample and out-of-time, because in-sample results look better than they are.
Key formulas to remember
- Hit rate (true positive rate)
- Hit rate = defaulters flagged ÷ total defaulters
- Vertical axis of the ROC curve.
- False alarm rate (false positive rate)
- False alarm rate = non-defaulters flagged ÷ total non-defaulters
- Horizontal axis of the ROC curve.
- Gini from AUC
- Gini = AR = 2 × AUC − 1
- Inverse: AUC = (Gini + 1) ÷ 2.
- Benchmarks
- Random model: AUC = 0.5, Gini = 0. Perfect model: AUC = 1, Gini = 1
- AUC below 0.5 means the ranking is reversed.
- Accuracy ratio
- AR = area between model CAP and random line ÷ area between perfect CAP and random line
- The perfect CAP depends on the portfolio default rate.
- Binomial back-test
- Observed defaults ~ Binomial(n, PD); z ≈ (D − n × PD) ÷ √(n × PD × (1 − PD))
- Normal approximation for large n. Assumes independent defaults.
How to solve Validation and Performance Testing of Rating Models questions
Use this order for any validation question.
- 1Decide what is being tested: ranking (discriminatory power) or PD accuracy (calibration).
- 2For ranking, identify the measure given: ROC/AUC, Gini, CAP or accuracy ratio.
- 3Convert between measures with Gini = AR = 2 × AUC − 1.
- 4Check benchmarks: 0.5 AUC or 0 Gini is no better than random; higher is better.
- 5For calibration, compare predicted PD with realised default rate per grade, using the binomial test or z-score.
- 6Check the sample: is it out-of-sample and out-of-time, and are the defaults independent?
- 7State the interpretation: good ranking does not prove good calibration, and vice versa.
Quickest way: Convert and compare
When to use it: When the question gives one measure and asks for another or asks which model is better.
- Write Gini = 2 × AUC − 1 and AR = Gini.
- Convert every model to the same measure.
- Higher value means better ranking; ignore the other options that confuse AUC and Gini.
- If the question is about PD levels, switch to back-testing and compute the z-score.
Common mistakes in Validation and Performance Testing of Rating Models
Treating Gini and accuracy ratio as different numbers
They have different names and are drawn on different curves (ROC vs CAP).
Fix: Remember they are equal. Both equal 2 × AUC − 1.
Saying AUC of 0.5 is a Gini of 0.5
Mixing up the two scales.
Fix: AUC 0.5 gives Gini 0. Always compute 2 × 0.5 − 1.
Assuming high discriminatory power means PDs are accurate
Ranking and calibration are blended together.
Fix: A model can rank perfectly and still understate every PD. Test both.
Putting the wrong rate on the ROC axes
Swapping hit rate and false alarm rate.
Fix: Hit rate is vertical; false alarm rate is horizontal.
Trusting in-sample results
The model was fitted to that same data, so performance looks strong.
Fix: Insist on out-of-sample and out-of-time validation.
Ignoring default correlation in the binomial test
The test looks simple and exact.
Fix: Say the independence assumption makes the test too strict on excess defaults when correlation exists, so it can wrongly reject the model.
Worked examples
Example 1
A bank's rating model has an AUC of 0.82. Calculate the Gini coefficient and the accuracy ratio.
Show the solution
- Gini = 2 × AUC − 1.
- Gini = 2 × 0.82 − 1 = 1.64 − 1 = 0.64.
- The accuracy ratio equals the Gini coefficient, so AR = 0.64.
Answer: Gini = 0.64 and accuracy ratio = 0.64.
Example 2
A grade has a predicted PD of 2%. In the year, 400 borrowers were in the grade and 14 defaulted. Using the normal approximation, compute the z-score. Is the observed count significantly above the prediction at the 95% one-sided level (critical value 1.645)?
Show the solution
- Expected defaults = 400 × 0.02 = 8.
- Standard deviation = √(400 × 0.02 × 0.98) = √7.84 = 2.8.
- z = (14 − 8) ÷ 2.8 = 2.14.
- 2.14 is greater than 1.645.
Answer: z ≈ 2.14, which exceeds 1.645. The PD looks underestimated at the 95% one-sided level, assuming independent defaults.
Exam tips
- Memorise Gini = AR = 2 × AUC − 1 and the two benchmark points (random and perfect).
- When options list AUC and Gini values, convert first, then compare.
- Read carefully whether the question asks about ranking or about PD accuracy.
- Look for the trap in sample design: out-of-sample, out-of-time and independence assumptions are favourite distractors.
Practice questions from Credit Scoring and Rating
- A bank observes that the one-year default rate for a speculative-grade rating class was 4% in year one and 6% in year two, among surviving i…
- A risk analyst computes the Gini (accuracy ratio) for a retail scorecard and finds an AUC of 0.80 from the ROC curve. What is the correspond…
- A one-year transition matrix has three states: A, B and Default (D). From A: A 90%, B 8%, D 2%. From B: A 10%, B 80%, D 10%. D is absorbing.…
- A bank uses an external-ratings-based approach in which risk weights step up sharply as ratings fall. In a recession, many of its corporate …
- A portfolio manager notes that agency transition matrices estimated across the business cycle show that the probability of a BBB issuer bein…
Validation and Performance Testing of Rating Models in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Validation and Performance Testing of Rating Models: frequently asked questions
Is the accuracy ratio the same as the Gini coefficient?
Yes. The accuracy ratio from the CAP curve equals the Gini coefficient from the ROC curve. Both equal 2 × AUC − 1.
What AUC is a random credit model?
A random model has an AUC of 0.5, shown as the diagonal on the ROC curve. Its Gini is 0.
How do you back-test a credit rating model?
Compare the predicted PD of each rating grade with the realised default rate over the period. Use a test such as the binomial test. Also use out-of-sample data and allow for default correlation.
What is the difference between discrimination and calibration?
Discrimination is how well the model separates defaulters from non-defaulters. Calibration is how close predicted PDs are to actual default rates. You need both.