FRM Exam Part I · External and Internal Credit Ratings
Validation and Backtesting of Internal Rating Systems
Updated 11 October 2026 · Fact-checked
Validation checks whether a rating system works. Discriminatory power asks whether it ranks risky borrowers worse than safe ones, measured by the CAP curve, accuracy ratio or AUC. Calibration asks whether predicted PDs match realised default rates, tested by backtesting with a binomial test. Solve by naming which property the question tests.
Understand Validation and Backtesting of Rating Systems
A rating system assigns each borrower a grade and each grade a probability of default (PD). Validation is the process of checking that the system keeps doing its job. It has two core questions, and you must keep them apart.
Discriminatory power (also called rank ordering or separation) asks: do the worst grades contain more defaulters than the best grades? A system can rank perfectly and still quote the wrong PD levels. Calibration asks: are the PDs attached to each grade close to the default rates you actually observe? A system can have the right levels but rank poorly.
The main tool for discriminatory power is the Cumulative Accuracy Profile (CAP). Sort borrowers from worst to best score. Plot the cumulative share of all borrowers on the x-axis and the cumulative share of all defaulters captured on the y-axis. A perfect model captures all defaults in the worst-rated borrowers. A random model gives the 45-degree diagonal. The accuracy ratio (AR) is the area between your model's CAP and the diagonal, divided by the same area for the perfect model. AR runs from 0 (random) to 1 (perfect). It links to the ROC curve: AR = 2 × AUC − 1, where AUC is the area under the ROC curve.
Backtesting PD estimates is a calibration exercise. For one grade with n borrowers and predicted PD p, if defaults are independent, the number of defaults follows a binomial distribution with mean np and standard deviation √(np(1 − p)). If observed defaults are far above np, the PD is too low (underestimated risk). Default correlation and economic cycles break the independence assumption, so simple binomial tests can reject good models too often. Validation also covers data quality, stability of grades over time, and the use of out-of-sample and out-of-time samples to avoid overfitting.
Key formulas to remember
- Accuracy ratio
- AR = (Area between model CAP and diagonal) ÷ (Area between perfect CAP and diagonal)
- Ranges from 0 (no better than random) to 1 (perfect). Also called the Gini coefficient or Somers' D in this context.
- AR and AUC link
- AR = 2 × AUC − 1
- AUC is the area under the ROC curve. AUC = 0.5 is random, so AR = 0.
- Perfect model CAP area
- Area above diagonal = 0.5 × (1 − D), where D = overall default rate
- Used to compute AR from a model area. A lower default rate means a larger perfect-model area.
- Binomial mean and standard deviation of defaults
- E(defaults) = n × p; σ = √(n × p × (1 − p))
- Assumes independent defaults and one PD for a grade of n borrowers.
- Normal approximation z-score
- z = (Observed defaults − n × p) ÷ √(n × p × (1 − p))
- Reasonable when n × p is not small. Compare with the critical value, for example 1.645 for a one-sided 95% test.
- Observed default rate
- Default rate = Number of defaults ÷ Number of borrowers at start of period
- Compare this with the predicted PD for each grade.
How to solve Validation and Backtesting of Rating Systems questions
Use this sequence for any validation or backtesting question.
- 1Identify what is being tested: discriminatory power (ranking) or calibration (PD level). The wording tells you: 'rank', 'separate', 'CAP', 'AR', 'AUC' means discrimination; 'predicted versus realised', 'PD too low', 'binomial' means calibration.
- 2Write the relevant formula before touching the numbers.
- 3For discrimination: read the area or AR given. Convert using AR = 2 × AUC − 1, or AR = model area ÷ perfect area.
- 4For calibration: compute expected defaults n × p, then the standard deviation √(n × p × (1 − p)).
- 5Compute the z-score and compare with the critical value, or compare observed defaults with the expected range.
- 6State the conclusion in direction: PD underestimated if defaults are too high, overestimated if too low.
- 7Check whether the question hints at correlation, cycle effects or small samples, and note that these weaken a simple binomial test.
Quickest way: Rank or level? Then one calculation
When to use it: Use this for multiple-choice questions where you have under two minutes.
- Decide in five seconds: ranking (AR, AUC, CAP) or level (PD versus default rate).
- If ranking, use AR = 2 × AUC − 1 and check the answer lies between 0 and 1.
- If level, compute n × p and √(n × p × (1 − p)), then z.
- Eliminate options that confuse the two properties, such as saying a high AR proves PDs are accurate.
- Sanity check the sign: more defaults than expected means PDs are too low.
Common mistakes in Validation and Backtesting of Rating Systems
Treating a high accuracy ratio as proof that PDs are correct
Students merge the two properties because both are called 'model performance'.
Fix: AR measures ranking only. Calibration needs a separate test of PD against default rate.
Using AR = AUC
Both are summary numbers for the same curves, so they look interchangeable.
Fix: Remember AR = 2 × AUC − 1. Random gives AUC 0.5 and AR 0.
Reading the CAP diagonal as the perfect model
Both are straight reference lines on the chart.
Fix: The diagonal is the random model. The perfect model rises steeply to 100% of defaults at the default-rate share of borrowers, then goes flat.
Forgetting the (1 − p) term in the binomial standard deviation
Students recall np and drop part of the variance formula.
Fix: Write σ = √(np(1 − p)) every time. For small p the result is close to √(np), but the exam may test the exact value.
Concluding the model is wrong from one binomial rejection without considering correlation
The test looks mechanical, so students ignore its independence assumption.
Fix: Defaults cluster in downturns. Positive correlation widens the true spread of defaults, so the simple test over-rejects. Say so when asked about limitations.
Validating only on the development sample
It is the data you already have.
Fix: Use out-of-sample and out-of-time data. In-sample results overstate performance because of overfitting.
Worked examples
Example 1
A bank's rating model has an area under the ROC curve of 0.82. What is its accuracy ratio, and what does it say about the model?
Show the solution
- Formula: AR = 2 × AUC − 1.
- Substitute: AR = 2 × 0.82 − 1.
- Compute: 1.64 − 1 = 0.64.
Answer: AR = 0.64. The model ranks borrowers well, clearly better than random (AR = 0), but it is short of perfect (AR = 1). This says nothing about whether the PD levels are accurate.
Example 2
A rating grade has 400 borrowers and a predicted one-year PD of 2%. At year end 14 borrowers defaulted. Assuming independent defaults, test at a one-sided 95% level (critical z = 1.645) whether the PD is underestimated.
Show the solution
- Expected defaults = n × p = 400 × 0.02 = 8.
- Standard deviation = √(400 × 0.02 × 0.98) = √7.84 = 2.8.
- z = (14 − 8) ÷ 2.8 = 6 ÷ 2.8 = 2.143.
- Compare: 2.143 is greater than 1.645, so reject the null that the PD is correct.
Answer: z ≈ 2.14, above 1.645. There is evidence at the 95% one-sided level that the 2% PD is underestimated. Because defaults may be correlated, this simple test may overstate the evidence.
Exam tips
- Always label the question as discrimination or calibration first. Many wrong options swap the two.
- Memorise AR = 2 × AUC − 1 and that random means AR 0 and AUC 0.5.
- Show expected defaults, standard deviation and z in that order. The arithmetic is quick and easy to check.
- Expect conceptual options about default correlation, procyclicality, point-in-time versus through-the-cycle PDs and out-of-sample testing.
- If observed defaults exceed expected, the answer direction is PD underestimated. Check the direction before choosing.
Practice questions from External and Internal Credit Ratings
- Under the S&P Global Ratings long-term issuer credit rating scale, which of the following is the lowest rating that is still classified as i…
- Which of the following is a commonly cited criticism of the issuer-pays model used by major credit rating agencies?
- An analyst compares a corporate bond rated BBB with a structured-finance tranche also rated BBB. Which statement best reflects a known limit…
- A bank validates its internal rating system by comparing the ranking of obligors with subsequent default outcomes. The validation team compu…
- A bank's risk committee notes that external agency ratings tend to change slowly even when a borrower's market-implied credit quality deteri…
Validation and Backtesting of Rating Systems in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Validation and Backtesting of Rating Systems: frequently asked questions
What is the difference between discriminatory power and calibration?
Discriminatory power is the ability to rank borrowers so that defaulters sit in worse grades. Calibration is how close the predicted PD of each grade is to the realised default rate. A model can have one without the other.
How do I read a CAP curve?
Borrowers are sorted from worst to best on the x-axis and the cumulative share of defaulters is on the y-axis. The further the curve bows above the diagonal, the better the model ranks. The accuracy ratio summarises this as a number between 0 and 1.
How is backtesting of PD done?
For each grade you compare the observed default rate with the predicted PD. A binomial test, or its normal approximation, checks whether the gap is bigger than chance. Correlation between defaults means the test should be read with caution.
Is the accuracy ratio the same as the Gini coefficient?
In credit scoring the accuracy ratio is equal to the Gini coefficient. Both equal 2 × AUC − 1. Check the wording of the question, but the values are the same.