Skip to content

FRM Exam Part II · Credit Scoring and Retail Credit Risk Management

Scorecard Validation: ROC, AUC, Gini, KS and PSI

Updated 11 October 2026 · Fact-checked

Scorecard validation tests whether a credit score ranks borrowers correctly and whether the population it scores is still the one it was built on. Discrimination is measured by ROC, AUC, Gini (2 × AUC − 1) and KS (maximum gap between cumulative good and bad distributions). Stability is measured by the population stability index (PSI).

Understand Scorecard Validation and Performance Metrics

A scorecard gives each borrower a score. A good scorecard gives bad accounts (those that default) low scores and good accounts high scores. Validation asks two questions: does it rank well, and is it still stable?

Discrimination is the power to separate goods from bads. The ROC curve plots the true positive rate (share of bads caught) against the false positive rate (share of goods wrongly flagged) as you move the cut-off score. A random model gives the 45-degree diagonal. A better model bows towards the top-left corner.

The AUC is the area under the ROC curve. It equals the probability that a randomly chosen bad has a lower score than a randomly chosen good. AUC of 0.5 means no skill. AUC of 1 means perfect ranking. The Gini coefficient (also called the accuracy ratio) rescales AUC so that 0 means no skill and 1 means perfect: Gini = 2 × AUC − 1. They carry the same information.

The KS statistic is the largest vertical gap between the cumulative distribution of bads and the cumulative distribution of goods across scores. It is read in percentage points or as a number between 0 and 1. A higher KS means better separation at the best cut-off.

Stability is a different question. The population stability index (PSI) compares the score distribution of today's applicants with the development sample, bucket by bucket. A large PSI says the population has shifted, so the scorecard may no longer work as designed. Full validation also includes backtesting: comparing predicted PD with realised default rates, and checking calibration, not just ranking. A model can rank well and still be badly calibrated.

Key formulas to remember

Gini coefficient
Gini = 2 × AUC − 1
Equivalent to AUC. AUC = (Gini + 1) ÷ 2. Ranges 0 to 1 for a model that ranks in the right direction.
True positive rate (sensitivity)
TPR = bads below cut-off ÷ total bads
Y-axis of the ROC curve. Defined with low score meaning high risk.
False positive rate
FPR = goods below cut-off ÷ total goods
X-axis of the ROC curve.
KS statistic
KS = max over scores | F_bad(s) − F_good(s) |
F is the cumulative share of each group at or below score s. Equals the maximum of TPR − FPR.
Population stability index
PSI = Σ (A_i − E_i) × ln(A_i ÷ E_i)
A_i is the actual (recent) share in bucket i, E_i the expected (development) share. Shares are fractions that sum to 1.
PSI rule of thumb
< 0.10 stable; 0.10 to 0.25 some shift; > 0.25 significant shift
Industry convention, not a regulatory rule. Cut-offs vary by firm.
AUC interpretation
AUC = P(score of random bad < score of random good)
Ties are usually counted as half.

How to solve Scorecard Validation and Performance Metrics questions

Use this order for any question on scorecard testing.

  1. 1Identify what is being tested: ranking power (ROC, AUC, Gini, KS), calibration (predicted vs realised PD) or stability (PSI).
  2. 2Fix the direction: low score = higher risk, so bads should sit at low scores. Check how cumulative columns are built.
  3. 3For KS, compute cumulative % of bads and cumulative % of goods at each score band, take the absolute difference, and pick the maximum.
  4. 4For AUC and Gini, use the given AUC or Gini and convert with Gini = 2 × AUC − 1. If given ROC points, use the trapezium rule.
  5. 5For PSI, convert counts to shares, compute (A − E) × ln(A ÷ E) for each bucket, and sum.
  6. 6Compare the result with the benchmark thresholds and state a conclusion: keep, monitor, recalibrate or redevelop.
  7. 7Check the answer's logic: AUC must lie between 0.5 and 1 for a useful model, and each PSI bucket term must be zero or positive.

Quickest way: Convert and compare, do not rebuild

When to use it: When the question gives AUC, Gini, KS or PSI values and asks for a conversion or a conclusion.

  1. If Gini is given, AUC = (Gini + 1) ÷ 2. If AUC is given, Gini = 2 × AUC − 1.
  2. Eliminate options with AUC below 0.5 or Gini above 1 or below 0.
  3. For KS, scan the table for the biggest gap between the two cumulative columns. It is often in the middle bands.
  4. For PSI, compute each bucket term quickly: sign of (A − E) and ln(A ÷ E) always match, so every term is positive.
  5. Apply the thresholds 0.10 and 0.25 for the verdict.

Common mistakes in Scorecard Validation and Performance Metrics

  • Treating Gini and AUC as two different measures of performance.

    They have different scales and different names.

    Fix: Remember Gini = 2 × AUC − 1. AUC 0.80 means Gini 0.60. They carry the same information.

  • Saying AUC of 0.5 is the worst possible result.

    Students assume the scale starts at zero like Gini.

    Fix: AUC 0.5 is random ranking. AUC below 0.5 means the scores are inverted, which you can fix by flipping the sign.

  • Taking KS as the difference at one chosen cut-off instead of the maximum across all scores.

    Credit officers think in terms of a single approval cut-off.

    Fix: KS is the maximum gap. Compute the gap at every score band and report the largest.

  • Using PSI to judge whether the scorecard ranks well.

    PSI is reported alongside Gini in monitoring packs.

    Fix: PSI only compares score distributions. It says nothing about default outcomes. Use Gini, AUC and KS for discrimination and backtesting for calibration.

  • Calculating PSI with counts or percentages that do not sum to the same base, or dropping the natural log.

    Rushing the arithmetic.

    Fix: Convert to shares that sum to 1 in both columns. Use ln, not log base 10.

  • Assuming a high Gini means the PDs are accurate.

    Ranking and calibration are confused.

    Fix: Gini is unchanged if every PD is doubled. Check calibration separately by comparing predicted PD with the observed default rate.

Worked examples

Example 1

A scorecard has 1,000 accounts: 100 bads and 900 goods. Cumulative shares at four score cut-offs (lowest scores first) are: Band 1: bads 40%, goods 5%. Band 2: bads 70%, goods 20%. Band 3: bads 90%, goods 50%. Band 4: bads 100%, goods 100%. Find the KS statistic.

Show the solution
  1. Compute the absolute gap at each band.
  2. Band 1: 40% − 5% = 35 points.
  3. Band 2: 70% − 20% = 50 points.
  4. Band 3: 90% − 50% = 40 points.
  5. Band 4: 100% − 100% = 0 points.
  6. The maximum is 50 points, at Band 2.

Answer: KS = 50 percentage points (0.50), reached at Band 2.

Example 2

The development sample score distribution across three buckets is 30%, 50%, 20%. This quarter's applicants are 20%, 50%, 30%. Calculate the PSI and state the conclusion.

Show the solution
  1. Bucket 1: (0.20 − 0.30) × ln(0.20 ÷ 0.30) = (−0.10) × ln(0.6667) = (−0.10) × (−0.4055) = 0.04055.
  2. Bucket 2: (0.50 − 0.50) × ln(1) = 0.
  3. Bucket 3: (0.30 − 0.20) × ln(0.30 ÷ 0.20) = 0.10 × ln(1.5) = 0.10 × 0.4055 = 0.04055.
  4. PSI = 0.04055 + 0 + 0.04055 = 0.0811, about 0.081.
  5. 0.081 is below 0.10, so the shift is small under the usual rule of thumb.

Answer: PSI ≈ 0.081. The population is stable under the common convention (below 0.10), so no redevelopment is triggered, though monitoring continues.

Exam tips

  • Convert freely between AUC and Gini. Many questions are one-line conversions: Gini 0.60 means AUC 0.80.
  • Read the direction of the score carefully. Check whether bads are cumulated from the lowest score or the highest.
  • When asked what PSI tells you, answer stability of the population, not predictive power.
  • If a question mixes high Gini with poor default-rate matching, the answer is a calibration problem, not a ranking problem.
  • State PSI thresholds as industry rules of thumb. Do not call them Basel requirements.

Practice questions from Credit Scoring and Retail Credit Risk Management

Scorecard Validation and Performance Metrics: frequently asked questions

What is the difference between Gini coefficient and AUC?

Both measure ranking power from the same ROC curve. Gini = 2 × AUC − 1, so AUC 0.5 equals Gini 0, and AUC 1 equals Gini 1. Gini is the area between the ROC curve and the diagonal, scaled to a maximum of 1.

How do you calculate the KS statistic for a scorecard?

Build cumulative distributions of bads and goods across score bands. Take the absolute difference at each band. The largest difference is the KS statistic.

What is a good PSI value for a credit scorecard?

A common convention is below 0.10 for stable, 0.10 to 0.25 for a moderate shift and above 0.25 for a significant shift. These are rules of thumb that firms may adjust.

Does a high Gini mean the scorecard is well calibrated?

No. Gini measures ranking only. Calibration compares predicted PDs with realised default rates, usually by score band, and needs separate backtesting.