Skip to content

FRM Exam Part II · Credit Scoring and Retail Credit Risk Management

Model Risk, Fair Lending and Regulation in Credit Scoring

Updated 11 October 2026 · Fact-checked

Model risk in retail credit scoring is the chance of loss or poor decisions because a scorecard is built on bad data, wrongly specified, misused or left unvalidated. Fair lending adds a legal limit: models must not discriminate on protected traits. Regulators expect governance, validation, documentation, monitoring and explainable decisions.

Understand Model Risk, Fair Lending and Regulatory Considerations

A credit scoring model turns borrower data into a score that drives approval, pricing and limits. It is only as good as its data and its assumptions. Model risk is the risk of adverse outcomes from decisions based on a model that is incorrect or used incorrectly. It has two sources: errors in the model itself, and misuse of a sound model.

Data quality is the first weak point. Typical problems are missing values, input errors, inconsistent definitions of default, and stale data. Another is sample selection bias (reject inference problem): you only observe repayment for applicants who were accepted, so the development sample misses the rejected population. A model built on accepted accounts may mis-rank the people it now has to judge. Population drift is a related issue: the borrowers today differ from those in the development sample, or the economy has changed, so the score loses power. Models built in benign periods may understate defaults in a downturn.

Fair lending rules prohibit discrimination on protected characteristics such as race, sex, religion, national origin, age or marital status (the exact list depends on the jurisdiction). Two ideas matter. Disparate treatment is intentional different treatment of a protected group, for example using the characteristic directly. Disparate impact is a neutral-looking variable or rule that disproportionately hurts a protected group without a legitimate business need, or when a less discriminatory alternative exists. Proxy variables such as postcode can act as stand-ins for protected traits. Removing the protected variable is therefore not enough. You must test outcomes.

Regulatory expectations focus on governance. Under Basel internal ratings-based rules, rating systems need sound design, data history, validation and use in management (the use test). Supervisory model risk guidance (such as SR 11-7) asks for clear model definition, independent validation with effective challenge, documentation, ongoing monitoring and board oversight. Customers should also get adverse action reasons, so models must be explainable. Complex machine learning models raise the explainability burden.

More broadly, scoring has limits. It is backward looking, assumes the future resembles the past, captures only the factors in the data, and can be gamed or become less predictive over time. Overrides and judgement need controls too.

Key formulas to remember

Model risk definition
Model risk = risk of adverse outcomes from decisions based on incorrect or misused model outputs
Two sources: fundamental errors in the model, and incorrect or inappropriate use.
Adverse impact ratio (selection rate test)
Impact ratio = approval rate of protected group ÷ approval rate of reference group
A common screen; a ratio well below 1 flags possible disparate impact. The 0.80 level (four-fifths rule) is a US employment-testing rule of thumb, not a universal lending law. It is a screen, not proof.
Disparate treatment vs disparate impact
Treatment = intent or direct use of protected trait; Impact = neutral rule with unjustified disproportionate effect
Impact can arise through proxy variables even when the protected trait is excluded.
Core validation components
Conceptual soundness + ongoing monitoring + outcomes analysis (backtesting)
These are the three pillars of model validation in supervisory guidance such as SR 11-7.
Population stability index (PSI)
PSI = Σ (Actual% − Expected%) × ln(Actual% ÷ Expected%)
Summed across score bands. Higher values mean a larger shift in the score distribution; common rules of thumb treat below 0.10 as stable and above 0.25 as a major shift.

How to solve Model Risk, Fair Lending and Regulatory Considerations questions

Use this sequence for any question on model risk, fair lending or regulation of retail scoring models.

  1. 1Identify the issue type: data quality, model specification, model use, discrimination, or governance and regulation.
  2. 2Name the precise concept: sample selection bias, population drift, proxy variable, disparate treatment, disparate impact, effective challenge, use test.
  3. 3Check the facts: is a protected trait used directly (treatment), or does a neutral variable produce unequal outcomes (impact)?
  4. 4If numbers are given, compute the metric (approval rates, impact ratio, PSI) and compare with the stated benchmark.
  5. 5Interpret: what does the result imply for model reliability or legal exposure? Remember a screen is not proof.
  6. 6Choose the remedy that matches the cause: better data or reject inference, redevelopment or recalibration, removing or replacing a proxy, independent validation, or stronger governance.
  7. 7Eliminate options that overstate, such as 'removing the protected variable guarantees fairness'.

Quickest way: Label the problem, then match the fix

When to use it: Use when you have about one minute per multiple-choice question and the options are close.

  1. Underline the cause in the stem: missing data, rejected applicants, shifted population, protected trait, proxy, no independent review.
  2. Map it: rejected applicants → selection bias; shifted scores → drift or PSI; neutral variable with unequal outcomes → disparate impact; direct use → disparate treatment; no challenge → weak governance.
  3. Pick the answer that fixes that exact cause.
  4. Reject any option with absolute words such as 'always', 'eliminates' or 'guarantees'.

Common mistakes in Model Risk, Fair Lending and Regulatory Considerations

  • Believing that dropping the protected variable makes a model fair.

    It sounds like the obvious fix for discrimination.

    Fix: Remember proxies. Test outcomes by group and look for less discriminatory alternatives.

  • Confusing disparate treatment with disparate impact.

    Both involve unequal results for groups.

    Fix: Treatment concerns intent or direct use of the trait. Impact concerns a neutral rule with unjustified unequal effect.

  • Treating the four-fifths ratio as a legal pass/fail test for lending.

    It is quoted often and gives a clean number.

    Fix: Call it a screening rule of thumb. Business justification and alternatives still matter.

  • Ignoring sample selection bias because the model has a high measured accuracy.

    Accuracy is measured only on accepted accounts that have outcomes.

    Fix: Recognise that rejected applicants have no observed performance and consider reject inference.

  • Treating validation as the developer checking their own work.

    Candidates mix up testing and validation.

    Fix: Validation needs independence and effective challenge from staff with competence and authority.

  • Assuming a model that worked at development will keep working.

    Past performance feels like evidence.

    Fix: Monitor drift, backtest outcomes and recalibrate or redevelop when performance decays.

Worked examples

Example 1

A bank approves 360 of 600 applicants from Group A (reference group) and 210 of 500 applicants from Group B (protected group). Compute the impact ratio and interpret it against the four-fifths screen.

Show the solution
  1. Approval rate for Group A = 360 ÷ 600 = 0.60.
  2. Approval rate for Group B = 210 ÷ 500 = 0.42.
  3. Impact ratio = 0.42 ÷ 0.60 = 0.70.
  4. Compare with 0.80: 0.70 is below the screen.
  5. Interpretation: this flags possible disparate impact and calls for review of variables, proxies, business justification and less discriminatory alternatives. It does not prove discrimination.

Answer: Impact ratio = 0.70, below 0.80, so the model is flagged for disparate impact review but is not proven discriminatory.

Example 2

A bank's scorecard was built only on accepted applicants from a boom period. Which risk does this most directly create, and what is the best response? Options: A) Sample selection bias; use reject inference and monitor performance. B) Disparate treatment; remove the age variable. C) Settlement risk; add collateral. D) Liquidity risk; raise the cash buffer.

Show the solution
  1. The development sample contains only accepted applicants, so performance of rejected applicants is unobserved.
  2. This is sample selection bias, which can make the model mis-rank the full applicant population.
  3. The boom-period data adds a second concern: defaults may be understated in a downturn.
  4. Option B concerns intent or direct trait use, which the facts do not describe.
  5. Options C and D concern unrelated risks.
  6. The best response is reject inference, stress or downturn calibration, and ongoing monitoring.

Answer: A. Sample selection bias; use reject inference and monitor performance.

Exam tips

  • Questions often hide the concept in a short case. Name it first, then pick the option.
  • Expect distractors that overstate: 'eliminates bias', 'guarantees accuracy', 'proves discrimination'. Reject them.
  • Learn the difference between treatment and impact and the role of proxies. It is a common test.
  • For regulation, link scoring models to Basel IRB design, validation and use test, and to SR 11-7 style governance.
  • If given approval rates, compute the ratio carefully and divide protected by reference.

Practice questions from Credit Scoring and Retail Credit Risk Management

Model Risk, Fair Lending and Regulatory Considerations: frequently asked questions

What is model risk in credit scoring?

It is the risk of poor decisions or losses because a scorecard is flawed or misused. Causes include bad data, selection bias, drift, weak validation and use outside the model's intended purpose.

What is the difference between disparate treatment and disparate impact?

Disparate treatment is intentional different treatment or direct use of a protected trait. Disparate impact is a neutral-looking rule that disproportionately harms a protected group without adequate business justification. Proxy variables often cause impact.

What are the main limitations of credit scoring models?

They look backward, assume stable relationships, depend on data quality and miss rejected applicants' outcomes. Scores also lose power as populations and economic conditions change, so they need monitoring and recalibration.

What do regulators expect for retail credit models under Basel?

Under the internal ratings-based approach, rating systems need sound design, adequate data, validation and real use in risk management. Supervisory model risk guidance adds independent validation, documentation, monitoring and board oversight.