Skip to content

FRM Exam Part II · The Financial Stability Implications of Artificial Intelligence

AI Model Risk, Data Quality and Governance for FRM Part II

Updated 11 October 2026 · Fact-checked

AI model risk is the chance of loss or poor decisions because an AI model is wrong, opaque, trained on poor data or misused. You manage it with good data controls, independent validation, explainability tools, human oversight, monitoring and clear board-level governance. In exam questions, match each weakness to its control.

Understand Model Risk, Data Quality and Governance

A model is a simplified tool that turns data into a decision or estimate. Model risk is the harm that arises when the model is wrong or used badly. AI and machine learning (ML) models add to this risk because they are often complex and learn patterns from data rather than from rules that a person writes down.

The first issue is opacity. Deep neural networks and large ensembles have thousands or millions of parameters. You can see the inputs and the output, but not easily why the output came out that way. Explainability is the ability to describe, in terms a user or supervisor accepts, what drives a result. Tools such as feature importance, partial dependence plots, SHAP and LIME give approximate explanations. They do not make the model transparent. They can be unstable, and a good explanation of a model is not proof that the model is correct.

The second issue is data quality. A model can only learn from what it is given. Missing values, errors, stale data, unrepresentative samples and historical bias all flow into the output. If past lending decisions were biased, a model trained on them can repeat the bias, even if protected attributes are removed, because other variables act as proxies. Data drift (inputs change) and concept drift (the relationship changes) make a once-good model degrade. Models also overfit: they fit noise in the training data and fail on new data. In stress, history may hold no similar period at all.

The third issue is generative AI. Large language models can produce hallucinations: fluent, confident output that is false or invented. They can also leak confidential data, give different answers to the same prompt, and depend on a vendor's model that the bank cannot inspect. Third-party and concentration risk arise when many firms use the same provider or foundation model.

The answer is governance. This means an inventory of all models, risk-tiering by materiality, clear ownership, documented development, independent validation and effective challenge, ongoing monitoring with thresholds, human-in-the-loop review for high-impact decisions, change control, fallback plans and board reporting. Principles from supervisory guidance such as SR 11-7 apply to AI too, but validators must adapt them: more focus on data lineage, out-of-sample and out-of-time testing, benchmarking against simpler models, bias testing and stability. Financial stability authorities add a system view: common models, data and vendors can make firms act alike and amplify shocks.

Key formulas to remember

Model risk drivers
Model risk = f(model error, data error, misuse)
A conceptual rule, not a calculation. Use it to sort a case into design, data or use problems.
Out-of-sample performance check
Generalisation gap = in-sample performance − out-of-sample performance
A large positive gap signals overfitting. Test on held-out and out-of-time data.
Drift monitoring idea
Alert if monitored metric breaches its pre-set threshold
Metrics include accuracy, population stability and input distribution shifts. Thresholds must be set before deployment.
Proportionality rule
Validation depth rises with model materiality and complexity
High-impact, opaque models need the strongest validation and oversight.

How to solve Model Risk, Data Quality and Governance questions

Use this sequence for any question on AI model risk, data or governance.

  1. 1Identify the source of the problem: model design, data, use, or third-party dependence.
  2. 2Name the specific risk: opacity, overfitting, drift, bias, hallucination, or concentration.
  3. 3Ask what harm follows: wrong decision, unfair outcome, loss, compliance breach or systemic herding.
  4. 4Pick the control that targets that source: data checks, validation, explainability tools, monitoring, human review or vendor oversight.
  5. 5Check the governance layer: ownership, inventory, independent challenge, escalation and board reporting.
  6. 6Apply proportionality: more material and more opaque means stronger controls.
  7. 7Eliminate options that overclaim, such as saying a tool removes risk or that explainability proves accuracy.

Quickest way: Match the weakness to the control

When to use it: Use it for short MCQs where you must pick the best control or the main risk.

  1. Underline the symptom in the stem (for example, confident but false output, or performance fell after market change).
  2. Map it: false output means hallucination, so human review and grounding; fall after change means drift, so monitoring and recalibration; unfair outcomes means bias, so data review and fairness testing.
  3. Prefer independent, ongoing and proportionate controls over one-off or self-review answers.
  4. Reject absolute words like always, eliminates or guarantees.

Common mistakes in Model Risk, Data Quality and Governance

  • Treating explainability tools as proof that a model is correct.

    A clear chart feels like validation.

    Fix: Remember that explanations are approximations. Validation still needs performance, stability and conceptual soundness testing.

  • Assuming removing protected attributes removes bias.

    It seems logical that no input means no discrimination.

    Fix: Proxy variables and biased historical labels can still carry bias. Test outcomes across groups.

  • Confusing data drift with concept drift.

    Both cause performance to fall.

    Fix: Data drift is a change in input distribution. Concept drift is a change in the relationship between inputs and outcome.

  • Letting the model developer validate the model.

    Developers know the model best.

    Fix: Validation needs independence and effective challenge, usually from a second line function.

  • Thinking vendor models need less oversight.

    The vendor is seen as the expert.

    Fix: The bank stays accountable. Demand documentation, testing access, performance monitoring and exit plans, and watch concentration.

  • Treating hallucination as a rare software bug.

    It looks like a simple fault.

    Fix: It is an inherent feature of generative models. Use human review, grounding in verified sources and usage limits.

Worked examples

Example 1

A bank uses a gradient-boosted model for credit approvals. In development it classifies 92% of cases correctly in-sample, but only 78% on a later out-of-time sample. Which problem is most likely, and what is the best response?

Show the solution
  1. Compute the gap: 92% − 78% = 14 percentage points.
  2. A large drop on new data signals overfitting or drift, not a data entry error alone.
  3. Because the weakness appears on an out-of-time sample, check whether the model fit noise or conditions changed.
  4. Best response: independent validation, simplify or regularise, compare with a simpler benchmark, and set monitoring thresholds before release.

Answer: Overfitting (possibly with drift). The generalisation gap is 14 percentage points, so require independent out-of-sample validation, regularisation or a simpler benchmark, and ongoing monitoring.

Example 2

An analyst uses a generative AI assistant to draft a regulatory summary. It cites a rule that does not exist. Name the risk and the governance controls the firm should apply.

Show the solution
  1. The output is fluent but false, so this is a hallucination.
  2. The harm is a compliance and reputational breach if the text is used unchecked.
  3. Controls: human review of all outputs used externally, grounding in approved source documents, usage policy, logging of prompts and outputs.
  4. Governance: place the tool in the model inventory, risk-tier it, and validate and monitor it with vendor oversight if third-party.

Answer: Hallucination. Apply mandatory human review, source grounding, usage policy and logging, and bring the tool under model inventory, validation, monitoring and vendor oversight.

Exam tips

  • Expect scenario stems. Identify the root cause first, then choose the control that fits it.
  • Words like independent, ongoing and proportionate usually mark the best governance answer.
  • Watch for overclaims: explainability tools do not eliminate model risk and vendor models do not shift accountability.
  • Link to financial stability: common models, data or providers can cause herding and concentration.

Practice questions from The Financial Stability Implications of Artificial Intelligence

Model Risk, Data Quality and Governance in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Model Risk, Data Quality and Governance: frequently asked questions

What is the difference between explainability and interpretability?

Interpretability usually means a model is understandable by design, such as a linear model. Explainability often means using extra tools to describe a complex model after the fact. Exams focus on the fact that post-hoc explanations are approximate.

How do you validate an AI model?

Review conceptual soundness, data lineage and quality, and test out-of-sample and out-of-time performance. Benchmark against a simpler model, test stability and bias, and set up ongoing monitoring. Validation must be independent and proportionate to materiality.

Why are hallucinations a governance issue?

Generative models can produce confident but false output, and this cannot be fully removed. Firms need human review, grounded sources, use limits and logging. Without these, false output can reach clients or regulators.

Does SR 11-7 apply to AI models?

Its principles of sound development, validation, effective challenge and governance apply to any model, including AI. Validators must adapt techniques to the greater complexity, opacity and data dependence of AI.