Skip to content

FRM Exam Part II · Advances in Artificial Intelligence: Implications for Capital Markets Activities

AI Model Risk, Explainability and Data Quality

Updated 11 October 2026 · Fact-checked

AI model risk is the chance of loss or bad decisions because an AI model is wrong, misused or not understood. Main drivers are opacity (black box), overfitting, bias, poor data and hallucinations. You solve questions by naming the source, linking it to the control (validation, explainability, data governance, human oversight) and judging residual risk.

Understand Model Risk, Explainability and Data Quality Challenges

A model turns inputs into estimates used for decisions. Model risk, as SR 11-7 frames it, comes from two places: errors in the model itself (wrong design, data or implementation) and incorrect or inappropriate use of a correct model. AI and machine learning (ML) keep both sources and add new ones.

The black box problem means a complex model, such as a deep neural network or large ensemble, gives outputs but you cannot easily trace why. This makes validation harder. A validator cannot check that the logic is conceptually sound if the logic is not visible. Explainability tools, such as feature importance, partial dependence, SHAP or LIME, give approximate explanations after the fact. They help, but they describe the model; they do not make it inherently transparent. A simpler, interpretable model may be preferred where decisions must be justified, for example credit denial.

Overfitting is when a model learns noise in the training data, so it looks excellent in-sample and performs badly on new data. Flexible AI models overfit easily, especially with short or regime-specific histories. Controls: out-of-sample and out-of-time testing, cross-validation, regularisation and monitoring live performance. Model drift is the later decay when markets or customer behaviour change from the training period (data drift or concept drift).

Bias arises when training data or design leads to systematically unfair or skewed outputs, for example historical lending data reflecting past discrimination, or proxy variables standing in for protected traits. Bias is one type of model risk, not the same thing. Model risk is the wider category; it also covers instability, wrong use and implementation errors.

Data quality drives everything: inaccurate, incomplete, stale, unrepresentative or poorly labelled data feeds directly into outputs (garbage in, garbage out). Hallucinations occur in generative AI and large language models, which can produce fluent but false or invented content. This matters for research summaries, compliance and client communication. Governance needs inventory of models, independent validation with effective challenge, documentation, monitoring, human oversight and clear accountability, including for vendor and third-party models.

Key formulas to remember

SR 11-7 model risk sources
Model risk = fundamental errors in the model + incorrect or inappropriate use
Both count. A sound model used outside its intended purpose is still model risk.
Overfitting signature
In-sample error much lower than out-of-sample error
A large gap suggests the model memorised noise. Fix with out-of-sample testing, regularisation, simpler models.
Three pillars of model governance
Development and use → Validation with effective challenge → Governance, policies and controls
Validation must be independent of those who build the model.
Data quality dimensions
Accuracy, completeness, timeliness, consistency, representativeness
Use as a checklist when a question describes a data problem.
Drift types
Data drift = input distribution changes; Concept drift = input-to-output relationship changes
Both lead to degraded live performance and call for monitoring and retraining.

How to solve Model Risk, Explainability and Data Quality Challenges questions

Use this sequence for any scenario question on AI model risk, explainability or data.

  1. 1Identify the symptom in the scenario: unexplained output, great backtest but poor live results, skewed outcomes, bad inputs or false text.
  2. 2Map it to the source: opacity, overfitting, drift, bias, data quality or hallucination.
  3. 3Decide whether the problem is in the model itself or in how it is used (SR 11-7 distinction).
  4. 4Pick the control that fits the source: explainability tools or simpler model, out-of-sample testing, monitoring, bias testing, data governance, or human review.
  5. 5Check governance: independent validation, effective challenge, documentation, inventory, accountability, vendor oversight.
  6. 6Judge the residual risk and the proportionate response, for example limits on use, human override, or retraining.
  7. 7Eliminate options that overclaim, such as 'explainability tools remove opacity' or 'more data always fixes bias'.

Quickest way: Symptom-to-control matching

When to use it: Use when you have about 90 seconds per MCQ and the options list several plausible controls.

  1. Underline the symptom word: 'cannot explain', 'out-of-sample', 'unfair', 'stale', 'fabricated'.
  2. Match it: cannot explain = explainability; out-of-sample gap = overfitting; unfair = bias; stale or missing = data quality; fabricated = hallucination.
  3. Choose the option that treats that exact source, not a generic one.
  4. Reject absolutes like 'eliminates', 'guarantees' or 'always'.

Common mistakes in Model Risk, Explainability and Data Quality Challenges

  • Treating model risk and AI bias as the same thing

    Both appear in fairness discussions.

    Fix: Bias is one source of model risk. Model risk also covers errors, instability, drift and misuse.

  • Believing explainability tools make a black box transparent

    Tools like SHAP sound definitive.

    Fix: They give approximate post-hoc explanations. Validation still needs challenge, testing and sometimes a simpler model.

  • Judging a model by in-sample fit

    High accuracy looks like quality.

    Fix: Look for out-of-sample and out-of-time results. A big gap signals overfitting.

  • Assuming that more data fixes bias

    Volume feels like quality.

    Fix: Biased or unrepresentative data stays biased at scale. Fix needs data review, bias testing and design changes.

  • Ignoring model use when the model itself is sound

    Candidates focus on technical errors.

    Fix: SR 11-7 treats inappropriate use as model risk. Check intended purpose and limits.

  • Exempting vendor or third-party AI from validation

    Candidates think the vendor carries the risk.

    Fix: The bank remains accountable. It needs documentation, testing and monitoring of vendor models as far as it can.

Worked examples

Example 1

A bank's ML credit-scoring model shows 96% accuracy on training data and 78% on a later holdout sample. The model is a deep ensemble that credit officers cannot explain to rejected applicants. Identify the two main model risk issues and one control for each.

Show the solution
  1. Compare the accuracy: 96% in-sample versus 78% out-of-sample is a gap of 18 percentage points.
  2. A large gap points to overfitting: the model learned noise in the training data.
  3. Control for overfitting: out-of-time validation, cross-validation, regularisation, or a simpler model, plus live performance monitoring.
  4. Officers cannot explain decisions: this is the black box, or explainability, problem.
  5. Control for opacity: explainability tools such as feature importance or SHAP for local explanations, or an interpretable model for decisions needing justification.
  6. Validation should be independent and apply effective challenge to both issues.

Answer: Overfitting (control: out-of-sample testing and regularisation) and black-box opacity (control: explainability tools or a simpler interpretable model), under independent validation.

Example 2

A bank deploys a generative AI assistant to summarise regulatory updates for compliance staff. It cites a regulation that does not exist. Which risk is this, and which control is most appropriate: (A) retrain with more market data, (B) mandatory human review of outputs before use, (C) increase the model's size, (D) remove validation to speed up release?

Show the solution
  1. A confident but fabricated citation is a hallucination.
  2. Hallucinations come from how generative models produce plausible text, not from lack of market data, so A does not address it.
  3. Larger size does not guarantee removal of hallucinations, so C is not reliable.
  4. Removing validation raises model risk, so D is wrong.
  5. Human review before use catches fabricated content and keeps accountability with staff, so B fits.

Answer: B. This is hallucination risk, controlled by mandatory human review, alongside validation and defined limits on use.

Exam tips

  • Start with the symptom in the stem; the question usually points to exactly one source of model risk.
  • Watch for absolute words. Answers saying a tool eliminates opacity or bias are usually wrong.
  • Link SR 11-7 ideas: sound development, independent validation with effective challenge, and governance.
  • Distinguish model risk (broad) from bias (one cause) and data quality (an input problem).
  • For generative AI, expect hallucination and human oversight as the paired concept.

Practice questions from Advances in Artificial Intelligence: Implications for Capital Markets Activities

Model Risk, Explainability and Data Quality Challenges: frequently asked questions

What is the black box problem in machine learning finance?

It is the difficulty of tracing why a complex model produced a particular output. This hampers validation, regulatory explanation and customer disclosure. Explainability tools and simpler models partly reduce it.

What is the difference between model risk and AI bias?

Model risk is the broad risk of loss from model errors or misuse. Bias is one source, where data or design produce systematically skewed outcomes. Model risk also covers overfitting, drift and wrong use.

Does SR 11-7 apply to AI models?

SR 11-7 is supervisory guidance on model risk management that defines a model broadly, so it is widely applied to AI and ML. Its principles of sound development, validation with effective challenge and governance carry over, though AI makes them harder to apply.

How is overfitting detected?

Compare performance on training data with performance on data the model has not seen, including later periods. A large drop out-of-sample signals overfitting. Ongoing monitoring of live results also helps.