FRM Exam Part II · Advances in Artificial Intelligence: Implications for Capital Markets Activities
AI in Risk Management, Compliance and Surveillance
Updated 11 October 2026 · Fact-checked
AI in risk management means using machine learning to find patterns in large datasets. Banks apply it to credit and market risk models, fraud detection, RegTech compliance and market abuse surveillance. It improves prediction and speed, but brings explainability, data quality, model risk and governance challenges you must address.
Understand AI in Risk Management, Compliance and Surveillance
Artificial intelligence (AI) in finance mostly means machine learning (ML). ML models learn patterns from data instead of following fixed rules written by a person. They can handle many more inputs than a traditional model and can capture nonlinear relationships and interactions.
In credit risk, ML models such as gradient-boosted trees and neural networks predict probability of default using traditional data and, sometimes, alternative data like transaction histories. They often rank borrowers better than a logistic regression scorecard. The cost is lower transparency. Supervisors and validators want to know why a borrower was declined.
In market risk, ML can help with volatility forecasting, scenario generation, risk factor selection and proxying illiquid positions. Models trained on calm periods may fail in a regime change. This is model drift or non-stationarity.
In fraud detection and surveillance, ML finds unusual behaviour. Supervised learning uses labelled fraud cases. Unsupervised learning (clustering, anomaly detection) finds outliers without labels. Surveillance systems also use natural language processing (NLP) on chats, emails and voice to spot market abuse such as insider dealing, spoofing and manipulation. Fraud is rare, so classes are imbalanced and false positives are a big operational cost.
In RegTech, AI automates compliance: KYC and AML screening, transaction monitoring, regulatory reporting and reading rule changes. The benefit is efficiency and fewer missed alerts. The risks are opacity, biased or poor data, reliance on third-party vendors, and the need for human oversight and clear accountability.
Key formulas to remember
- Precision
- Precision = TP ÷ (TP + FP)
- Share of alerts that are real cases. Low precision means many false positives and heavy review workload.
- Recall (detection rate)
- Recall = TP ÷ (TP + FN)
- Share of real cases caught. Low recall means missed fraud or abuse.
- False positive rate
- FPR = FP ÷ (FP + TN)
- Share of legitimate cases wrongly flagged.
- Accuracy
- Accuracy = (TP + TN) ÷ (TP + TN + FP + FN)
- Misleading when classes are imbalanced, as in fraud.
- Overfitting rule
- In-sample performance ≫ out-of-sample performance
- A large gap signals overfitting. Test on held-out data.
How to solve AI in Risk Management, Compliance and Surveillance questions
Use this method for any scenario question on AI in risk, compliance or surveillance.
- 1Identify the use case: credit, market risk, fraud, compliance or surveillance.
- 2Identify the learning type: supervised, unsupervised or NLP.
- 3State the benefit the scenario points to, such as better prediction, speed or coverage.
- 4Find the key risk: explainability, overfitting, data quality or bias, drift, imbalance, or third-party dependence.
- 5If numbers are given, compute precision, recall or false positive rate from the confusion matrix.
- 6Choose the control: validation, out-of-sample testing, human oversight, monitoring, documentation or governance.
- 7Match the answer to the exact wording and eliminate options that overstate AI as always better or fully replacing humans.
Quickest way: Use case, risk, control
When to use it: For conceptual multiple-choice questions with limited time.
- Name the use case in a few words.
- Ask what could go wrong: opacity, bias, drift, overfitting or imbalance.
- Pick the option that pairs that risk with the matching control.
- Reject absolute words such as always, eliminates or removes the need for.
Common mistakes in AI in Risk Management, Compliance and Surveillance
Using accuracy to judge a fraud model.
Accuracy feels like the natural score.
Fix: Fraud is rare, so a model flagging nothing can look highly accurate. Use precision, recall and false positive rate.
Treating better predictive power as sufficient for approval.
Performance is the visible metric.
Fix: Supervisors also require explainability, stability, fairness and documentation. Strong accuracy alone does not meet governance needs.
Confusing supervised and unsupervised learning.
Both detect anomalies.
Fix: Supervised needs labelled outcomes. Unsupervised finds patterns or outliers without labels, useful when new fraud types appear.
Assuming AI removes human responsibility.
Automation sounds complete.
Fix: Accountability stays with the firm. Human review, challenge and escalation remain required.
Ignoring training data limits.
Models seem objective.
Fix: Models inherit bias and gaps in their data and can fail when conditions shift. Monitor drift and retrain with validation.
Forgetting vendor and concentration risk in RegTech.
Focus is on the algorithm itself.
Fix: Outsourced AI creates third-party dependence. Banks need due diligence, exit plans and oversight.
Worked examples
Example 1
A surveillance model reviews 10,000 trades. 50 are genuine abuse cases. The model flags 100 trades, of which 40 are genuine. Calculate precision, recall and the false positive rate.
Show the solution
- TP = 40. FP = 100 − 40 = 60. FN = 50 − 40 = 10. TN = 10,000 − 40 − 60 − 10 = 9,890.
- Precision = 40 ÷ (40 + 60) = 40 ÷ 100 = 0.40.
- Recall = 40 ÷ (40 + 10) = 40 ÷ 50 = 0.80.
- FPR = 60 ÷ (60 + 9,890) = 60 ÷ 9,950 ≈ 0.0060.
Answer: Precision is 40%, recall is 80% and the false positive rate is about 0.6%. The model catches most abuse but 60% of alerts are false, so analyst workload is high.
Example 2
A bank's neural network PD model has far higher out-of-sample ranking power than its scorecard on 2019-2022 data, but the validator rejects it for use in lending decisions. Give two reasons and a remedy for each.
Show the solution
- Reason one: lack of explainability. The bank cannot easily give reasons for declines or show economic sense of drivers.
- Remedy: use interpretability tools, simpler challenger models or constrained features, and document driver behaviour.
- Reason two: training data covers a benign period, so the model may drift or fail in stress.
- Remedy: stress and out-of-time testing, ongoing monitoring, and retraining triggers.
- Add that human oversight and model governance should stay in place.
Answer: The validator is concerned with opacity and stability. Explainability tools and benchmarking address the first. Stress testing and monitoring address the second. Predictive power alone does not justify approval.
Exam tips
- Expect scenario questions pairing a use case with its main risk and control.
- Know precision, recall and false positive rate. Be ready to compute them from counts.
- Remember fraud and abuse are rare events, so imbalance and false positives matter.
- Watch for absolute wording. Answers saying AI eliminates risk or human review are usually wrong.
- Link to model risk guidance: validation, effective challenge and governance apply to ML models too.
Practice questions from Advances in Artificial Intelligence: Implications for Capital Markets Activities
- A bank compares a transparent logistic regression with a deep neural network for a trade surveillance alert system. The neural network has h…
- A firm adds an L1 (lasso) penalty to a regression-based machine learning model with 200 candidate predictors. Compared with an unpenalized m…
- A bank trains a machine-learning credit-trading model on ten years of market data in which volatility was persistently low. During a sudden …
- A trading desk deploys a deep neural network to generate short-term signals. Validators find that the model performs well in backtests but i…
- A trading firm uses a reinforcement-learning algorithm for execution. A risk manager worries that the algorithm could learn behaviour that r…
AI in Risk Management, Compliance and Surveillance: frequently asked questions
How is AI used in market abuse surveillance?
Firms use ML and NLP to scan trades, orders, emails and chats for patterns such as spoofing or insider dealing. Anomaly detection helps find new behaviours. Analysts then review alerts.
What are the limits of machine learning credit risk models?
They can be hard to explain, may overfit, and can inherit bias from data. They may also perform poorly when conditions differ from the training period. Validation and monitoring are needed.
What is RegTech?
RegTech uses technology, often AI, to help firms meet regulatory requirements. Examples are KYC and AML screening, transaction monitoring and automated reporting.
Why is accuracy a poor measure for fraud models?
Fraud is rare, so a model that never flags anything can still score high accuracy. Precision and recall show how well real cases are found.