FRM Exam Part II · Validating Bank Holding Companies' Value-at-Risk Models for Market Risk
Limitations of VaR Validation and Supervisory Findings
Updated 11 October 2026 · Fact-checked
VaR validation limits are the reasons backtesting alone cannot prove a model is good. Tests have low power at 99% confidence, ignore tail size, and can be distorted by changing portfolios. Supervisors therefore expect extra tools: benchmarking, sensitivity analysis, stress tests, clean P&L and documented, independent challenge.
Understand Limitations of VaR Validation and Supervisory Findings
Start with what validation does. A bank holding company compares its VaR with actual trading results and asks whether the model is reliable. The usual evidence is the count of exceptions, which are days when the loss is larger than VaR. At 99% VaR you expect about 1% of days, or roughly 2.5 in 250, to be exceptions.
The first weakness is low test power. With about 250 observations a year and only a few expected exceptions, a statistical test often cannot tell a good model from a bad one. A model that understates risk can still show an acceptable count. This is a Type II error (accepting a bad model). Rejecting a good model is a Type I error. Few data points make both errors hard to control.
The second weakness is portfolio change. VaR is computed on today's positions, but the backtest uses results from a portfolio that moves during the day. Intraday trading, fees and reserves enter actual P&L but are not in the VaR. Positions also change over the test window, so the history is not a single stable portfolio. Exceptions can then come from trading and not from model error. Using hypothetical (clean) P&L, which holds positions fixed, helps separate the two.
The third weakness is tail risk. An exception count says how often VaR is breached, not by how much. A model with few but huge breaches can pass. VaR also says nothing about losses beyond the threshold. Expected shortfall, stress tests and EVT address this, but they are harder to backtest.
Supervisory reviews of bank holding companies typically find: backtesting that relies only on exception counts, weak documentation, poor data and proxy handling, limited use of benchmark models, unclear explanations of exceptions, and validation that is not independent or does not challenge assumptions. The recommendation is a broader programme: several tests, several confidence levels and horizons, sub-portfolio backtests, benchmarking, stress testing, and ongoing monitoring with clear governance.
Key formulas to remember
- Expected exceptions
- Expected exceptions = N × (1 − c)
- N is the number of days and c the VaR confidence level. At 99% over 250 days, expect 2.5.
- Exception rate
- Exception rate = x ÷ N
- Compare with 1 − c. A rate near the expected value does not prove the model is right, because test power is low.
- Type I and Type II errors
- Type I = reject a correct model; Type II = accept an incorrect model
- Low power at 99% means a high chance of Type II error.
- Core limitations
- Low power + portfolio change + tail blindness
- A memory aid for the three main weaknesses of exceedance-based validation.
How to solve Limitations of VaR Validation and Supervisory Findings questions
Use this approach for any question on validation weaknesses or supervisory findings.
- 1Identify what the question describes: a count of exceptions, a P&L definition, a portfolio change or a tail loss.
- 2Name the limitation it links to: low power, P&L contamination, changing positions, or tail size not captured.
- 3State the consequence in words, such as a bad model being accepted or exceptions being misattributed.
- 4Check the numbers if given. Compute expected exceptions as N × (1 − c) and compare.
- 5Pick the remedy that fits: clean P&L, sub-portfolio tests, lower confidence level, benchmark model, stress test or expected shortfall.
- 6Match the remedy to a supervisory recommendation, such as independent validation or better documentation.
- 7Eliminate options that say a passing backtest proves accuracy, or that treat one test as sufficient.
Quickest way: Weakness-to-remedy match
When to use it: Use when the question is a short conceptual MCQ and time is tight.
- Underline the key phrase: few exceptions, intraday trading, size of loss, or one test only.
- Map it: few data to low power; intraday or fees to dirty P&L; size to tail blindness; one test to need for benchmarking and stress tests.
- Choose the option that adds information, not one that claims a pass settles the issue.
- Reject absolutes such as always, proves or guarantees.
Common mistakes in Limitations of VaR Validation and Supervisory Findings
Believing a passing backtest proves the model is accurate.
Pass or fail results feel conclusive.
Fix: Remember low power: a bad model can pass. Passing only means the evidence did not reject it.
Confusing Type I and Type II errors.
Both words sound similar and the null is the model being good.
Fix: Type I rejects a good model. Type II accepts a bad one. Low power raises Type II.
Treating exception count as a measure of tail loss size.
Students link more exceptions to bigger losses.
Fix: Counts show frequency only. Size needs expected shortfall, stress tests or exception severity analysis.
Using actual P&L as if it were purely the model's result.
Actual P&L is the easy number to obtain.
Fix: Actual P&L includes intraday trades, fees and reserves. Hypothetical P&L on fixed positions isolates model error.
Assuming a higher confidence level makes backtesting stronger.
Higher confidence sounds more rigorous.
Fix: At 99% there are fewer expected exceptions, so power falls. A lower level such as 95% gives more observations, though it says less about the far tail.
Proposing a single test as the fix for weak validation.
Students look for one decisive method.
Fix: Supervisors favour a combination: backtests, benchmarking, sensitivity, stress testing and independent review.
Worked examples
Example 1
A bank backtests a 99% one-day VaR over 250 trading days and records 4 exceptions. The risk manager says the model is therefore proven accurate. Give the expected number of exceptions and explain why the claim is weak.
Show the solution
- Expected exceptions = 250 × (1 − 0.99) = 250 × 0.01 = 2.5.
- Observed is 4, which is above 2.5 but not by much.
- With so few expected exceptions, a statistical test has low power. A model that understates risk could produce a similar count.
- Also, the count says nothing about how large the losses on those 4 days were.
Answer: Expected exceptions are 2.5. The claim is weak because of low test power and because exception counts ignore the size of tail losses.
Example 2
A bank's backtest uses actual daily P&L. A desk made large intraday trades and earned fees on several days with exceptions. Which is the most appropriate response? A) Raise the VaR confidence level to 99.9%. B) Backtest using hypothetical P&L on fixed positions as well. C) Drop backtesting. D) Use only a longer observation window.
Show the solution
- Identify the issue: actual P&L includes intraday trading and fees, which VaR does not model.
- This contaminates the link between exceptions and model quality.
- Hypothetical P&L holds end-of-day positions fixed, so it isolates model error.
- A raises the confidence level, which reduces expected exceptions and lowers power, so it does not fix contamination.
- C removes a key tool. D does not address P&L definition.
Answer: B. Add a backtest on hypothetical (clean) P&L, which separates model error from trading and fee effects.
Exam tips
- Expect scenario questions: map the described problem to low power, P&L contamination, portfolio change or tail blindness.
- Be ready to compute expected exceptions as N × (1 − c) before judging results.
- Watch for options that say a pass proves the model; they are usually wrong.
- Know the remedies: clean P&L, sub-portfolio tests, benchmarking, stress testing, expected shortfall and independent validation.
- Keep Type I and Type II straight; low power means more Type II errors.
Practice questions from Validating Bank Holding Companies' Value-at-Risk Models for Market Risk
- A bank's 99% one-day VaR model is backtested over 250 trading days and produces 9 exceptions. The validator is deciding how to interpret thi…
- A bank's trading desk shows more exceptions when VaR is compared with actual (clean-plus-dirty) P&L than with hypothetical P&L. Which is the…
- A validator benchmarks a bank's 99% one-day VaR of USD 12.0 million against a vendor model giving USD 15.0 million. Over 250 days, the bank'…
- A bank holding company's validation team reviews a VaR model that uses a one-year window of daily historical data and finds that the data fo…
- A bank's validation group finds that backtesting exceptions are clustered in a few consecutive weeks, although the total count of 4 over 250…
Limitations of VaR Validation and Supervisory Findings: frequently asked questions
Why is VaR backtesting low in power?
At 99% confidence, a year of data gives only about 2 to 3 expected exceptions. With so few events, the test struggles to separate a good model from a flawed one. This makes it easy to accept a bad model.
Why does tail risk limit VaR validation?
VaR only marks a threshold, and exception counts only say how often it is crossed. They do not show how large the losses beyond it are. Expected shortfall and stress tests are needed to see tail size.
What are typical supervisory findings on bank VaR validation?
Common findings are reliance on exception counts alone, weak documentation, poor data and proxy treatment, limited benchmarking and unexplained exceptions. Supervisors recommend broader testing, clean P&L, stress testing and independent challenge.
How does portfolio change affect backtesting?
VaR uses current positions, but actual results come from a portfolio that changes intraday and over time. Exceptions may reflect trading, fees or reserves instead of model error. Hypothetical P&L on fixed positions reduces this problem.