FRM Exam Part II · Backtesting VaR
Type I and Type II Errors in VaR Backtesting
Updated 11 October 2026 · Fact-checked
A type I error rejects a VaR model that is actually correct. A type II error accepts a model that is actually wrong. Backtests use exception counts, so they cannot cut both errors at once. Higher VaR confidence levels and short samples give few exceptions and weak power, so type II errors rise.
Understand Type I and Type II Errors in Backtesting
A VaR backtest compares the number of days when the loss exceeded VaR (exceptions) with the number the model predicts. If the model is correct at 99%, you expect exceptions on 1% of days. Real samples are noisy, so the count will rarely match exactly.
Because of this noise, you need a rule to decide when the count is too high or too low. That rule creates two ways to be wrong. A type I error means rejecting a correct model. A type II error means failing to reject an incorrect model. In FRM language, the type I error rate is the significance level of the test. Power is the probability of rejecting a model that is wrong. Power = 1 − P(type II error).
There is a trade-off. If you set a tight cutoff, you reject more good models (more type I errors) but catch more bad ones (fewer type II errors). If you loosen the cutoff, the reverse happens. For a fixed sample, you cannot shrink both together.
Two things drive power. First, the VaR confidence level. At 99%, the expected exceptions in 250 days are only 2.5. A model that truly has a 2% exception rate gives about 5 exceptions. That is hard to separate from luck. At 95%, the expected count is 12.5, so differences are easier to see. Second, sample size. More observations narrow the spread of the exception rate and raise power. Backtesting 99% VaR with one year of data has low power. Using a longer window helps, but only if the model and portfolio stayed stable.
The practical lesson: failing to reject does not prove the model is good. It may just mean the test is weak. The Basel traffic light approach reflects this trade-off, with zones that balance the cost of wrongly penalising a bank against the cost of accepting a weak model.
Key formulas to remember
- Type I error
- P(reject H0 | model is correct) = α
- The significance level of the test. Rejecting a good model.
- Type II error
- β = P(do not reject H0 | model is wrong)
- Accepting a bad model.
- Power
- Power = 1 − β
- Probability of rejecting a wrong model. Higher is better.
- Expected exceptions
- E[x] = T × p, where p = 1 − confidence level
- T is the number of backtest days.
- Std deviation of exception count
- σ = √(T × p × (1 − p))
- Under the binomial model with independent exceptions. Used to judge how far a count is from expected.
How to solve Type I and Type II Errors in Backtesting questions
Use this order for any question on errors in backtesting.
- 1Identify the null hypothesis: the model is correct, so the true exception rate is p = 1 − confidence level.
- 2Decide what the question describes: rejecting a correct model (type I) or accepting a wrong model (type II).
- 3Compute the expected exceptions T × p and compare with the observed count.
- 4Check the confidence level and sample size. Few expected exceptions means low power.
- 5Apply the trade-off: a stricter rejection cutoff lowers type II errors and raises type I errors, and the reverse.
- 6State the interpretation: failure to reject does not prove the model is accurate.
- 7Match your result to the answer option, checking direction words such as increase or decrease.
Quickest way: Two-question shortcut
When to use it: Use when an MCQ asks which error is more likely or how a change affects the errors.
- Ask: is the model actually good or bad in the scenario? Good and rejected means type I. Bad and accepted means type II.
- Ask: did the change add information (more data) or just move the cutoff? More data raises power. Moving the cutoff trades one error for the other.
- Remember: higher VaR confidence means fewer expected exceptions, so lower power.
Common mistakes in Type I and Type II Errors in Backtesting
Swapping type I and type II errors.
The labels are easy to mix up under time pressure.
Fix: Type I is about the null being true: rejecting a good model. Type II is accepting a bad one.
Saying a model that passes is proven accurate.
Passing feels like confirmation.
Fix: A pass only means no rejection. With low power, bad models also pass often.
Thinking a higher confidence level makes backtests more powerful.
Higher confidence sounds more rigorous.
Fix: Higher confidence means fewer expected exceptions, so differences from the true rate are harder to detect.
Believing both errors can be reduced by changing the cutoff.
Students treat the cutoff as a free choice.
Fix: At fixed sample size, moving the cutoff lowers one error and raises the other. Only more data helps both.
Confusing the significance level with the VaR confidence level.
Both use percentages such as 1% and 5%.
Fix: VaR confidence sets the expected exception rate. The significance level is the type I error rate of the test.
Worked examples
Example 1
A bank backtests 99% one-day VaR over 250 days. The true exception rate of the model is actually 2%. A test rejects the model only if there are 7 or more exceptions. Which statement is correct?
Show the solution
- Expected exceptions under the null: 250 × 0.01 = 2.5.
- Expected exceptions under the true rate: 250 × 0.02 = 5.
- The true mean of 5 is below the rejection cutoff of 7, so rejection is not the typical outcome.
- The model is wrong, so failing to reject is a type II error.
- Power is low because the true count of 5 sits close to the cutoff of 7 given the noise.
Answer: The test will often fail to reject the wrong model. That is a type II error, and power is low.
Example 2
A risk team backtests a correct 95% VaR model. Over 500 days they expect 25 exceptions. The model is rejected after a sample showing 38 exceptions. Standard deviation of the count is √(500 × 0.05 × 0.95). Classify the outcome and explain whether more data would help.
Show the solution
- Standard deviation = √(23.75) ≈ 4.87.
- Excess over expected: 38 − 25 = 13, which is about 2.7 standard deviations.
- The model is stated as correct and was rejected, so this is a type I error.
- This is a chance outcome of the test, at a rate set by the significance level.
- More data narrows the spread of the exception rate and improves the test, so a longer sample would help confirm or reverse the decision.
Answer: This is a type I error: a correct model was rejected. A longer sample would reduce the chance of such a misleading result.
Exam tips
- Read the scenario for whether the model is truly good or bad. That fixes the error type.
- Expect conceptual MCQs on how confidence level and sample size change power. Higher confidence lowers power; more data raises it.
- If an option says a passed backtest proves accuracy, it is almost certainly wrong.
- Link the trade-off to the Basel traffic light zones, which balance wrongly penalising a bank against missing a weak model.
Practice questions from Backtesting VaR
- A risk manager backtests a 99% one-day VaR model over 250 days and wants to test both the correct number of exceptions and whether exception…
- A bank backtests its 95% one-day VaR over 500 days and records 35 exceptions. Using the normal approximation to the binomial, what is the z-…
- A bank backtests a 99% one-day VaR over 250 days. Assume exceptions are independent with probability 1% each. The bank has 5 exceptions and …
- A bank's 99% one-day VaR model has 8 exceptions over 250 days in the yellow zone. Investigation shows 3 of them occurred because traders' po…
- A bank backtests its 99% one-day VaR over 500 days using hypothetical (clean) P&L, where positions are held fixed, and separately using actu…
Type I and Type II Errors in Backtesting in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Type I and Type II Errors in Backtesting: frequently asked questions
What is the difference between type I and type II error in VaR backtesting?
A type I error rejects a VaR model that is correct. A type II error accepts a model that is wrong. The first harms a bank with a good model; the second leaves a weak model in use.
Why is the power of a backtest low at high VaR confidence levels?
At 99%, only about 1% of days are expected to be exceptions. With a typical sample you see just a few, so a wrong model looks similar to a correct one by chance.
How can I reduce both errors at once?
Use more data. With a fixed sample, a tighter cutoff reduces type II errors but increases type I errors. A longer sample, or a better test, raises power without raising type I errors.
Does passing a backtest mean the VaR model is accurate?
No. It means the test did not find enough evidence to reject. If power is low, many inaccurate models would pass too.