Skip to content

FRM Exam Part II · Beyond Exceedance-Based Backtesting of Value-at-Risk Models

Limitations of Exceedance-Based VaR Backtesting

Updated 11 October 2026 · Fact-checked

Exceedance-based backtesting counts days when the loss beats VaR and compares the count with the expected number. It ignores how large the losses are, when they occur and whether they cluster. Tests like Kupiec also have low power in small samples, so wrong models often pass. Fix this by adding independence tests, size-aware measures and longer samples.

Understand Limitations of Exceedance-Based VaR Backtesting

A VaR exception (exceedance) is a day when the actual loss is greater than the VaR forecast. At 99% confidence, a good model should produce exceptions on about 1% of days. Backtesting counts the exceptions and asks whether the count is close enough to that target.

The Kupiec test (unconditional coverage) and the Basel traffic light approach both work this way. Each day becomes a simple yes/no: exception or not. That is the source of all the limitations. Once you reduce a day to a yes/no, you throw away information.

Three things are lost. Size: an exception of ₹1 lakh over VaR and one of ₹100 crore over VaR count the same. A model with severe tail losses can look as good as one with mild ones. Timing and clustering: five exceptions spread over a year are very different from five in one week. Clusters signal that the model reacts too slowly to volatility. The Kupiec test only checks the total, so it cannot see clusters. The Christoffersen independence and conditional coverage tests were built to check this. Direction and context: the count says nothing about why exceptions happened, such as bad data, wrong mapping or a regime change.

The second problem is low power. Power is the chance a test rejects a model that is truly wrong. At 99% VaR with 250 days, you expect only 2.5 exceptions. The gap between a good model (1% true rate) and a bad one (say 2% true rate) is only a few observations, which random noise can hide. So the test often fails to reject a bad model (Type II error). Power is better at lower confidence levels such as 95%, and with longer samples.

The practical lesson: passing an exception test is necessary, not sufficient. Use it with independence tests, tail-loss measures (such as expected shortfall backtests or PIT-based tests), benchmarking and P&L analysis.

Key formulas to remember

Expected number of exceptions
E[N] = T × p, where p = 1 − confidence level
T is the number of days. At 99% over 250 days, E[N] = 2.5.
Unconditional coverage (Kupiec) likelihood ratio
LR_uc = −2 ln[(1 − p)^(T−N) × p^N] + 2 ln[(1 − N/T)^(T−N) × (N/T)^N]
Compared with a chi-square distribution with 1 degree of freedom. The 5% critical value is 3.84. It tests only the frequency of exceptions.
Independence idea (Christoffersen)
LR_cc = LR_uc + LR_ind, chi-square with 2 degrees of freedom
LR_ind tests whether today's exception depends on yesterday's. Conditional coverage tests frequency and independence together. The 5% critical value is 5.99.
Basel traffic light zones (250 days, 99% VaR)
Green: 0–4 exceptions; Yellow: 5–9; Red: 10 or more
Yellow-zone exceptions raise the capital multiplier. Red usually means the model is presumed flawed.

How to solve Limitations of Exceedance-Based VaR Backtesting questions

For any question on limits of exception-based backtesting, identify what information the test uses and what it ignores, then match the weakness to the fix.

  1. 1Read what is given: confidence level, sample length, exception count and whether timing or size is described.
  2. 2Compute the expected exceptions: T × (1 − confidence level).
  3. 3Compare the actual count with the expected count and with the test or zone thresholds if the question asks for a result.
  4. 4Ask what the test ignores: size of losses, clustering in time, or both.
  5. 5Check sample size. If T is small and the confidence level is high, flag low power and the risk of a Type II error.
  6. 6Pick the remedy that matches the weakness: independence or conditional coverage test for clustering; tail-loss or ES-based measures for size; longer samples or lower confidence levels for power.
  7. 7State the conclusion carefully: passing means no evidence against the model, not proof that it is accurate.

Quickest way: Weakness-to-fix matching

When to use it: Use for conceptual MCQs where options list different limitations or remedies.

  1. Total count looks fine but exceptions are bunched: the problem is independence. Answer: Christoffersen.
  2. Count is fine but losses beyond VaR are huge: the problem is size. Answer: tail or ES-based measures.
  3. Few observations, 99% level: the problem is power. Answer: Type II error; use longer sample or additional tests.
  4. Question says the model passed: eliminate any option claiming it is proven correct.

Common mistakes in Limitations of Exceedance-Based VaR Backtesting

  • Saying a model that passes the Kupiec test is accurate.

    Passing sounds like confirmation.

    Fix: Failing to reject only means the data give no strong evidence against the model. With low power, bad models can pass.

  • Believing Kupiec detects clustering of exceptions.

    Both Kupiec and Christoffersen are called coverage tests.

    Fix: Kupiec checks only frequency (unconditional coverage). Clustering needs the independence test or the combined conditional coverage test.

  • Thinking larger exceptions count more in the traffic light approach.

    Intuition says bigger losses should matter more.

    Fix: Each exception counts once regardless of size. That is a core limitation.

  • Confusing Type I and Type II errors when discussing low power.

    Both terms sound alike.

    Fix: Low power means a high chance of a Type II error: failing to reject a wrong model. Type I is rejecting a correct model.

  • Using the wrong degrees of freedom: 1 for conditional coverage.

    Students reuse the Kupiec value.

    Fix: Unconditional coverage and independence each use 1 degree of freedom. Conditional coverage uses 2.

  • Claiming power is greater at 99% than at 95%.

    A higher confidence level feels more rigorous.

    Fix: At 99% there are very few expected exceptions, so power is lower. Backtesting at 95% has more observations in the tail of interest.

Worked examples

Example 1

A bank backtests its one-day 99% VaR over 250 trading days and records 4 exceptions. All 4 occurred in the same week, and each loss was more than three times the VaR. Under the Basel traffic light approach, what zone is the model in, and what does this reveal about the approach?

Show the solution
  1. Expected exceptions = 250 × 0.01 = 2.5.
  2. The count of 4 falls in the green zone (0–4 exceptions).
  3. The approach counts only the number of exceptions, so the bunching in one week is invisible to it.
  4. It also ignores that each loss was far above VaR, since size does not matter in the count.
  5. Both facts suggest the model may be slow to react to volatility and understate tail risk.

Answer: The model is in the green zone, yet the clustering and the size of the losses point to a weakness. This shows that exception counting ignores timing and size.

Example 2

A risk manager tests a 99% VaR model over 250 days. The true exception probability is actually 2%, double the stated 1%. Observed exceptions are 5. Explain whether the sample is likely to reliably detect the problem and what this says about the test.

Show the solution
  1. Expected exceptions under the model = 250 × 0.01 = 2.5.
  2. Expected exceptions under the true rate = 250 × 0.02 = 5.
  3. The difference between the two cases is only 2.5 exceptions.
  4. Random variation in counts of this size is comparable to that gap, so a count of 5 can arise often under either rate.
  5. The Kupiec test is therefore unlikely to reject reliably. The probability of a Type II error is high, which means low power.
  6. Remedies: use a longer sample, test at a lower confidence level, or add other tests.

Answer: No. A model with double the intended exception rate can often pass over 250 days, because the expected counts of 2.5 and 5 are too close relative to noise. This is the low-power limitation.

Exam tips

  • Know the three losses of information: size, timing and clustering. Questions often ask for one that a given test misses.
  • Link each limitation to its remedy: clustering to Christoffersen, size to ES or PIT-based tests, small samples to longer data.
  • Be exact on the language: 'fail to reject' is not 'accept'.
  • Remember the zones for 250 days at 99%: green 0–4, yellow 5–9, red 10 or more.
  • If numbers are given, compute expected exceptions first. It anchors every other step.

Practice questions from Beyond Exceedance-Based Backtesting of Value-at-Risk Models

Limitations of Exceedance-Based VaR Backtesting in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Limitations of Exceedance-Based VaR Backtesting: frequently asked questions

Why is exception counting not enough for VaR backtesting?

It reduces each day to exception or no exception. It ignores how big the loss was, when it happened and whether exceptions bunch together. Two models with the same count can have very different risk.

Why does the Kupiec test have low power in small samples?

At 99% confidence only a few exceptions are expected, for example 2.5 in 250 days. Random noise in such small counts hides the difference between a good and a bad model, so wrong models often pass.

What is the difference between unconditional and conditional coverage?

Unconditional coverage checks only whether the share of exceptions matches the expected rate. Conditional coverage also checks independence, meaning exceptions should not cluster. The conditional test combines both and uses 2 degrees of freedom.

How can I overcome these limitations?

Add an independence test, use backtests that look at the size of tail losses such as expected shortfall or PIT-based tests, use longer samples, and combine with benchmarking and P&L analysis.