FRM Exam Part II · Validating Bank Holding Companies' Value-at-Risk Models for Market Risk
Statistical Tests of VaR Accuracy and Independence
Updated 11 October 2026 · Fact-checked
VaR backtests check whether exceptions occur at the promised rate and without clustering. The Kupiec unconditional coverage test compares the exception count with the expected count, using a chi-square likelihood ratio with 1 degree of freedom. The Christoffersen independence test checks clustering (1 df). Their sum is the conditional coverage test (2 df). Reject when the statistic exceeds the critical value.
Understand Statistical Tests of VaR Accuracy and Independence
A VaR model at confidence level c says that a loss larger than VaR should happen with probability p = 1 − c. For a 99% one-day VaR, you expect an exception (a day when the loss exceeds VaR) on about 1% of days. Backtesting asks whether the observed exceptions are consistent with that claim.
There are two separate properties. Unconditional coverage asks: is the number of exceptions about right? Independence asks: are exceptions spread out, or do they bunch together? A model can pass the first and fail the second. For example, it may give the right count overall, but all exceptions arrive in one stressed fortnight. That means the model reacts too slowly to changing volatility.
The Kupiec test (proportion of failures) tests unconditional coverage. It is a likelihood ratio test comparing the hypothesised rate p with the observed rate N/T. The Christoffersen independence test looks at a two-state sequence (exception or not) and asks whether today's exception probability depends on yesterday's. The conditional coverage test combines both: the exception rate must be right and the exceptions must be independent.
The tests have limited power, meaning the ability to reject a bad model. At 99% confidence with about 250 days you expect only 2 or 3 exceptions. A model that is too low can easily look acceptable. Independence tests are weaker still, because clustering needs consecutive exceptions and those are rare. A pass is therefore not proof that the model is good. A failure is a stronger signal than a pass. Larger samples, or a lower confidence level such as 95%, give more power.
Key formulas to remember
- Exception indicator and expected rate
- I(t) = 1 if loss(t) > VaR(t), else 0; N = Σ I(t); p = 1 − c; expected exceptions = p × T
- T is the number of backtest days, N the observed exceptions. Under a correct model, N is Binomial(T, p).
- Kupiec unconditional coverage statistic
- LR_uc = −2 ln[(1 − p)^(T − N) × p^N] + 2 ln[(1 − N/T)^(T − N) × (N/T)^N]
- Under H0 (exception probability = p), LR_uc follows chi-square with 1 degree of freedom. Critical values: 3.84 at 5% and 6.63 at 1%. Reject H0 if LR_uc is larger.
- Transition probabilities
- π01 = n01 ÷ (n00 + n01); π11 = n11 ÷ (n10 + n11); π = (n01 + n11) ÷ (n00 + n01 + n10 + n11)
- nij = number of days in state j following state i (0 = no exception, 1 = exception). Independence means π01 = π11.
- Christoffersen independence statistic
- LR_ind = −2 ln[(1 − π)^(n00 + n10) × π^(n01 + n11)] + 2 ln[(1 − π01)^n00 × π01^n01 × (1 − π11)^n10 × π11^n11]
- Chi-square with 1 degree of freedom under H0 (independence). Same critical values as Kupiec: 3.84 at 5%.
- Conditional coverage statistic
- LR_cc = LR_uc + LR_ind
- Chi-square with 2 degrees of freedom. Critical values: 5.99 at 5% and 9.21 at 1%.
How to solve Statistical Tests of VaR Accuracy and Independence questions
Use this order for any question on formal VaR backtests. It works for calculation questions and for interpretation questions.
- 1Identify T (days), N (exceptions) and the VaR confidence level. Compute p = 1 − c and the expected exceptions p × T.
- 2Decide which property the question tests: count (unconditional coverage), clustering (independence), or both (conditional coverage).
- 3Pick the matching statistic and its degrees of freedom: Kupiec 1, independence 1, conditional coverage 2.
- 4Compute the likelihood ratio, or use the value given. Remember the conditional coverage statistic is the sum of the other two.
- 5Compare with the chi-square critical value at the stated significance level. Reject H0 only if the statistic is larger.
- 6State the conclusion in words: model rejected or not rejected, and which property failed.
- 7If asked about reliability, comment on power: few expected exceptions mean a pass is weak evidence, and the test can fail to catch a bad model.
Quickest way: Compare, sum, then read the chi-square table
When to use it: Use when the question gives LR statistics or counts and asks whether the model is rejected, under time pressure.
- Memorise the 5% critical values: 3.84 for 1 df and 5.99 for 2 df. Also know 6.63 and 9.21 for 1%.
- If one statistic is given for uncoverage and one for independence, add them for conditional coverage and use the 2 df value.
- For a rough check, compare N with p × T. If N is close to expected, Kupiec will not reject. If it is more than about double, expect a rejection.
- If exceptions appear in consecutive days, think independence failure even before calculating.
- Eliminate options that use the wrong degrees of freedom or reverse the reject rule.
Common mistakes in Statistical Tests of VaR Accuracy and Independence
Using the 1 df critical value (3.84) for the conditional coverage test.
All three tests are chi-square, so students reuse the familiar number.
Fix: Conditional coverage adds two components, so it has 2 degrees of freedom. Use 5.99 at 5%.
Saying a model that passes Kupiec has independent exceptions.
Students treat the correct count as proof of a correct model.
Fix: Kupiec ignores timing. Only the independence or conditional coverage test addresses clustering.
Treating non-rejection as proof that the VaR model is accurate.
Pass/fail language suggests certainty.
Fix: With few expected exceptions the tests have low power. Non-rejection means the data are not inconsistent with the model, nothing more.
Counting too many or too few transitions when building the 2 × 2 table.
Students count days instead of day-to-day pairs.
Fix: T days give T − 1 transitions. The four counts n00, n01, n10 and n11 must sum to T − 1.
Rejecting the model when the statistic is below the critical value.
Confusion with p-value logic.
Fix: Reject H0 only when the likelihood ratio is above the critical value. A larger statistic means a worse fit.
Mixing up π01 and π11 in the interpretation.
Both look similar, and the subscripts are read backwards.
Fix: First subscript is yesterday's state, second is today's. π11 is the chance of an exception today given one yesterday. Much larger than π01 signals clustering.
Worked examples
Example 1
A bank backtests a 99% one-day VaR over T = 250 days and records N = 7 exceptions. Using the Kupiec test, should the model be rejected at the 5% and the 1% significance levels? (Critical values: 3.84 and 6.63.)
Show the solution
- p = 1 − 0.99 = 0.01. Expected exceptions = 0.01 × 250 = 2.5. Observed N/T = 7 ÷ 250 = 0.028.
- Log-likelihood under H0: 243 × ln(0.99) + 7 × ln(0.01) = 243 × (−0.010050) + 7 × (−4.605170) = −2.442 − 32.236 = −34.678.
- Log-likelihood at the observed rate: 243 × ln(0.972) + 7 × ln(0.028) = 243 × (−0.028399) + 7 × (−3.575550) = −6.901 − 25.029 = −31.930.
- LR_uc = 2 × (−31.930 − (−34.678)) = 2 × 2.749 ≈ 5.50.
- Compare: 5.50 > 3.84, so reject at 5%. 5.50 < 6.63, so do not reject at 1%.
Answer: LR_uc ≈ 5.50. The model is rejected at the 5% level but not at the 1% level. Exceptions are too frequent (7 versus about 2.5 expected), but the evidence is not overwhelming.
Example 2
A separate 250-day backtest of a 99% one-day VaR gives these transition counts: n00 = 220, n01 = 12, n10 = 12, n11 = 5. Test independence at 5% (critical value 3.84) and interpret.
Show the solution
- Check the table: 220 + 12 + 12 + 5 = 249 = T − 1 transitions. Good. This backtest has 17 exceptions, since n01 + n11 = 17.
- π01 = 12 ÷ (220 + 12) = 12 ÷ 232 ≈ 0.0517. π11 = 5 ÷ (12 + 5) = 5 ÷ 17 ≈ 0.2941. π = 17 ÷ 249 ≈ 0.0683.
- Log-likelihood under independence: 232 × ln(0.9317) + 17 × ln(0.0683) ≈ 232 × (−0.07072) + 17 × (−2.6842) ≈ −16.407 − 45.632 = −62.039.
- Log-likelihood with separate probabilities: 220 × ln(0.94828) + 12 × ln(0.05172) + 12 × ln(0.70588) + 5 × ln(0.29412) ≈ −11.684 − 35.545 − 4.180 − 6.119 = −57.528.
- LR_ind = 2 × (−57.528 + 62.039) ≈ 9.02.
- 9.02 > 3.84, so reject independence (it also exceeds 6.63).
Answer: LR_ind ≈ 9.0, so independence is rejected. The chance of an exception after an exception (about 29%) is far above the chance after a quiet day (about 5%). Exceptions cluster, which suggests the VaR model is slow to adapt to changing volatility. The model also fails conditional coverage, because LR_cc = LR_uc + LR_ind, and this sum is far above the 5.99 critical value, whatever the Kupiec result.
Exam tips
- Know the three degrees of freedom by heart: 1, 1 and 2. Many wrong options differ only in this.
- Questions often give LR values and ask for the conclusion. Add the two for conditional coverage before comparing.
- When a question says exceptions occurred in consecutive days or in a single stress period, the answer is about independence, not coverage.
- Be ready to explain low power: at 99% over about a year you expect 2 to 3 exceptions, so a poor model can pass.
- When asked what to do after a rejection, link the failure type: too many exceptions suggests understated risk, clustering suggests slow volatility updating.
Practice questions from Validating Bank Holding Companies' Value-at-Risk Models for Market Risk
- A bank holding company wants to assess its internal VaR model by comparing its output with an independent measure. It runs the same trading …
- A bank backtests its one-day 99% VaR model over 250 trading days and records 9 exceptions. Under the Basel traffic-light framework, which zo…
- A validator runs a sensitivity analysis on a bank's VaR model by changing the lookback window from 250 to 500 days while holding all else co…
- During validation, a bank holding company is found to scale its one-day 99% VaR to a ten-day horizon by multiplying by the square root of te…
- A bank holding company's validation team reviews a trading desk's VaR model and finds that several illiquid corporate bonds are mapped to a …
Statistical Tests of VaR Accuracy and Independence: frequently asked questions
What is the difference between unconditional and conditional coverage in VaR backtesting?
Unconditional coverage only checks that the overall exception rate matches 1 − c. Conditional coverage checks the rate and also that exceptions are independent over time. Conditional coverage is the stricter test, and its statistic is the sum of the Kupiec and independence statistics.
How do I interpret the Kupiec likelihood ratio test?
Compare LR_uc with the chi-square critical value for 1 degree of freedom (3.84 at 5%). If it is larger, reject the hypothesis that the model's exception probability equals p. If it is smaller, you cannot reject, but this is not proof the model is accurate.
What does the Christoffersen independence test look for?
It tests whether the probability of an exception today depends on whether there was one yesterday. If π11 is much larger than π01, exceptions cluster. This usually signals a model that responds too slowly to changes in market volatility.
Why do VaR backtests have low power with small samples?
At high confidence levels, exceptions are rare. A 250-day sample at 99% has only about 2.5 expected exceptions, so the difference between a good and a bad model is small in count terms. The tests then often fail to reject a flawed model.