FRM Exam Part II · Beyond Exceedance-Based Backtesting of Value-at-Risk Models
Kolmogorov-Smirnov, Anderson-Darling and Berkowitz Tests for VaR Backtesting
Updated 11 October 2026 · Fact-checked
These tests check whether a model's probability integral transform (PIT) values behave like independent uniform(0,1) draws. Kolmogorov-Smirnov and Anderson-Darling test uniformity; Anderson-Darling weights the tails more. Berkowitz converts PITs to normal scores and uses a likelihood-ratio test of zero mean, unit variance and no autocorrelation.
Understand Distribution Tests: Kolmogorov-Smirnov, Anderson-Darling, Berkowitz
A VaR exception count uses only one point of the forecast distribution. Distribution backtests use the whole forecast. Each day the model gives a forecast distribution F_t for the return. You take the realised return r_t and compute the PIT value u_t = F_t(r_t). If the model is correct, the u_t are independent and uniform on (0,1).
That gives two things to test: uniformity (is the distribution of u_t flat?) and independence (does u_t tell you anything about u_t−1?). If u_t cluster near 0, realised returns fall in the lower tail more often than forecast, so losses are larger than the model forecast. If they cluster near 1, realised returns fall in the upper tail more often than forecast, so gains are larger. A U-shape, with mass near both 0 and 1, means the forecast tails are too thin and risk is understated. If they cluster near 0.5, the model overstates risk.
The Kolmogorov-Smirnov (KS) test compares the empirical CDF of the PIT values with the uniform CDF and uses the largest gap. It weights all gaps equally. But the variance of the empirical CDF under the null is greatest near 0.5, so the largest gap is most likely to arise there by chance. This makes KS relatively less sensitive to tail deviations than AD, and it can miss tail errors. The Anderson-Darling (AD) test also compares the two CDFs, but it weights each squared gap by the inverse of u(1 − u), as in A² = n ∫ (F_n(u) − u)² ÷ [u(1 − u)] du. This puts far more weight on the tails, which is where VaR matters. Both tests assume the PITs are independent, so neither tests independence.
The Berkowitz test first applies the inverse normal CDF: z_t = Φ⁻¹(u_t). Under a correct model, z_t are iid N(0,1). You fit z_t − μ = ρ(z_t−1 − μ) + ε_t with ε_t ~ N(0, σ²) and compare its likelihood with the restricted model μ = 0, σ² = 1, ρ = 0. In the standard test the likelihood-ratio statistic is chi-square with 3 degrees of freedom. Because it tests mean, variance and autocorrelation together, it checks independence as well as distribution. A tail-focused version of the test also exists. It looks only at the tail region, has a different likelihood and its own test setup, so the 3 degrees of freedom belong to the standard test.
Key formulas to remember
- PIT value
- u_t = F_t(r_t)
- F_t is the model's forecast CDF for day t. Under a correct model the u_t are iid U(0,1).
- Kolmogorov-Smirnov statistic
- D = max over u of |F_n(u) − u|
- F_n is the empirical CDF of the PITs. With sorted values u(1) ≤ … ≤ u(n), D = max of [i/n − u(i)] and [u(i) − (i−1)/n]. A large D means rejection. Every gap has equal weight.
- Anderson-Darling statistic
- A² = −n − (1/n) Σ (2i − 1) × [ln u(i) + ln(1 − u(n+1−i))]
- Sum over i = 1 to n on sorted PITs. This is the computational form of A² = n ∫ (F_n(u) − u)² ÷ [u(1 − u)] du, which is where the weight 1/[u(1 − u)] comes from. Tail-sensitive. Large A² means rejection. Critical values come from tables.
- Berkowitz transform
- z_t = Φ⁻¹(u_t)
- Under the null, z_t are iid N(0,1).
- Berkowitz alternative model
- z_t − μ = ρ(z_t−1 − μ) + ε_t, ε_t ~ N(0, σ²)
- Null hypothesis: μ = 0, σ² = 1, ρ = 0.
- Berkowitz LR statistic
- LR = −2 × [L(0, 1, 0) − L(μ̂, σ̂², ρ̂)] ~ χ²(3)
- L is the log-likelihood. For the standard Berkowitz test, three restrictions give 3 degrees of freedom. The 5% critical value is 7.815. The tail version has a different likelihood and its own test setup.
How to solve Distribution Tests: Kolmogorov-Smirnov, Anderson-Darling, Berkowitz questions
Use this order for any question on PIT-based distribution tests.
- 1Identify the forecast distribution and compute or read the PIT values u_t = F_t(r_t). State the null: PITs are iid U(0,1).
- 2Decide what the question tests: uniformity only (KS or AD) or uniformity plus independence (Berkowitz).
- 3For KS, sort the PITs and find the largest gap between the empirical CDF and the uniform line, checking both sides of each step. For AD, note that it weights the tails.
- 4For Berkowitz, transform with Φ⁻¹, then compare restricted and unrestricted log-likelihoods and compute LR = 2 × (L unrestricted − L restricted).
- 5Compare the statistic with the critical value. For Berkowitz use χ² with 3 degrees of freedom. Reject if the statistic exceeds it.
- 6Interpret in risk terms: where do the PITs cluster, what does that say about tail or variance forecasts, and what should be done with the model?
Quickest way: Pick the test from the wording
When to use it: Use when the question asks which test to use or what a result implies, and no long calculation is needed.
- Tail errors or tail weighting mentioned: choose Anderson-Darling over KS.
- Autocorrelation, mean, variance or likelihood ratio mentioned: choose Berkowitz, with 3 degrees of freedom.
- Only the shape of the distribution mentioned: KS or AD, and remember they ignore independence.
- For a Berkowitz LR question, double the log-likelihood difference and compare with 7.815 at 5%, or 11.345 at 1%.
- Reject means the model is misspecified. Do not reject means the data gave no evidence against it.
Common mistakes in Distribution Tests: Kolmogorov-Smirnov, Anderson-Darling, Berkowitz
Using KS or AD to conclude the PITs are independent.
The tests are called distribution tests and students assume they cover everything about the PIT series.
Fix: KS and AD assume independence and test only uniformity. Use Berkowitz or a separate autocorrelation test for independence.
Applying the Berkowitz likelihood to raw PIT values.
Students forget the transformation step.
Fix: Always convert with z_t = Φ⁻¹(u_t) first. The test is built on normal scores.
Using 2 degrees of freedom for the Berkowitz LR test.
Students count only mean and variance, or only the two obvious parameters.
Fix: The null of the standard test fixes three parameters: μ = 0, σ² = 1 and ρ = 0. Use χ²(3). The tail version has its own setup.
Saying KS is as good as AD for tail risk.
Both use the gap between two CDFs, so they look alike.
Fix: KS weights every gap equally, but the empirical CDF has the most variance near 0.5, so KS is relatively less sensitive to tail deviations. AD weights by 1/[u(1 − u)], so it is more sensitive in the tails.
Computing the KS gap on one side of each step only.
The empirical CDF jumps at each observation, and students compare only after or only before the jump.
Fix: Check both i/n − u(i) and u(i) − (i−1)/n for every sorted value and take the largest.
Reading a failure to reject as proof the model is correct.
Students treat the test like a pass mark.
Fix: Failure to reject means insufficient evidence against the model. With few observations the tests have low power.
Worked examples
Example 1
A bank backtests a model on five daily PIT values: 0.10, 0.25, 0.40, 0.70, 0.90. Compute the KS statistic and say what a large value would imply.
Show the solution
- The values are already sorted, with n = 5.
- Compute i/n − u(i): 0.2 − 0.10 = 0.10; 0.4 − 0.25 = 0.15; 0.6 − 0.40 = 0.20; 0.8 − 0.70 = 0.10; 1.0 − 0.90 = 0.10. The largest is 0.20.
- Compute u(i) − (i−1)/n: 0.10 − 0 = 0.10; 0.25 − 0.2 = 0.05; 0.40 − 0.4 = 0; 0.70 − 0.6 = 0.10; 0.90 − 0.8 = 0.10. The largest is 0.10.
- D is the larger of the two: D = 0.20.
- Compare D with a KS critical value for n = 5 from a table. A D above that value would reject uniformity. With so few observations the test has little power.
Answer: D = 0.20. A D above the critical value would mean the PITs are not uniform, so the forecast distribution is misspecified.
Example 2
Suppose a Berkowitz test on a VaR model's PIT-based normal scores gives these hypothetical log-likelihoods: −346.5 under the restricted model (μ = 0, σ² = 1, ρ = 0) and −340.2 under the unrestricted model. The figures are illustrative, not from a real fit. Test at 5% and 1% and interpret.
Show the solution
- LR = 2 × (L unrestricted − L restricted) = 2 × (−340.2 − (−346.5)).
- −340.2 + 346.5 = 6.3, so LR = 12.6.
- The test has 3 restrictions, so the statistic is χ²(3). The 5% critical value is 7.815 and the 1% critical value is 11.345.
- 12.6 exceeds both critical values, so reject the null at 5% and at 1%.
- Interpretation: the normal scores are not iid N(0,1). The mean, variance or autocorrelation differs from the null. The model forecasts are misspecified and need review, for example the volatility estimate or the return distribution assumption.
Answer: LR = 12.6, which is above 7.815 and 11.345. Reject the null at both levels. The normal scores are not iid N(0,1), so the PIT series is not iid uniform.
Exam tips
- Know which test checks what. KS and AD test uniformity only. Berkowitz tests distribution and independence jointly.
- Expect questions on why AD beats KS for VaR: tail weighting. Do not claim AD tests independence.
- Memorise the Berkowitz setup: normal transform, AR(1) alternative, LR with 3 degrees of freedom.
- In interpretation questions, link PIT clustering to model errors. PITs near 0 mean returns fall in the lower tail more often than forecast. U-shaped PITs, with mass near both 0 and 1, suggest tails are too thin and risk is understated. PITs piled near 0.5 suggest risk is overestimated.
- Remember that a tail-focused version of the Berkowitz test exists and looks at the tail region, which is relevant for VaR.
Practice questions from Beyond Exceedance-Based Backtesting of Value-at-Risk Models
- An analyst transforms PIT values u_t into z_t = Φ^-1(u_t) and wants to test the joint hypothesis of correct distribution and independence us…
- A bank backtests its 97.5% one-day VaR and ES. Over a test sample, VaR was exceeded on five days. The realised losses on those days exceeded…
- A risk manager has PIT values u_t from a VaR model and wants to test the full distribution rather than just tail exceedances. She transforms…
- A risk manager backtests a 99% one-day VaR model over 250 trading days and finds 2 exceedances, a result that passes the standard traffic-li…
- A bank has a 99% VaR model that produced zero exceedances in 250 days. Which interpretation is most consistent with the limitations of excee…
Distribution Tests: Kolmogorov-Smirnov, Anderson-Darling, Berkowitz: frequently asked questions
What is the difference between Anderson-Darling and Kolmogorov-Smirnov in backtesting?
Both compare the empirical CDF of the PIT values with the uniform CDF. KS uses the single largest gap and weights all gaps equally, so it is relatively less sensitive to tail deviations. AD weights squared gaps by the inverse of u(1 − u), so it is more sensitive in the tails.
How do you apply the KS test to VaR PIT values?
Compute each day's PIT from the forecast distribution, sort the values and find the largest distance between the empirical CDF and the uniform line. Compare it with the critical value. A larger statistic rejects uniformity.
How does the Berkowitz likelihood ratio test work?
Transform PITs with the inverse normal CDF. Fit a normal AR(1) model and compare its likelihood with the restricted model of zero mean, unit variance and zero autocorrelation. In the standard test the LR statistic is χ² with 3 degrees of freedom.
Do KS and AD test independence of PITs?
No. They assume the PITs are independent and test only whether their distribution is uniform. Berkowitz includes an autocorrelation parameter, so it tests independence as well.