Skip to content

FRM Part II · FRM Exam Part II

Beyond Exceedance-Based Backtesting of Value-at-Risk Models: formula sheet

Full chapter guide

Key formulas

Expected number of exceptions
E[N] = T × p, where p = 1 − confidence level
T is the number of days. At 99% over 250 days, E[N] = 2.5.
Unconditional coverage (Kupiec) likelihood ratio
LR_uc = −2 ln[(1 − p)^(T−N) × p^N] + 2 ln[(1 − N/T)^(T−N) × (N/T)^N]
Compared with a chi-square distribution with 1 degree of freedom. The 5% critical value is 3.84. It tests only the frequency of exceptions.
Independence idea (Christoffersen)
LR_cc = LR_uc + LR_ind, chi-square with 2 degrees of freedom
LR_ind tests whether today's exception depends on yesterday's. Conditional coverage tests frequency and independence together. The 5% critical value is 5.99.
Basel traffic light zones (250 days, 99% VaR)
Green: 0–4 exceptions; Yellow: 5–9; Red: 10 or more
Yellow-zone exceptions raise the capital multiplier. Red usually means the model is presumed flawed.
PIT value
u_t = F_t(x_t)
F_t is the forecast CDF for day t's P&L. x_t is the realized P&L. Under a correct model, u_t ~ iid U(0,1).
Kolmogorov-Smirnov statistic
D = max | F_emp(u) − u |
Largest vertical distance between the empirical CDF of the PIT values and the uniform CDF. Larger D means stronger evidence against the model.
Berkowitz normal transform
z_t = Φ⁻¹(u_t)
Under a correct model, z_t is iid N(0,1). Φ⁻¹ is the inverse standard normal CDF.
Berkowitz alternative model
z_t − μ = ρ(z_{t−1} − μ) + ε_t, with ε_t ~ N(0, σ²)
Null hypothesis: μ = 0, σ = 1, ρ = 0.
Berkowitz likelihood ratio test
LR = −2[L(0, 1, 0) − L(μ̂, σ̂², ρ̂)] ~ χ²(3)
Three restrictions under the null, so 3 degrees of freedom. Reject if LR exceeds the chi-square critical value.
Link to VaR exceptions
Exception at confidence c ⇔ u_t < 1 − c
For 99% VaR, an exception is u_t < 0.01. Exceedance tests use only this one region.
PIT value
u_t = F_t(r_t)
F_t is the model's forecast CDF for day t. Under a correct model the u_t are iid U(0,1).
Kolmogorov-Smirnov statistic
D = max over u of |F_n(u) − u|
F_n is the empirical CDF of the PITs. With sorted values u(1) ≤ … ≤ u(n), D = max of [i/n − u(i)] and [u(i) − (i−1)/n]. A large D means rejection. Every gap has equal weight.
Anderson-Darling statistic
A² = −n − (1/n) Σ (2i − 1) × [ln u(i) + ln(1 − u(n+1−i))]
Sum over i = 1 to n on sorted PITs. This is the computational form of A² = n ∫ (F_n(u) − u)² ÷ [u(1 − u)] du, which is where the weight 1/[u(1 − u)] comes from. Tail-sensitive. Large A² means rejection. Critical values come from tables.
Berkowitz transform
z_t = Φ⁻¹(u_t)
Under the null, z_t are iid N(0,1).
Berkowitz alternative model
z_t − μ = ρ(z_t−1 − μ) + ε_t, ε_t ~ N(0, σ²)
Null hypothesis: μ = 0, σ² = 1, ρ = 0.
Berkowitz LR statistic
LR = −2 × [L(0, 1, 0) − L(μ̂, σ̂², ρ̂)] ~ χ²(3)
L is the log-likelihood. For the standard Berkowitz test, three restrictions give 3 degrees of freedom. The 5% critical value is 7.815. The tail version has a different likelihood and its own test setup.
Expected shortfall
ES(α) = E[L | L ≥ VaR(α)]
Average loss in the tail beyond VaR at confidence level α. Always at least VaR.
Quantile (pinball) score for VaR
S(v, y) = (1{y ≤ v} − α)(v − y)
Here y is the realised P&L (or return) and v is the quantile forecast at lower tail level α. Minimised in expectation at the true quantile, so VaR is elicitable.
Elicitability rule
VaR: elicitable. ES: not elicitable alone. (VaR, ES): jointly elicitable.
Memorise all three statements. Exams test this directly.
Tail excess ratio (simple tail check)
Average realised loss on exception days ÷ forecast ES
A ratio well above 1 suggests ES is understated. A ratio near 1 is consistent with the forecast. It is a heuristic and has low power with few exceptions.
FRTB capital confidence level
Internal models ES at 97.5%
Backtesting under FRTB uses 99% and 97.5% VaR exceptions over 250 days, assessed at desk level.
Quantile (pinball) loss for VaR at level α
S = (α − 1{L ≤ VaR}) × (L − VaR)
L is the realised loss and 1{L ≤ VaR} equals 1 if L ≤ VaR, otherwise 0. This equals α × (L − VaR) on an exception and (1 − α) × (VaR − L) otherwise. The next entry shows the same rule as two cases.
Two-case quantile loss (losses as positive numbers)
If L > VaR: S = α × (L − VaR). If L ≤ VaR: S = (1 − α) × (VaR − L)
Here L is the realised loss, α is the VaR confidence level (e.g. 0.99). Exceptions get weight α, non-exceptions get weight 1 − α.
Average score for a model
Average S = (1 ÷ T) × Σ S_t
Sum over the same T days for each model. Lower is better.
Consistency and elicitability
True α-quantile minimises E[S(x, Y)]
This holds for the quantile loss. It makes VaR elicitable, so models can be ranked by score.
Basel traffic light zones (250 days, 99% VaR)
Green: 0–4 exceptions. Yellow: 5–9. Red: 10 or more.
This is an exception-count rule, not a score. It ignores the size of the exceptions.

Quick revision

  • Exceedance backtests use one binary signal per day and ignore how large the loss was.
  • A correct model gives PIT values that are independent and uniform on [0, 1].
  • PIT value = forecast CDF evaluated at the realised return.
  • Kolmogorov-Smirnov uses the largest gap between empirical and theoretical CDF.
  • Anderson-Darling is similar but puts more weight on the tails.
  • Berkowitz maps PIT values to standard normal and uses a likelihood ratio test.
  • Berkowitz can test for mean, variance and autocorrelation in the transformed series.
  • Too many PIT values near 0 or 1 suggest the model understates tail risk.
  • ES looks at the average loss beyond VaR, so it captures tail severity.
  • A statistic is elicitable if a scoring function exists that the true value minimises in expectation.
  • VaR is elicitable; ES on its own is not, but it can be elicited jointly with VaR.
  • Scoring functions rank competing models; a lower expected score (for loss-type scores) means a better forecast.

Common mistakes

  • Saying a model that passes the Kupiec test is accurate. Fix: Failing to reject only means the data give no strong evidence against the model. With low power, bad models can pass.
  • Believing Kupiec detects clustering of exceptions. Fix: Kupiec checks only frequency (unconditional coverage). Clustering needs the independence test or the combined conditional coverage test.
  • Saying PIT values should be standard normal. Fix: u_t is uniform on [0,1]. Only after applying Φ⁻¹ do you get z_t ~ N(0,1).
  • Testing uniformity but ignoring independence. Fix: KS checks the marginal distribution only. Independence needs a separate check, such as autocorrelation or the ρ term in Berkowitz.
  • Using KS or AD to conclude the PITs are independent. Fix: KS and AD assume independence and test only uniformity. Use Berkowitz or a separate autocorrelation test for independence.
  • Applying the Berkowitz likelihood to raw PIT values. Fix: Always convert with z_t = Φ⁻¹(u_t) first. The test is built on normal scores.
  • Saying ES cannot be backtested because it is not elicitable. Fix: Elicitability concerns ranking by a score. ES can still be tested jointly with VaR or by tail-based tests.
  • Saying VaR is not elicitable because it is not coherent. Fix: VaR is elicitable through the quantile score but not subadditive. ES is coherent but not elicitable alone. They are separate properties.
  • Using the same weight for exceptions and non-exceptions Fix: Use α on exceedances and 1 − α on the rest. The asymmetry is what makes the quantile the optimal forecast.
  • Thinking the model with fewer exceptions always wins Fix: A very high VaR has few exceptions but pays many small penalties. Compute the full score.

Exam tips

  • Know the three losses of information: size, timing and clustering. Questions often ask for one that a given test misses.
  • Link each limitation to its remedy: clustering to Christoffersen, size to ES or PIT-based tests, small samples to longer data.
  • Be exact on the language: 'fail to reject' is not 'accept'.
  • Remember the zones for 250 days at 99%: green 0–4, yellow 5–9, red 10 or more.
  • If numbers are given, compute expected exceptions first. It anchors every other step.
  • Know the sentence: if the model is right, PIT values are iid uniform. Many options are built from it.
  • Berkowitz degrees of freedom are 3. Memorize this.
  • Expect interpretation questions: U-shaped PIT means risk understated, hump-shaped means overstated.