FRM Exam Part II · Beyond Exceedance-Based Backtesting of Value-at-Risk Models
Backtesting VaR Models with the Probability Integral Transform
Updated 11 October 2026 · Fact-checked
The probability integral transform (PIT) backtest converts each realized P&L into its cumulative probability under that day's forecast distribution. If the model is right, these PIT values are iid uniform on [0,1]. You test uniformity (for example Kolmogorov-Smirnov) and independence, or apply the Berkowitz normal transform and a likelihood ratio test.
Understand Backtesting with Probability Integral Transform (PIT)
Exceedance-based backtesting looks at one quantile. It counts how often losses beat VaR. A model can pass that test and still misstate the rest of the distribution, such as the tail beyond VaR or the centre. The PIT approach tests the whole forecast distribution.
Here is the idea. Each day t, your model gives a forecast cumulative distribution function F_t for P&L. When the day ends, you observe the realized P&L, x_t. Compute u_t = F_t(x_t). This is the probability integral transform value. It tells you where the actual outcome fell in the forecast distribution. u_t = 0.03 means the outcome was in the worst 3% of what you predicted.
If F_t is the true conditional distribution, then u_t is uniform on [0,1]. It is also independent across days. So the series u_1, u_2, ... should look like iid U(0,1) draws. Too many values near 0 or 1 means the tails are too thin. Too many near 0.5 means the distribution is too wide. A hump-shaped or U-shaped histogram is the classic visual check. A VaR exception at 99% is just u_t < 0.01, so exceedance counting uses a tiny slice of the information.
To test it formally, you have two routes. First, test uniformity directly with distribution tests such as Kolmogorov-Smirnov (largest gap between the empirical CDF of the u_t and the 45-degree line) or Anderson-Darling (which weights the tails more). Second, the Berkowitz approach maps each u_t to z_t = Φ⁻¹(u_t), which should be iid standard normal. Then you fit an AR(1) with mean and variance and use a likelihood ratio test of mean 0, variance 1 and no autocorrelation.
Two cautions. The test checks the model jointly: distribution shape and independence. Rejection tells you something is wrong, not what. Also, with limited data, tests have low power, especially in the tails.
Key formulas to remember
- PIT value
- u_t = F_t(x_t)
- F_t is the forecast CDF for day t's P&L. x_t is the realized P&L. Under a correct model, u_t ~ iid U(0,1).
- Kolmogorov-Smirnov statistic
- D = max | F_emp(u) − u |
- Largest vertical distance between the empirical CDF of the PIT values and the uniform CDF. Larger D means stronger evidence against the model.
- Berkowitz normal transform
- z_t = Φ⁻¹(u_t)
- Under a correct model, z_t is iid N(0,1). Φ⁻¹ is the inverse standard normal CDF.
- Berkowitz alternative model
- z_t − μ = ρ(z_{t−1} − μ) + ε_t, with ε_t ~ N(0, σ²)
- Null hypothesis: μ = 0, σ = 1, ρ = 0.
- Berkowitz likelihood ratio test
- LR = −2[L(0, 1, 0) − L(μ̂, σ̂², ρ̂)] ~ χ²(3)
- Three restrictions under the null, so 3 degrees of freedom. Reject if LR exceeds the chi-square critical value.
- Link to VaR exceptions
- Exception at confidence c ⇔ u_t < 1 − c
- For 99% VaR, an exception is u_t < 0.01. Exceedance tests use only this one region.
How to solve Backtesting with Probability Integral Transform (PIT) questions
Use this sequence for any PIT backtesting question, whether it asks you to compute, interpret or choose a test.
- 1Identify the forecast distribution F_t for each day and the realized P&L x_t.
- 2Compute u_t = F_t(x_t). For a normal forecast, u_t = Φ((x_t − μ_t) ÷ σ_t).
- 3State what a correct model implies: u_t are iid U(0,1), so about 10% fall in each decile and there is no serial dependence.
- 4Choose the test: Kolmogorov-Smirnov or Anderson-Darling for uniformity, Berkowitz LR for normality, mean, variance and autocorrelation of z_t = Φ⁻¹(u_t).
- 5Compute the statistic and compare it with the critical value. Reject the model if the statistic is larger.
- 6Interpret the pattern: clustering near 0 and 1 means tails too thin or volatility too low. Clustering near 0.5 means the forecast is too wide.
- 7Remember what it does not show: rejection signals misspecification but not its cause.
Quickest way: Read the PIT pattern first
When to use it: Use when a question shows a PIT histogram, a summary of u_t values, or Berkowitz estimates and asks what is wrong with the model.
- Check the shape: flat is good. U-shaped means too many extreme outcomes, so the model underestimates risk. Hump-shaped means the model overestimates risk.
- Check the Berkowitz estimates: μ̂ not 0 means biased mean. σ̂ above 1 means volatility understated. ρ̂ not 0 means dependence the model ignores.
- Count degrees of freedom: Berkowitz LR uses 3, so compare with χ²(3). Its 5% critical value is about 7.81.
- If the statistic exceeds the critical value, reject. Otherwise do not reject, which is not proof the model is right.
Common mistakes in Backtesting with Probability Integral Transform (PIT)
Saying PIT values should be standard normal.
Students mix up the PIT value u_t with the Berkowitz transform z_t.
Fix: u_t is uniform on [0,1]. Only after applying Φ⁻¹ do you get z_t ~ N(0,1).
Testing uniformity but ignoring independence.
The KS test sounds complete, so students stop there.
Fix: KS checks the marginal distribution only. Independence needs a separate check, such as autocorrelation or the ρ term in Berkowitz.
Using the wrong degrees of freedom for the Berkowitz LR test.
Students count only mean and variance.
Fix: Three restrictions (μ = 0, σ² = 1, ρ = 0), so χ²(3).
Concluding that a non-rejected model is proven correct.
Confusing failure to reject with acceptance.
Fix: Say there is no evidence against the model. Power is low with small samples, especially in the tails.
Using one fixed distribution for all days.
Students compute PIT with an unconditional CDF.
Fix: Use each day's own forecast CDF F_t. If volatility changes, the forecast changes.
Reading a U-shaped PIT histogram as a model that is too conservative.
Confusion over what mass near 0 and 1 means.
Fix: Mass at both ends means outcomes land in the forecast tails too often. The model understates risk.
Worked examples
Example 1
A bank forecasts tomorrow's P&L as normal with mean 0 and standard deviation USD 2 million. Realized P&L is −USD 3.29 million. (a) Compute the PIT value. (b) Is this a 99% VaR exception? Use Φ(−1.645) = 0.05 and Φ(−2.326) = 0.01.
Show the solution
- Standardize: z = (−3.29 − 0) ÷ 2 = −1.645.
- PIT value u = Φ(−1.645) = 0.05.
- A 99% VaR exception requires u < 0.01.
- 0.05 is not below 0.01, so it is not an exception.
Answer: u = 0.05. Not a 99% VaR exception, although the outcome is in the worst 5% of the forecast distribution, so PIT carries more information than the exception flag.
Example 2
A backtest of 250 days gives Berkowitz maximum-likelihood estimates μ̂ = 0, σ̂ = 1.3, ρ̂ = 0. The log-likelihood at the estimates is higher than under the null by 6.5. The 5% critical value of χ²(3) is 7.81. What do you conclude?
Show the solution
- LR = 2 × (difference in log-likelihood) = 2 × 6.5 = 13.0.
- Degrees of freedom = 3, critical value 7.81.
- 13.0 > 7.81, so reject the null at 5%.
- σ̂ = 1.3 > 1 means realized outcomes are more dispersed than forecast, so the model understates volatility.
Answer: Reject the model at 5%. The likely issue is understated volatility, since μ̂ and ρ̂ are consistent with the null.
Exam tips
- Know the sentence: if the model is right, PIT values are iid uniform. Many options are built from it.
- Berkowitz degrees of freedom are 3. Memorize this.
- Expect interpretation questions: U-shaped PIT means risk understated, hump-shaped means overstated.
- Be ready to say why PIT beats exception counting: it uses the whole distribution, not one quantile.
- Choose between KS and Anderson-Darling: Anderson-Darling puts more weight on the tails.
Practice questions from Beyond Exceedance-Based Backtesting of Value-at-Risk Models
- When using a scoring function to compare competing VaR forecasts, which property makes a scoring function 'strictly consistent' for the quan…
- A risk analyst evaluates a bank's daily 99% VaR model using the probability integral transform (PIT). For each day, she computes the value o…
- Two models forecast 97.5% ES for a portfolio. Model A and Model B both pass a standard exceedance-count test on VaR at 97.5%. Over the sampl…
- A risk manager wants to test whether a bank's daily trading P&L is consistent with the full predictive distribution produced by its VaR mode…
- A bank's two VaR models both pass the Kupiec unconditional coverage test at 99%, each with 2 or 3 exceedances in 250 days. A validator wants…
Backtesting with Probability Integral Transform (PIT): frequently asked questions
What is the probability integral transform in VaR backtesting?
It is the forecast CDF evaluated at the realized P&L. The result is a number between 0 and 1 showing where the outcome fell in the predicted distribution. If the model is correct, these numbers are iid uniform.
How do I test PIT uniformity?
Use the Kolmogorov-Smirnov test, which measures the largest gap between the empirical CDF of the PIT values and the uniform CDF. Anderson-Darling is an alternative that is more sensitive to the tails. Plot a histogram as a visual check.
What is the Berkowitz test?
It converts PIT values to z = Φ⁻¹(u), which should be iid standard normal under a correct model. A likelihood ratio test then checks that the mean is 0, the variance is 1 and there is no autocorrelation. The statistic follows a chi-square with 3 degrees of freedom.
Why is PIT backtesting better than counting exceptions?
Exception counting uses only whether a loss passed one quantile. PIT uses every observation and tests the entire forecast distribution. It can detect a bad tail or a wrong centre even when the exception count looks fine.