Skip to content

FRM Part II · FRM Exam Part II

Beyond Exceedance-Based Backtesting of Value-at-Risk Models

Exceedance-based backtesting only counts how often losses exceed VaR. Beyond it, you test the whole forecast distribution using the probability integral transform (PIT), which should be uniform if the model is right. You then apply KS, Anderson-Darling or Berkowitz tests, backtest Expected Shortfall, and compare models with scoring functions.

What this chapter covers

This chapter extends VaR backtesting. The basic method counts exceedances and checks them against the confidence level, for example with a binomial or Kupiec-type test. That method uses one yes/no signal per day. It ignores how large the losses are and how the model performs in the rest of the distribution.

The chapter fixes this in steps. First you see why exceedance counting has low power and misses clustering and tail severity. Then you learn the probability integral transform (PIT): each realised return is mapped to its percentile under the forecast distribution. If the model is correct, these values are independent and uniform on [0, 1]. The distribution tests (Kolmogorov-Smirnov, Anderson-Darling, Berkowitz) check that claim in different ways. The chapter then moves to Expected Shortfall (ES) backtesting and the idea of elicitability, and ends with scoring functions for ranking competing models.

This links directly to the Market Risk topic in Part II, where VaR and ES models are built, and to the Basel framework, which moved the trading book measure from VaR to ES. It also supports model risk questions in other topics. Expect applied questions that give you a result and ask what it says about the model.

Backtesting questions are practical and case-like, which suits the 80-question Part II format. This chapter lets you answer them with precision: you must pick the right test for the flaw described, read the result correctly, and explain the limits of each method. Candidates who only memorise names lose marks on interpretation, so the effort pays off in market risk and in model validation questions elsewhere in the paper.

Beyond Exceedance-Based Backtesting of Value-at-Risk Models: topics in the order to study them

  1. 1Limitations of Exceedance-Based VaR BacktestingStart here because it explains the problems (low power, no severity, ignores clustering) that every later method tries to solve.
  2. 2Backtesting with Probability Integral Transform (PIT)PIT is the building block for all the distribution tests, so you need it before them.
  3. 3Distribution Tests: Kolmogorov-Smirnov, Anderson-Darling, BerkowitzThese tests are applied to PIT values, so they come after PIT and show how to judge uniformity and independence.
  4. 4Expected Shortfall Backtesting and ElicitabilityOnce you can test a full distribution, you can see why ES is harder to backtest directly and what elicitability means.
  5. 5Scoring Functions and Comparing VaR ModelsLast, because it builds on elicitability: scoring functions rank models rather than just accept or reject them.

How to prepare Beyond Exceedance-Based Backtesting of Value-at-Risk Models

Treat this as a short logical chain. Each method answers a weakness of the one before it. Aim to explain why each tool exists, not just what it does.

  1. Write down three weaknesses of exceedance counting in your own words, such as ignoring loss size and clustering.
  2. Learn PIT as a procedure: forecast distribution, realised return, percentile. State what a correct model implies for the PIT series.
  3. Make a one-line comparison of KS, Anderson-Darling and Berkowitz: what each tests and where it is strongest, for example Anderson-Darling giving more weight to the tails.
  4. Learn the Berkowitz idea: transform PIT values to a standard normal and test with a likelihood ratio, which can also check independence.
  5. Define elicitability simply and note why VaR is elicitable while ES alone is not, and what that means for backtesting and scoring.
  6. Practise interpreting scenarios: given a test result or score, say what is wrong with the model and what you would do next.
  7. Do timed mixed questions on this chapter and review each wrong answer by the concept it tested.

Common mistakes in Beyond Exceedance-Based Backtesting of Value-at-Risk Models

  • Saying a model passes because exceedances match the expected count.

    Fix: Always check whether the question hints at clustering or large tail losses, then name the method that captures it.

  • Confusing what the PIT series should look like.

    Fix: Remember: PIT values are uniform; the Berkowitz step converts them to standard normal before testing.

  • Mixing up KS and Anderson-Darling.

    Fix: Link KS to the maximum distance and Anderson-Darling to extra weight on the tails.

  • Claiming ES cannot be backtested at all because it is not elicitable.

    Fix: State that ES lacks a stand-alone scoring function, yet it can be backtested and jointly scored with VaR.

  • Treating a test as proof the model is right.

    Fix: Say that not rejecting means no evidence against the model at that power, and that tests have limited power with small samples.

  • Reading a scoring function the wrong way round.

    Fix: Check the definition given in the question; for loss-type scoring functions, lower average score means the better model.

Last-day revision: Beyond Exceedance-Based Backtesting of Value-at-Risk Models

  • Exceedance backtests use one binary signal per day and ignore how large the loss was.
  • A correct model gives PIT values that are independent and uniform on [0, 1].
  • PIT value = forecast CDF evaluated at the realised return.
  • Kolmogorov-Smirnov uses the largest gap between empirical and theoretical CDF.
  • Anderson-Darling is similar but puts more weight on the tails.
  • Berkowitz maps PIT values to standard normal and uses a likelihood ratio test.
  • Berkowitz can test for mean, variance and autocorrelation in the transformed series.
  • Too many PIT values near 0 or 1 suggest the model understates tail risk.
  • ES looks at the average loss beyond VaR, so it captures tail severity.
  • A statistic is elicitable if a scoring function exists that the true value minimises in expectation.
  • VaR is elicitable; ES on its own is not, but it can be elicited jointly with VaR.
  • Scoring functions rank competing models; a lower expected score (for loss-type scores) means a better forecast.

Beyond Exceedance-Based Backtesting of Value-at-Risk Models practice questions

Beyond Exceedance-Based Backtesting of Value-at-Risk Models in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Beyond Exceedance-Based Backtesting of Value-at-Risk Models: frequently asked questions

Why is exceedance-based backtesting not enough?

It only counts breaches of VaR, so it ignores how large the losses are and whether breaches cluster. It also has low power with typical sample sizes. The methods in this chapter use more of the information in the forecast.

What is the PIT in VaR backtesting?

It is the forecast cumulative probability of each realised return. If the model is correct, the PIT values are independent and uniformly distributed on [0, 1]. Departures from this point to a misspecified model.

What does elicitability mean and why does it matter?

A statistic is elicitable if some scoring function is minimised in expectation by its true value. This lets you compare forecasts objectively. VaR is elicitable, while ES alone is not, which shapes how ES is backtested.

How should I prepare this chapter for FRM Part II?

Learn the logic chain from exceedance limits to PIT, to distribution tests, to ES and scoring. Focus on interpreting results in short scenarios rather than on derivations. Finish with timed practice questions.