Skip to content

FRM Exam Part II · Beyond Exceedance-Based Backtesting of Value-at-Risk Models

Expected Shortfall Backtesting and Elicitability Explained

Updated 11 October 2026 · Fact-checked

Expected shortfall (ES) is the average loss beyond VaR. It is not elicitable alone, because no scoring function has ES as its sole minimiser. But the pair (VaR, ES) is jointly elicitable, so you can backtest ES with joint scoring or tail-based tests. Under FRTB, banks backtest VaR exceptions, not ES directly.

Understand Expected Shortfall Backtesting and Elicitability

A backtest compares a model's forecasts with realised outcomes. For VaR it is easy: count the days when the loss exceeded VaR and compare with the expected number. Each day is a yes/no event.

ES is harder. ES is the expected loss given that the loss exceeds VaR. It depends on the size of tail losses, not just whether they happened. So you need loss magnitudes in the tail, and tail events are rare. Data are thin.

Elicitability asks: is there a scoring function so that the true value of the risk measure is the forecast that minimises the expected score? If yes, you can rank competing forecasts by average score. VaR (a quantile) is elicitable, using the quantile (pinball) loss. The mean is elicitable with squared error. ES is not elicitable on its own. Its score would need to depend on information that cannot be pinned down without also knowing VaR.

The fix is a key result: the pair (VaR, ES) is jointly elicitable. A joint scoring function exists that is minimised only when both forecasts are correct. You can then compare models by their average joint score, and a lower score is better. Do not read this as 'ES cannot be backtested'. The correct point is that ES cannot be ranked by a score on its own, but it can be tested.

Other approaches are tail-based tests. These use the losses beyond VaR, for example comparing the average realised tail loss with forecast ES, or using the probability integral transform on tail observations. Under Basel FRTB, internal models use 97.5% ES for capital, but the formal backtesting is on 99% and 97.5% VaR exceptions at desk level. ES is covered through P&L attribution and VaR backtests rather than a direct ES test.

Key formulas to remember

Expected shortfall
ES(α) = E[L | L ≥ VaR(α)]
Average loss in the tail beyond VaR at confidence level α. Always at least VaR.
Quantile (pinball) score for VaR
S(v, y) = (1{y ≤ v} − α)(v − y)
Here y is the realised P&L (or return) and v is the quantile forecast at lower tail level α. Minimised in expectation at the true quantile, so VaR is elicitable.
Elicitability rule
VaR: elicitable. ES: not elicitable alone. (VaR, ES): jointly elicitable.
Memorise all three statements. Exams test this directly.
Tail excess ratio (simple tail check)
Average realised loss on exception days ÷ forecast ES
A ratio well above 1 suggests ES is understated. A ratio near 1 is consistent with the forecast. It is a heuristic and has low power with few exceptions.
FRTB capital confidence level
Internal models ES at 97.5%
Backtesting under FRTB uses 99% and 97.5% VaR exceptions over 250 days, assessed at desk level.

How to solve Expected Shortfall Backtesting and Elicitability questions

Use this method for any question on ES backtesting, elicitability or the FRTB link.

  1. 1Identify what is asked: ranking forecasts, testing ES, or regulatory treatment.
  2. 2Check which measure is involved. VaR alone, ES alone, or the pair.
  3. 3If elicitability is asked, apply the rule: VaR yes, ES alone no, (VaR, ES) jointly yes.
  4. 4If a score comparison is given, remember lower average score means a better forecast, and the score must be the joint VaR-ES function for ES questions.
  5. 5If tail losses are given, compute average realised loss on exception days and compare with forecast ES, noting the small sample.
  6. 6For regulatory questions, recall that FRTB uses ES for capital but backtests VaR exceptions at desk level.
  7. 7State the interpretation in one line: what the result says about model accuracy and why.

Quickest way: Three-line elicitability check

When to use it: Use it on multiple-choice questions asking which statement about ES backtesting is correct.

  1. Reject any option saying ES is elicitable by itself, or that it is impossible to backtest.
  2. Prefer the option saying the pair (VaR, ES) is jointly elicitable.
  3. For FRTB options, pick the one with ES for capital and VaR exceptions for backtesting.

Common mistakes in Expected Shortfall Backtesting and Elicitability

  • Saying ES cannot be backtested because it is not elicitable.

    Students merge elicitability with testability.

    Fix: Elicitability concerns ranking by a score. ES can still be tested jointly with VaR or by tail-based tests.

  • Saying VaR is not elicitable because it is not coherent.

    Coherence and elicitability are confused.

    Fix: VaR is elicitable through the quantile score but not subadditive. ES is coherent but not elicitable alone. They are separate properties.

  • Claiming FRTB backtests ES exceptions.

    Capital uses ES, so students assume the test does too.

    Fix: FRTB backtests VaR exceptions at 99% and 97.5% for desks. ES drives capital.

  • Treating a lower joint score as worse.

    Scores feel like points to maximise.

    Fix: Scoring functions are losses. Lower average score means better forecasts.

  • Trusting a tail ratio near 1 as strong proof.

    Few exceptions are seen, so the sample is small.

    Fix: Say the test has low power with few tail observations and the result is only consistent with the model.

Worked examples

Example 1

A bank's 97.5% one-day ES forecast is USD 8.0 million. Over a year, VaR was exceeded on 7 days, with losses of USD 7, 9, 10, 8, 12, 9 and 11 million. Compute the tail excess ratio and interpret it.

Show the solution
  1. Sum the exception losses: 7 + 9 + 10 + 8 + 12 + 9 + 11 = 66.
  2. Average = 66 ÷ 7 = 9.43 million approximately.
  3. Ratio = 9.43 ÷ 8.0 = 1.18 approximately.
  4. A ratio above 1 suggests realised tail losses exceeded forecast ES by about 18%.
  5. With only 7 observations the test has low power, so the evidence is suggestive, not conclusive.

Answer: Ratio about 1.18. Forecast ES looks understated, but the small sample limits the conclusion.

Example 2

Which statement is correct? A) ES is elicitable alone via squared error. B) ES cannot be backtested at all. C) (VaR, ES) is jointly elicitable, so models can be ranked by a joint score. D) VaR is not elicitable. Explain.

Show the solution
  1. A is false. Squared error elicits the mean, not ES, and ES is not elicitable alone.
  2. B is false. ES can be tested using joint scoring or tail-based tests.
  3. D is false. VaR is elicitable using the quantile score.
  4. C is true. The pair has a joint scoring function minimised at the true values.

Answer: C. The pair (VaR, ES) is jointly elicitable, so competing models can be compared by lower average joint score.

Exam tips

  • Memorise the three-way rule: VaR elicitable, ES alone not, the pair jointly yes.
  • Watch for options that confuse elicitability with coherence. ES is coherent, VaR is not.
  • For FRTB questions, split the roles: ES at 97.5% for capital, VaR exceptions for backtesting.
  • In tail-ratio questions, compute the average on exception days only and mention the small sample.
  • Remember a lower joint score means a better model.

Practice questions from Beyond Exceedance-Based Backtesting of Value-at-Risk Models

Expected Shortfall Backtesting and Elicitability in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Expected Shortfall Backtesting and Elicitability: frequently asked questions

Is expected shortfall elicitable?

Not on its own. No scoring function has ES as its unique minimiser by itself. The pair (VaR, ES) is jointly elicitable, which allows joint scoring.

How do you backtest expected shortfall under FRTB?

FRTB does not directly test ES. It backtests VaR exceptions at 99% and 97.5% at desk level, and uses P&L attribution as another eligibility test. ES at 97.5% sets the internal model capital.

What is the difference between VaR and ES backtesting?

VaR backtesting counts exceptions, a yes/no outcome, and tests the count. ES backtesting must use the size of losses beyond VaR, so it needs joint scoring or tail-based tests and has less data.

Does non-elicitability make ES a poor risk measure?

No. ES is coherent and captures tail severity. Non-elicitability only limits direct ranking by a single score. You handle it by scoring VaR and ES together.