Skip to content

FRM Exam Part II · Beyond Exceedance-Based Backtesting of Value-at-Risk Models

Scoring Functions and Comparing VaR Models

Updated 11 October 2026 · Fact-checked

A scoring function gives each VaR forecast a penalty once the actual loss is known. A consistent scoring function, such as the quantile (pinball) loss, is minimised in expectation by the true quantile. You average the scores for each model over the same days. The lower average score ranks higher.

Understand Scoring Functions and Comparing VaR Models

Exception counting asks one question: did the loss exceed VaR? It ignores how far above VaR the loss went and how large VaR was on the other days. A model that sets VaR very high every day will pass an exception count with few exceptions, but it is useless for capital use. Exception counting cannot tell two passing models apart.

A scoring function fixes this. It takes a VaR forecast and the realised outcome and returns one penalty number. You add up the penalties over many days. The model with the lower total or average score is the better forecaster. This lets you rank competing models, not just pass or fail one.

Not every penalty works. A scoring function is consistent for a statistic (here, the α-quantile) if the true statistic minimises the expected score. A consistent function rewards honest forecasts. A statistic that has a consistent scoring function is called elicitable. The quantile is elicitable, so VaR can be compared this way. The standard choice is the quantile loss, also called the pinball loss. It penalises an exception heavily, with weight on the shortfall, and penalises a non-exception lightly, with weight on the unused buffer.

The score is asymmetric on purpose. For a 99% VaR, an exception costs 0.99 per unit of shortfall, while a non-exception costs only 0.01 per unit of excess cover. That asymmetry is what makes the 99% quantile the best forecast. Too low a VaR gets hit by large exception penalties. Too high a VaR gets hit by many small penalties.

In practice, you use scores alongside the Basel traffic light approach. The traffic light is a regulatory exception-count test that sets a capital multiplier. Scores are a loss-based view that measures how well each model tracks the tail. Models are compared on the same data and the same confidence level. Differences in average score can also be tested for statistical significance, since a small gap over a short sample may be noise.

Key formulas to remember

Quantile (pinball) loss for VaR at level α
S = (α − 1{L ≤ VaR}) × (L − VaR)
L is the realised loss and 1{L ≤ VaR} equals 1 if L ≤ VaR, otherwise 0. This equals α × (L − VaR) on an exception and (1 − α) × (VaR − L) otherwise. The next entry shows the same rule as two cases.
Two-case quantile loss (losses as positive numbers)
If L > VaR: S = α × (L − VaR). If L ≤ VaR: S = (1 − α) × (VaR − L)
Here L is the realised loss, α is the VaR confidence level (e.g. 0.99). Exceptions get weight α, non-exceptions get weight 1 − α.
Average score for a model
Average S = (1 ÷ T) × Σ S_t
Sum over the same T days for each model. Lower is better.
Consistency and elicitability
True α-quantile minimises E[S(x, Y)]
This holds for the quantile loss. It makes VaR elicitable, so models can be ranked by score.
Basel traffic light zones (250 days, 99% VaR)
Green: 0–4 exceptions. Yellow: 5–9. Red: 10 or more.
This is an exception-count rule, not a score. It ignores the size of the exceptions.

How to solve Scoring Functions and Comparing VaR Models questions

Use this method for any question on scoring VaR forecasts or comparing models.

  1. 1Identify the confidence level α (for example 0.99 for 99% VaR) and note whether losses or returns are given.
  2. 2Convert everything to losses as positive numbers so that the VaR and the outcome are comparable.
  3. 3For each day, check whether the loss exceeds VaR.
  4. 4Apply the right branch: α × (L − VaR) for an exception, (1 − α) × (VaR − L) otherwise.
  5. 5Add the daily scores and divide by the number of days for each model.
  6. 6Rank the models: the lowest average score is best. Check that all models use the same days and the same α.
  7. 7Interpret: say whether the winner is better because it has fewer or smaller exceptions, or because it is less conservative.
  8. 8If the question mentions a traffic light, count exceptions separately and give the zone.

Quickest way: Weights first, then compare totals

When to use it: Use this when the question gives a few days of data and asks which model scores better.

  1. Write the two weights: α for exceptions, 1 − α for non-exceptions.
  2. Mark each day as exception or not for each model.
  3. Compute only the gaps (L − VaR or VaR − L), multiply by the weight, and add.
  4. Compare totals; you do not need to divide by T if both models have the same number of days.
  5. Sanity check: a model with a big exception gap at α = 0.99 should be punished heavily.

Common mistakes in Scoring Functions and Comparing VaR Models

  • Using the same weight for exceptions and non-exceptions

    Students think of the score as a plain absolute error.

    Fix: Use α on exceedances and 1 − α on the rest. The asymmetry is what makes the quantile the optimal forecast.

  • Thinking the model with fewer exceptions always wins

    Exception counting is the habit from Basel and Kupiec tests.

    Fix: A very high VaR has few exceptions but pays many small penalties. Compute the full score.

  • Picking the highest average score as best

    Scores are confused with performance measures where higher is better.

    Fix: A score is a penalty. Lower is better.

  • Saying a scoring function is consistent because it is symmetric

    Symmetric loss like squared error feels natural.

    Fix: Squared error elicits the mean, not the quantile. For VaR you need the asymmetric quantile loss.

  • Comparing models over different periods or confidence levels

    Data sets are mixed without checking.

    Fix: Use the same days, the same α and the same loss definition for every model.

  • Treating the traffic light zone as a score

    Both are called backtests.

    Fix: The traffic light counts exceptions and sets a multiplier. A score uses the size of each gap too.

Worked examples

Example 1

A bank backtests two 99% one-day VaR models over three days. Realised losses are $4m, $9m and $1m. Model A VaR is $6m, $6m, $6m. Model B VaR is $3m, $8m, $2m. Using the quantile loss, which model scores better?

Show the solution
  1. α = 0.99, so exception weight = 0.99 and non-exception weight = 0.01.
  2. Model A day 1: L = 4 ≤ 6, score = 0.01 × (6 − 4) = 0.02.
  3. Model A day 2: L = 9 > 6, score = 0.99 × (9 − 6) = 2.97.
  4. Model A day 3: L = 1 ≤ 6, score = 0.01 × (6 − 1) = 0.05.
  5. Model A total = 0.02 + 2.97 + 0.05 = 3.04.
  6. Model B day 1: L = 4 > 3, score = 0.99 × (4 − 3) = 0.99.
  7. Model B day 2: L = 9 > 8, score = 0.99 × (9 − 8) = 0.99.
  8. Model B day 3: L = 1 ≤ 2, score = 0.01 × (2 − 1) = 0.01.
  9. Model B total = 0.99 + 0.99 + 0.01 = 1.99.
  10. 1.99 < 3.04, so Model B has the lower score.

Answer: Model B scores better (total 1.99 against 3.04), even though it has two exceptions to Model A's one. Its exceptions are small, while Model A has one large exception.

Example 2

Which statement about a consistent scoring function for VaR is correct? (A) Squared error is consistent for the 99% quantile. (B) The quantile loss is minimised in expectation by the true α-quantile. (C) A model with zero exceptions always has the lowest score. (D) Lower scores mean the model is less accurate.

Show the solution
  1. Check A: squared error elicits the mean, so it is not consistent for a quantile. A is wrong.
  2. Check B: B states the consistency property of the quantile loss. B is correct.
  3. Check C: a very high VaR has zero exceptions but pays non-exception penalties each day, so its score can be high. C is wrong.
  4. Check D: scores are penalties, so lower is better. D is wrong.

Answer: B

Exam tips

  • Always write the two weights α and 1 − α before you start any scoring calculation.
  • If an option says the model with the fewest exceptions is best, test it. The score may say otherwise.
  • Know the difference: traffic light counts exceptions; the scoring function uses exception size and the VaR level.
  • Link consistency to elicitability. Expected shortfall is not elicitable on its own, but it is jointly elicitable with VaR; VaR alone is elicitable.
  • Check the question says losses or returns before you apply the formula.

Practice questions from Beyond Exceedance-Based Backtesting of Value-at-Risk Models

Scoring Functions and Comparing VaR Models: frequently asked questions

What is the quantile loss in VaR backtesting?

It is an asymmetric penalty. On an exception you pay α times the shortfall, and otherwise you pay (1 − α) times the unused cover. The true α-quantile minimises its expected value.

Why is VaR called elicitable?

Because there is a scoring function, the quantile loss, that the true VaR minimises in expectation. This lets you rank competing VaR forecasts by their average score.

How is a scoring function different from the Basel traffic light approach?

The traffic light counts exceptions over a window and assigns a zone with a capital multiplier. A scoring function also uses how big each exception is and how much cover was unused, so it can rank models that both pass.

Does a lower score always mean a better model?

Lower average score means a better forecast over that sample. Over a short sample the gap can be noise, so the difference should be tested for significance before you rely on it.