Skip to content

FRM Exam Part II · Validating Bank Holding Companies' Value-at-Risk Models for Market Risk

VaR Benchmarking, Stress Testing and Sensitivity Analysis

Updated 11 October 2026 · Fact-checked

Benchmarking, stress testing and sensitivity analysis are validation tools used beyond backtesting. Benchmarking compares VaR with alternative models or hypothetical portfolio results. Stress tests probe extreme losses VaR does not capture. Sensitivity analysis changes one input or assumption at a time to see how VaR moves. Together they test whether a model is sound.

Understand Benchmarking, Stress Testing and Sensitivity Analysis

Backtesting asks one question: did actual losses exceed VaR too often? It is a strong test, but it has limits. At 99% confidence over 250 days you expect only about 2.5 exceptions a year. With so few observations, a bad model can look fine. Also, backtesting only judges the past, and only the portfolio the bank actually held.

So validators add other tools. Benchmarking compares the bank's VaR with an independent estimate. The benchmark may be an alternative model, such as historical simulation against a parametric model, or the same model built by a separate team. Large, persistent gaps are a flag. They do not tell you which model is right. They tell you to investigate.

Hypothetical portfolio testing runs the model on portfolios designed by the validator. Examples are single-risk-factor positions, concentrated positions, and portfolios with long and short legs or option payoffs. You know roughly what VaR should be, so you can check that the model responds correctly. It also lets you test risks that the actual book may not hold today. This helps when there are too few exceptions to judge the model.

Stress testing applies extreme but plausible moves, historical or hypothetical, and measures the loss. VaR says what is lost at a stated confidence level. It says little about the tail beyond it. Stress tests fill that gap. They complement VaR and do not validate its quantile.

Sensitivity analysis changes one assumption at a time, such as lookback window, decay factor, confidence level, holding period, correlation or volatility input, and records the change in VaR. If a small change in an assumption causes a big change in VaR, the model is fragile and the assumption needs support. The supervisory view is that no single test is enough, so a sound validation uses several.

Key formulas to remember

Expected exceptions
Expected exceptions = N × (1 − c)
N = number of days, c = VaR confidence level. 250 days at 99% gives 2.5. Shows why backtesting has low power.
Benchmark gap
Gap % = (VaR model − VaR benchmark) ÷ VaR benchmark × 100
A diagnostic only. A large gap means investigate, not that the model is wrong.
Sensitivity of VaR
Sensitivity = (VaR after change − VaR base) ÷ VaR base
Change one input at a time and keep everything else fixed.
Role of each tool
Backtest = outcomes; Benchmark = comparison; Hypothetical portfolio = known answer; Stress = tail losses; Sensitivity = assumption fragility
Use this to match a tool to the validation question.

How to solve Benchmarking, Stress Testing and Sensitivity Analysis questions

Use this method for any question asking which validation tool fits a situation or how to read its result.

  1. 1Identify the validation concern: too few exceptions, unknown tail loss, suspect assumption, model comparison, or untested risk.
  2. 2Match the tool: benchmarking for comparison, hypothetical portfolios for known-answer or untested risks, stress tests for tail loss, sensitivity analysis for assumption dependence.
  3. 3Check what the tool can and cannot conclude. A benchmark gap flags a problem but does not say which model is wrong.
  4. 4For sensitivity, confirm only one input changes at a time and compute the percentage change in VaR.
  5. 5Interpret the size of the result: small changes show robustness, large changes show fragility or dependence on a weak assumption.
  6. 6Link back to backtesting: state that these tools support but do not replace it, since they add evidence where exceptions are few.
  7. 7Choose the option that is precise and does not overstate what the test proves.

Quickest way: Match the problem to the tool

When to use it: For short conceptual MCQs under time pressure.

  1. Read the keyword: 'alternative model' means benchmarking; 'designed portfolio' means hypothetical; 'extreme scenario' means stress; 'change one input' means sensitivity.
  2. Eliminate options that say a test proves the model is correct or wrong.
  3. Eliminate options claiming stress tests validate the VaR quantile.
  4. Pick the option that treats the tools as complements to backtesting.

Common mistakes in Benchmarking, Stress Testing and Sensitivity Analysis

  • Saying a benchmark gap proves the bank's model is wrong.

    The benchmark looks like a reference answer.

    Fix: A benchmark is another estimate with its own errors. A gap means investigate the cause.

  • Treating stress testing as a way to check VaR's confidence level.

    Both measure potential loss.

    Fix: Stress tests look at extreme scenarios beyond the VaR quantile. They complement VaR and do not test its coverage.

  • Changing several inputs at once in a sensitivity analysis.

    It feels faster.

    Fix: Change one input at a time so you can attribute the VaR change to that input.

  • Assuming zero or few exceptions prove the model is good.

    Fewer exceptions look safer.

    Fix: Few exceptions can mean the model is too conservative or the test has low power. Add benchmarking and hypothetical portfolios.

  • Using only the actual portfolio to validate the model.

    It is the portfolio that matters to the bank.

    Fix: Hypothetical portfolios test concentrations, option payoffs and risk factors the current book may not hold.

Worked examples

Example 1

A bank's 99% one-day VaR is ₹42 crore from a parametric model. A validator's historical simulation benchmark gives ₹56 crore for the same portfolio. Compute the gap relative to the benchmark and state the correct conclusion.

Show the solution
  1. Gap = (42 − 56) ÷ 56 = −14 ÷ 56.
  2. −14 ÷ 56 = −0.25, so the gap is −25%.
  3. The bank's model gives a VaR 25% lower than the benchmark.
  4. This is a flag to investigate, for example fat tails or the lookback window, not proof that the bank model is wrong.

Answer: The gap is −25%. The bank's VaR is lower than the benchmark. This calls for investigation and does not prove the model wrong.

Example 2

A validator tests a USD portfolio's 99% VaR. The base case with a 250-day window gives USD 8.0 million. Changing only the window to 500 days gives USD 10.0 million. What is the sensitivity, and what does it suggest?

Show the solution
  1. Sensitivity = (10.0 − 8.0) ÷ 8.0.
  2. 2.0 ÷ 8.0 = 0.25, which is +25%.
  3. Only the window changed, so the change is attributable to window length.
  4. A 25% rise from one assumption shows VaR depends strongly on the lookback window, so that choice needs justification.

Answer: Sensitivity is +25%. VaR is sensitive to the lookback window, so the window choice must be justified and monitored.

Exam tips

  • Match the tool to the concern. Questions often describe a situation and ask which test fits.
  • Reject options that say a single test proves model validity.
  • When a question mentions too few exceptions, think benchmarking and hypothetical portfolios.
  • Remember stress tests cover losses beyond VaR and do not test the VaR quantile.
  • For sensitivity calculations, use the base VaR as the denominator.

Practice questions from Validating Bank Holding Companies' Value-at-Risk Models for Market Risk

Benchmarking, Stress Testing and Sensitivity Analysis: frequently asked questions

How do you validate a VaR model without enough exceptions?

Backtesting has low power when exceptions are rare. Add benchmarking against alternative models, hypothetical portfolios with known expected results, and sensitivity analysis of key assumptions. Use these together with backtesting.

What is the difference between stress testing and VaR validation?

VaR validation checks whether the model's quantile and assumptions are sound. Stress testing measures losses under extreme scenarios that VaR may not capture. Stress tests complement VaR but do not validate its confidence level.

Why use a hypothetical portfolio to test a VaR model?

You design the portfolio, so you can predict roughly what VaR should be and check the model's response. It also tests risks the actual book does not hold today, such as concentrations or option positions.

What does VaR sensitivity analysis show?

It shows how much VaR changes when you alter one assumption, such as window length, confidence level, holding period or correlation. Large changes show the model depends heavily on that assumption.