Skip to content

FRM Part II · FRM Exam Part II

Non-parametric Approaches: formula sheet

Full chapter guide

Key formulas

Historical return (simple)
R_t = (P_t − P_(t−1)) ÷ P_(t−1)
Use the same return definition consistently. Log returns are also common: ln(P_t ÷ P_(t−1)).
Scenario P&L
P&L_t = V × R_t
Apply each historical return to today's portfolio value V (or revalue the portfolio under each scenario's risk-factor changes).
Tail count
k = n × (1 − c)
Number of observations in the tail. For n = 500 and c = 99%, k = 5. Under the convention used on this page, VaR is the (k+1)-th worst outcome, so k outcomes are worse than the VaR loss. Other conventions interpolate between the k-th and (k+1)-th worst outcomes; follow the question's convention.
HS VaR
VaR(c) = −(percentile of P&L at level 1 − c)
Report as a positive loss. Equivalently, the loss ranked just beyond the tail count in the sorted list.
Horizon scaling (rule of thumb)
VaR(T days) ≈ VaR(1 day) × √T
Valid only under i.i.d. returns with zero mean; it is an approximation, not an HS output.
Bootstrap sample
Draw n observations from n historical returns, with replacement
Each sample has the same size as the original data. Repeats are expected.
Bootstrapped VaR estimate
VaR_boot = (1 ÷ B) × Σ VaR_b, for b = 1 to B
B is the number of bootstrap samples. VaR_b is the VaR from sample b.
Bootstrapped ES estimate
ES_boot = (1 ÷ B) × Σ ES_b, for b = 1 to B
ES_b is the average loss beyond VaR in sample b.
Bootstrap standard error
SE = √[ Σ (VaR_b − VaR_boot)² ÷ (B − 1) ]
This is the standard deviation of the B estimates. Dividing by B instead of B − 1 is also seen when B is large.
Approximate confidence interval
VaR_boot ± z × SE
Use z = 1.96 for 95% two-sided. Alternatively, use percentiles of the sorted bootstrap estimates.
Historical VaR position
VaR at confidence c is the loss at the (1 − c) tail percentile; with n observations, about n × (1 − c) losses lie beyond it
Interpolation conventions vary, so read the question's stated method.
BRW age weight
w(i) = λ^(i−1) × (1 − λ) ÷ (1 − λ^n)
i = 1 is the most recent day, n is the window length. The weights sum to 1. Each older day's weight is λ times the one before.
Weight ratio
w(i+1) ÷ w(i) = λ
Useful shortcut: compute w(1), then multiply by λ repeatedly.
Hull-White volatility-adjusted return
r*(t,i) = r(t,i) × σ(T+1) ÷ σ(t,i)
r(t,i) is the historical return, σ(t,i) the volatility estimated for that day, σ(T+1) the current volatility forecast. Each observation keeps probability 1/n.
Weighted VaR rule
VaR = loss at which cumulative weight, from the worst loss down, first reaches 1 − confidence level
Sort losses from largest to smallest and add their weights. Some texts interpolate between observations.
Filtered HS standardised residual
z(t) = r(t) ÷ σ(t)
σ(t) comes from the fitted volatility model (e.g. GARCH). Bootstrap the z values and multiply by the simulated σ for each future day.
Kernel density estimator
f̂(x) = (1 ÷ (n × h)) × Σ K((x − xᵢ) ÷ h)
Sum over all n observations xᵢ. h is the bandwidth. K is a kernel that is non-negative, symmetric and integrates to 1, so f̂ integrates to 1.
Triangular kernel
K(u) = 1 − |u| for |u| ≤ 1; K(u) = 0 otherwise
Only observations within h of x contribute. u = (x − xᵢ) ÷ h.
Epanechnikov kernel
K(u) = 0.75 × (1 − u²) for |u| ≤ 1; K(u) = 0 otherwise
Another bounded kernel. Weights fall as distance from x grows.
Gaussian kernel
K(u) = (1 ÷ √(2π)) × e^(−u² ÷ 2)
Never reaches zero, so every observation contributes a little to every x.
Plotting position (midpoint convention)
pᵢ = (i − 0.5) ÷ n
i is the rank of the observation when sorted from worst (smallest P&L) to best. Not i ÷ n.
Linear interpolation for a quantile
x_p = xᵢ + ((p − pᵢ) ÷ (pᵢ₊₁ − pᵢ)) × (xᵢ₊₁ − xᵢ)
Use the two neighbours with pᵢ ≤ p ≤ pᵢ₊₁. For VaR, p is the tail probability, for example 5% for 95% VaR.
Rank for a target tail probability
i* = p × n + 0.5
If i* is a whole number, that observation is the quantile. Otherwise interpolate between ranks ⌊i*⌋ and ⌊i*⌋ + 1, using the fractional part as the weight.
Rule-of-thumb bandwidth (Gaussian kernel)
h ≈ 1.06 × σ × n^(−1/5)
Silverman's rule. It suits data that are roughly normal. For fat-tailed data it tends to over-smooth.
Historical simulation VaR
VaR at confidence c = the loss at the (1 − c) quantile of the sorted P&L sample
With n observations at 99%, the VaR is near the (0.01 × n)th worst loss. Interpolation conventions vary, so follow the question.
Expected tail observations
Observations beyond VaR ≈ n × (1 − c)
For n = 500 at 99%, only about 5 observations lie in the tail. This shows why precision is poor.
Ghost effect timing
An observation affects VaR for exactly n days, then drops out
n is the window length. The VaR jump occurs on entry and again on exit.
Equal weighting
Each observation has weight 1 ÷ n
This is the standard historical simulation. Weighted variants replace it with declining weights.

Quick revision

  • Historical simulation applies past returns to today's portfolio and reads VaR from the sorted P&L.
  • It assumes no distribution, so fat tails and skew in the data are kept.
  • Basic HS gives all observations in the window equal weight.
  • Expected shortfall is the average loss beyond VaR, so it is never below VaR at the same confidence level.
  • Bootstrap resamples with replacement from the same data and averages the VaR estimates.
  • Bootstrap can also give a confidence interval for VaR and tends to make the estimate more stable.
  • Age-weighted HS uses weights that decay by λ, so recent data counts more and ghost effects fall.
  • Weights must sum to 1 before you find the tail quantile.
  • Volatility-weighted HS rescales past returns by current volatility relative to volatility at the time.
  • Kernel methods smooth the empirical distribution so quantiles are less sensitive to single points.
  • Main weakness: results depend heavily on the window, and nothing beyond the worst observed loss can be estimated.
  • Main strength: easy to explain and implement, with no distribution or correlation assumptions.

Common mistakes

  • Counting from the wrong end of the ranking Fix: Sort losses from worst to best and count k worst outcomes from the loss end.
  • Using 1 − c as the tail on the wrong side, e.g. 5% for a 99% VaR Fix: Tail probability = 1 − c. For 99%, the tail is 1%.
  • Resampling without replacement Fix: Remember that without replacement you would just reorder the same data and get the same VaR every time. Replacement is what creates variation.
  • Saying the bootstrap needs a normal distribution Fix: The bootstrap is non-parametric. It uses the empirical distribution and assumes no particular shape.
  • Giving the oldest observation the largest weight, or numbering days the wrong way. Fix: Label the most recent day as i = 1 before applying λ^(i−1). Check that weights fall as you go back in time.
  • Using weights that do not sum to 1. Fix: Always include the divisor, then add up your weights as a check.
  • Using i ÷ n instead of (i − 0.5) ÷ n as the cumulative probability of the i-th observation. Fix: When a question talks about observations as slice midpoints or interpolation, use (i − 0.5) ÷ n and rank = p × n + 0.5.
  • Forgetting the 1 ÷ (n × h) scaling when computing a kernel density, and just adding the K values. Fix: Write the full formula first. The density must integrate to 1, which is why you divide by n and by h.
  • Saying non-parametric methods make no assumptions at all. Fix: They make no distribution assumption, but they assume the past is representative and that returns are drawn from a stable process.
  • Believing a longer window is always better. Fix: Longer windows give more tail points but include stale data and respond slowly. It is a trade-off.

Exam tips

  • Read the convention in the question for the tail position. Different texts use the k-th or (k+1)-th worst loss, and options may be built to catch this.
  • Check that the answer is a positive currency amount at the right confidence level and horizon.
  • Questions often ask for advantages and disadvantages. Advantages: no distribution assumption, captures fat tails and non-linearity, simple. Disadvantages: equal weights, slow reaction, limited by window data, cannot exceed the worst observation.
  • When volatility has just risen, expect HS VaR to understate risk; when a crisis day leaves the window, expect VaR to drop suddenly.
  • Link to related fixes: weighted and bootstrap HS are the standard answers to HS weaknesses.
  • Expect conceptual MCQs. Know the three keywords: with replacement, same size, repeated many times.
  • Be ready to say what the bootstrap adds over plain historical simulation: a more precise estimate and a standard error or confidence interval.
  • Watch for distractors that say the bootstrap assumes normality or fixes volatility clustering. Both are wrong.