FRM Part II · FRM Exam Part II
Non-parametric Approaches: formula sheet
Key formulas
- Historical return (simple)
- R_t = (P_t − P_(t−1)) ÷ P_(t−1)
- Use the same return definition consistently. Log returns are also common: ln(P_t ÷ P_(t−1)).
- Scenario P&L
- P&L_t = V × R_t
- Apply each historical return to today's portfolio value V (or revalue the portfolio under each scenario's risk-factor changes).
- Tail count
- k = n × (1 − c)
- Number of observations in the tail. For n = 500 and c = 99%, k = 5. Under the convention used on this page, VaR is the (k+1)-th worst outcome, so k outcomes are worse than the VaR loss. Other conventions interpolate between the k-th and (k+1)-th worst outcomes; follow the question's convention.
- HS VaR
- VaR(c) = −(percentile of P&L at level 1 − c)
- Report as a positive loss. Equivalently, the loss ranked just beyond the tail count in the sorted list.
- Horizon scaling (rule of thumb)
- VaR(T days) ≈ VaR(1 day) × √T
- Valid only under i.i.d. returns with zero mean; it is an approximation, not an HS output.
- Bootstrap sample
- Draw n observations from n historical returns, with replacement
- Each sample has the same size as the original data. Repeats are expected.
- Bootstrapped VaR estimate
- VaR_boot = (1 ÷ B) × Σ VaR_b, for b = 1 to B
- B is the number of bootstrap samples. VaR_b is the VaR from sample b.
- Bootstrapped ES estimate
- ES_boot = (1 ÷ B) × Σ ES_b, for b = 1 to B
- ES_b is the average loss beyond VaR in sample b.
- Bootstrap standard error
- SE = √[ Σ (VaR_b − VaR_boot)² ÷ (B − 1) ]
- This is the standard deviation of the B estimates. Dividing by B instead of B − 1 is also seen when B is large.
- Approximate confidence interval
- VaR_boot ± z × SE
- Use z = 1.96 for 95% two-sided. Alternatively, use percentiles of the sorted bootstrap estimates.
- Historical VaR position
- VaR at confidence c is the loss at the (1 − c) tail percentile; with n observations, about n × (1 − c) losses lie beyond it
- Interpolation conventions vary, so read the question's stated method.
- BRW age weight
- w(i) = λ^(i−1) × (1 − λ) ÷ (1 − λ^n)
- i = 1 is the most recent day, n is the window length. The weights sum to 1. Each older day's weight is λ times the one before.
- Weight ratio
- w(i+1) ÷ w(i) = λ
- Useful shortcut: compute w(1), then multiply by λ repeatedly.
- Hull-White volatility-adjusted return
- r*(t,i) = r(t,i) × σ(T+1) ÷ σ(t,i)
- r(t,i) is the historical return, σ(t,i) the volatility estimated for that day, σ(T+1) the current volatility forecast. Each observation keeps probability 1/n.
- Weighted VaR rule
- VaR = loss at which cumulative weight, from the worst loss down, first reaches 1 − confidence level
- Sort losses from largest to smallest and add their weights. Some texts interpolate between observations.
- Filtered HS standardised residual
- z(t) = r(t) ÷ σ(t)
- σ(t) comes from the fitted volatility model (e.g. GARCH). Bootstrap the z values and multiply by the simulated σ for each future day.
- Kernel density estimator
- f̂(x) = (1 ÷ (n × h)) × Σ K((x − xᵢ) ÷ h)
- Sum over all n observations xᵢ. h is the bandwidth. K is a kernel that is non-negative, symmetric and integrates to 1, so f̂ integrates to 1.
- Triangular kernel
- K(u) = 1 − |u| for |u| ≤ 1; K(u) = 0 otherwise
- Only observations within h of x contribute. u = (x − xᵢ) ÷ h.
- Epanechnikov kernel
- K(u) = 0.75 × (1 − u²) for |u| ≤ 1; K(u) = 0 otherwise
- Another bounded kernel. Weights fall as distance from x grows.
- Gaussian kernel
- K(u) = (1 ÷ √(2π)) × e^(−u² ÷ 2)
- Never reaches zero, so every observation contributes a little to every x.
- Plotting position (midpoint convention)
- pᵢ = (i − 0.5) ÷ n
- i is the rank of the observation when sorted from worst (smallest P&L) to best. Not i ÷ n.
- Linear interpolation for a quantile
- x_p = xᵢ + ((p − pᵢ) ÷ (pᵢ₊₁ − pᵢ)) × (xᵢ₊₁ − xᵢ)
- Use the two neighbours with pᵢ ≤ p ≤ pᵢ₊₁. For VaR, p is the tail probability, for example 5% for 95% VaR.
- Rank for a target tail probability
- i* = p × n + 0.5
- If i* is a whole number, that observation is the quantile. Otherwise interpolate between ranks ⌊i*⌋ and ⌊i*⌋ + 1, using the fractional part as the weight.
- Rule-of-thumb bandwidth (Gaussian kernel)
- h ≈ 1.06 × σ × n^(−1/5)
- Silverman's rule. It suits data that are roughly normal. For fat-tailed data it tends to over-smooth.
- Historical simulation VaR
- VaR at confidence c = the loss at the (1 − c) quantile of the sorted P&L sample
- With n observations at 99%, the VaR is near the (0.01 × n)th worst loss. Interpolation conventions vary, so follow the question.
- Expected tail observations
- Observations beyond VaR ≈ n × (1 − c)
- For n = 500 at 99%, only about 5 observations lie in the tail. This shows why precision is poor.
- Ghost effect timing
- An observation affects VaR for exactly n days, then drops out
- n is the window length. The VaR jump occurs on entry and again on exit.
- Equal weighting
- Each observation has weight 1 ÷ n
- This is the standard historical simulation. Weighted variants replace it with declining weights.
Quick revision
- Historical simulation applies past returns to today's portfolio and reads VaR from the sorted P&L.
- It assumes no distribution, so fat tails and skew in the data are kept.
- Basic HS gives all observations in the window equal weight.
- Expected shortfall is the average loss beyond VaR, so it is never below VaR at the same confidence level.
- Bootstrap resamples with replacement from the same data and averages the VaR estimates.
- Bootstrap can also give a confidence interval for VaR and tends to make the estimate more stable.
- Age-weighted HS uses weights that decay by λ, so recent data counts more and ghost effects fall.
- Weights must sum to 1 before you find the tail quantile.
- Volatility-weighted HS rescales past returns by current volatility relative to volatility at the time.
- Kernel methods smooth the empirical distribution so quantiles are less sensitive to single points.
- Main weakness: results depend heavily on the window, and nothing beyond the worst observed loss can be estimated.
- Main strength: easy to explain and implement, with no distribution or correlation assumptions.
Common mistakes
- Counting from the wrong end of the ranking Fix: Sort losses from worst to best and count k worst outcomes from the loss end.
- Using 1 − c as the tail on the wrong side, e.g. 5% for a 99% VaR Fix: Tail probability = 1 − c. For 99%, the tail is 1%.
- Resampling without replacement Fix: Remember that without replacement you would just reorder the same data and get the same VaR every time. Replacement is what creates variation.
- Saying the bootstrap needs a normal distribution Fix: The bootstrap is non-parametric. It uses the empirical distribution and assumes no particular shape.
- Giving the oldest observation the largest weight, or numbering days the wrong way. Fix: Label the most recent day as i = 1 before applying λ^(i−1). Check that weights fall as you go back in time.
- Using weights that do not sum to 1. Fix: Always include the divisor, then add up your weights as a check.
- Using i ÷ n instead of (i − 0.5) ÷ n as the cumulative probability of the i-th observation. Fix: When a question talks about observations as slice midpoints or interpolation, use (i − 0.5) ÷ n and rank = p × n + 0.5.
- Forgetting the 1 ÷ (n × h) scaling when computing a kernel density, and just adding the K values. Fix: Write the full formula first. The density must integrate to 1, which is why you divide by n and by h.
- Saying non-parametric methods make no assumptions at all. Fix: They make no distribution assumption, but they assume the past is representative and that returns are drawn from a stable process.
- Believing a longer window is always better. Fix: Longer windows give more tail points but include stale data and respond slowly. It is a trade-off.
Exam tips
- Read the convention in the question for the tail position. Different texts use the k-th or (k+1)-th worst loss, and options may be built to catch this.
- Check that the answer is a positive currency amount at the right confidence level and horizon.
- Questions often ask for advantages and disadvantages. Advantages: no distribution assumption, captures fat tails and non-linearity, simple. Disadvantages: equal weights, slow reaction, limited by window data, cannot exceed the worst observation.
- When volatility has just risen, expect HS VaR to understate risk; when a crisis day leaves the window, expect VaR to drop suddenly.
- Link to related fixes: weighted and bootstrap HS are the standard answers to HS weaknesses.
- Expect conceptual MCQs. Know the three keywords: with replacement, same size, repeated many times.
- Be ready to say what the bootstrap adds over plain historical simulation: a more precise estimate and a standard error or confidence interval.
- Watch for distractors that say the bootstrap assumes normality or fixes volatility clustering. Both are wrong.