Skip to content

FRM Exam Part II · Non-parametric Approaches

Non-parametric Density Estimation and Kernels for VaR

Updated 11 October 2026 · Fact-checked

Non-parametric density estimation smooths the jagged empirical distribution of returns without assuming a normal or any other formula. Histograms, kernels and interpolation turn discrete observations into a continuous density, so you can read VaR at any percentile. To solve: rank the data, assign each point probability (i − 0.5) ÷ n, then interpolate linearly to the target percentile.

Understand Non-parametric Density Estimation and Kernels

Historical simulation uses the raw data as the distribution. That distribution is a set of spikes, one per observation. It has gaps. If you have 40 observations and want the 5% quantile, there is no single data point sitting exactly there.

Non-parametric density estimation fixes this by smoothing. You do not assume a normal or lognormal shape. You let the data speak, but you spread each observation's probability over a small neighbourhood so the result is a continuous curve.

The simplest smoother is a histogram. You split the return range into bins and count observations per bin. Bin width matters. Narrow bins give a spiky picture. Wide bins hide the shape, especially in the tails where you need it.

A kernel is a smoother histogram. You place a small symmetric bump (the kernel function) on each observation and add the bumps up. The bandwidth h sets how wide each bump is. A small h tracks the data closely but is noisy. A large h is smooth but blurs real features such as fat tails. The choice of bandwidth usually matters more than the choice of kernel shape.

Interpolation is the practical tool for VaR. Treat each ordered observation as sitting at the midpoint of its own probability slice, so the i-th smallest of n observations has cumulative probability (i − 0.5) ÷ n. Then draw a straight line between neighbouring points and read off the loss at the exact confidence level you need. This is one common convention for the plotting position. Note that interpolation works only inside the data range. It cannot tell you about losses worse than the worst observation.

Key formulas to remember

Kernel density estimator
f̂(x) = (1 ÷ (n × h)) × Σ K((x − xᵢ) ÷ h)
Sum over all n observations xᵢ. h is the bandwidth. K is a kernel that is non-negative, symmetric and integrates to 1, so f̂ integrates to 1.
Triangular kernel
K(u) = 1 − |u| for |u| ≤ 1; K(u) = 0 otherwise
Only observations within h of x contribute. u = (x − xᵢ) ÷ h.
Epanechnikov kernel
K(u) = 0.75 × (1 − u²) for |u| ≤ 1; K(u) = 0 otherwise
Another bounded kernel. Weights fall as distance from x grows.
Gaussian kernel
K(u) = (1 ÷ √(2π)) × e^(−u² ÷ 2)
Never reaches zero, so every observation contributes a little to every x.
Plotting position (midpoint convention)
pᵢ = (i − 0.5) ÷ n
i is the rank of the observation when sorted from worst (smallest P&L) to best. Not i ÷ n.
Linear interpolation for a quantile
x_p = xᵢ + ((p − pᵢ) ÷ (pᵢ₊₁ − pᵢ)) × (xᵢ₊₁ − xᵢ)
Use the two neighbours with pᵢ ≤ p ≤ pᵢ₊₁. For VaR, p is the tail probability, for example 5% for 95% VaR.
Rank for a target tail probability
i* = p × n + 0.5
If i* is a whole number, that observation is the quantile. Otherwise interpolate between ranks ⌊i*⌋ and ⌊i*⌋ + 1, using the fractional part as the weight.
Rule-of-thumb bandwidth (Gaussian kernel)
h ≈ 1.06 × σ × n^(−1/5)
Silverman's rule. It suits data that are roughly normal. For fat-tailed data it tends to over-smooth.

How to solve Non-parametric Density Estimation and Kernels questions

Use this order for any question on smoothing, kernels or interpolated VaR.

  1. 1Identify what is asked: a density value f̂(x), a bandwidth effect, or a VaR percentile.
  2. 2For a VaR question, sort the P&L from worst to best and number the observations 1 to n.
  3. 3Convert the tail probability to a rank: i* = p × n + 0.5. For 95% VaR use p = 5%; for 99% use p = 1%.
  4. 4If i* is a whole number, read the quantile straight from that observation. If not, find the two neighbours and compute the weight as the fractional part of i*.
  5. 5Interpolate: x_p = lower value + weight × (upper value − lower value). Then flip the sign if the question asks for VaR as a positive loss.
  6. 6For a kernel density value, compute u = (x − xᵢ) ÷ h for each observation, apply K(u), add them up, then divide by n × h.
  7. 7Check the answer: it should lie between the two neighbouring observations, and the density must be non-negative.
  8. 8If the question is conceptual, state the trade-off: small h or narrow bins give low bias but high noise; large h or wide bins give smooth curves but hide tails.

Quickest way: Rank-and-weight shortcut for interpolated VaR

When to use it: Use when the question gives a short list of ordered P&L values and asks for VaR at a stated confidence level.

  1. Compute i* = p × n + 0.5 in your head. For n = 40 and p = 5%, i* = 2.5.
  2. Take the whole part as the lower rank (2) and the fractional part as the weight (0.5).
  3. Move from the lower value towards the next value by weight × the gap. At weight 0.5, take the midpoint.
  4. Answer is that point. Report VaR as a positive loss unless told otherwise.
  5. Eliminate options outside the range of the two neighbours before calculating.

Common mistakes in Non-parametric Density Estimation and Kernels

  • Using i ÷ n instead of (i − 0.5) ÷ n as the cumulative probability of the i-th observation.

    i ÷ n is the familiar empirical CDF. The midpoint convention in these readings is less familiar.

    Fix: When a question talks about observations as slice midpoints or interpolation, use (i − 0.5) ÷ n and rank = p × n + 0.5.

  • Forgetting the 1 ÷ (n × h) scaling when computing a kernel density, and just adding the K values.

    Students focus on the kernel values and treat the sum as the answer.

    Fix: Write the full formula first. The density must integrate to 1, which is why you divide by n and by h.

  • Thinking a larger bandwidth always gives a better estimate.

    Smoother pictures look cleaner, so students assume they are more accurate.

    Fix: Remember the trade-off. Too large an h blurs tails and adds bias. Too small an h makes the tail noisy. Neither is free.

  • Believing interpolation or kernels can estimate VaR beyond the worst observation.

    Smoothing creates a continuous curve, which looks like it covers everything.

    Fix: Interpolation stays inside the data range. For very extreme quantiles you need a tail model such as extreme value theory.

  • Interpolating with the weight the wrong way round, so the answer sits nearer the wrong neighbour.

    Students mix up which observation is the lower and which is the upper, especially with negative P&L.

    Fix: Sort from worst to best. The weight is the fractional part of i*, applied as lower value + weight × (upper − lower). Check the result lies between the two.

  • Treating the kernel shape as the main driver of results.

    Kernel names such as Epanechnikov or Gaussian sound important.

    Fix: In practice the bandwidth drives the outcome far more than the kernel shape. Say this in conceptual answers.

Worked examples

Example 1

A risk manager has 40 daily P&L observations in USD, sorted from worst to best. The second-worst is −4.2 million and the third-worst is −3.6 million. Using the midpoint convention pᵢ = (i − 0.5) ÷ n and linear interpolation, what is the 95% one-day VaR? A) 3.6 million B) 3.9 million C) 4.2 million D) 4.5 million

Show the solution
  1. Tail probability for 95% VaR: p = 5% = 0.05.
  2. Cumulative probability of observation 2: (2 − 0.5) ÷ 40 = 0.0375.
  3. Cumulative probability of observation 3: (3 − 0.5) ÷ 40 = 0.0625.
  4. The target 0.05 lies between them. Weight = (0.05 − 0.0375) ÷ (0.0625 − 0.0375) = 0.0125 ÷ 0.025 = 0.5.
  5. Interpolated P&L = −4.2 + 0.5 × (−3.6 − (−4.2)) = −4.2 + 0.5 × 0.6 = −3.9 million.
  6. VaR is reported as a positive loss: 3.9 million.

Answer: B) 3.9 million

Example 2

Three observed returns are 1%, 2% and 4%. A triangular kernel K(u) = 1 − |u| for |u| ≤ 1 is used with bandwidth h = 2 (in percentage points). What is the estimated density f̂ at x = 2.5? A) 0.125 B) 0.208 C) 0.250 D) 0.417

Show the solution
  1. Formula: f̂(x) = (1 ÷ (n × h)) × Σ K((x − xᵢ) ÷ h), with n = 3 and h = 2, so n × h = 6.
  2. For xᵢ = 1: u = (2.5 − 1) ÷ 2 = 0.75, K = 1 − 0.75 = 0.25.
  3. For xᵢ = 2: u = (2.5 − 2) ÷ 2 = 0.25, K = 1 − 0.25 = 0.75.
  4. For xᵢ = 4: u = (2.5 − 4) ÷ 2 = −0.75, |u| = 0.75, K = 0.25.
  5. Sum of kernel values = 0.25 + 0.75 + 0.25 = 1.25.
  6. f̂(2.5) = 1.25 ÷ 6 = 0.2083, which is about 0.208.

Answer: B) 0.208

Exam tips

  • Questions are usually short calculations or conceptual choices. Practise the rank formula i* = p × n + 0.5 until it is automatic.
  • Read the convention in the question. If it says observations are midpoints of probability slices, use (i − 0.5) ÷ n.
  • For bandwidth questions, answer with the bias versus noise trade-off. Small h means a spiky, noisy estimate. Large h means oversmoothing.
  • Watch the sign. P&L is negative in the tail, but VaR is quoted as a positive loss.
  • If an option sits outside the two neighbouring observations in an interpolation question, drop it at once.

Practice questions from Non-parametric Approaches

Non-parametric Density Estimation and Kernels: frequently asked questions

What is non-parametric density estimation in risk management?

It estimates the shape of the return distribution directly from data, with no assumed formula such as the normal. Histograms and kernels smooth the data into a continuous density. You then read percentiles and VaR from that smoothed distribution.

How do I estimate VaR between two observations?

Sort the P&L from worst to best and assign each observation a cumulative probability of (i − 0.5) ÷ n. Find the two observations whose probabilities surround your tail probability. Interpolate linearly between their values and report the result as a positive loss.

What does the bandwidth do in kernel density estimation?

The bandwidth h sets how wide the bump on each observation is. A small h follows the data closely but gives a noisy curve. A large h gives a smooth curve but can blur features such as fat tails.

Can kernel smoothing fix the lack of data in the extreme tail?

No. Smoothing spreads probability around the observed points but does not add information beyond the worst observation. For very high confidence levels you need a tail model such as extreme value theory.