Risk Modelling and Survival Analysis · Introduction to extreme value theory
Peaks over Threshold and the Generalised Pareto Distribution
Updated 11 October 2026 · Fact-checked
Peaks over threshold (POT) models only the amounts by which losses exceed a high threshold u. For a high enough u, these exceedances are approximately generalised Pareto distributed, with shape ξ and scale σ. You fit the GPD to the exceedances, then use it to estimate tail probabilities and high quantiles.
Understand Peaks over Threshold and Generalised Pareto Distribution
Extreme value theory studies the tail of a distribution, where the rare, large losses sit. Reinsurers and risk managers care about this region because it drives capital and pricing. Ordinary distributions fitted to all the data are often poor in the tail, because the bulk of the data dominates the fit.
The peaks over threshold method fixes a high threshold u. You keep only the observations above u and look at the exceedances (excesses) Y = X − u, for X > u. Every large observation is used, not just one per period.
A key result (Pickands–Balkema–de Haan) says that for a wide class of distributions, as u becomes high, the distribution of the excess Y given X > u is approximated by the generalised Pareto distribution (GPD). It has a shape parameter ξ and a scale parameter σ > 0. The shape ξ controls tail heaviness. If ξ > 0 the tail is heavy (Pareto-like). If ξ = 0 it is exponential. If ξ < 0 the tail has a finite upper endpoint.
The block maxima method splits data into blocks (for example years) and models the maximum of each block with the generalised extreme value (GEV) distribution. POT uses all exceedances above u, so it usually uses data more efficiently. Block maxima throws away every large value except the biggest in each block. The shape parameter ξ is the same in both approaches when the underlying data are in the same domain of attraction.
Choosing the threshold is a trade-off. Too low, and the GPD approximation is poor, so the estimates are biased. Too high, and few exceedances remain, so the estimates have high variance. Common tools are the mean excess plot (look for the region where it is roughly linear) and parameter stability plots (look for ξ and the modified scale to settle down as u rises).
Key rules to remember
- Excess over threshold
- Y = X − u, given X > u
- Only observations above u are used. Subtract u before fitting.
- GPD distribution function
- H(y) = 1 − (1 + ξy/σ)^(−1/ξ) for ξ ≠ 0; H(y) = 1 − e^(−y/σ) for ξ = 0
- Valid for y ≥ 0 and 1 + ξy/σ > 0. If ξ < 0, y is bounded above by −σ/ξ.
- Tail probability of the original variable
- P(X > x) ≈ (N_u ÷ n) × (1 + ξ(x − u)/σ)^(−1/ξ), for x > u
- N_u is the number of exceedances and n the total sample size. N_u ÷ n estimates P(X > u).
- High quantile (value at risk)
- x_p = u + (σ/ξ) × [ ((n ÷ N_u) × (1 − p))^(−ξ) − 1 ], for ξ ≠ 0
- Valid for p high enough that x_p > u. For ξ = 0 use x_p = u + σ ln[(N_u ÷ n) ÷ (1 − p)].
- Mean of the GPD
- E[Y] = σ ÷ (1 − ξ), for ξ < 1
- The mean does not exist for ξ ≥ 1.
- Mean excess function
- e(u) = E[X − u | X > u]; for a GPD, e(u) = (σ + ξu) ÷ (1 − ξ) when ξ < 1
- For GPD data this is linear in u. A positive slope suggests ξ > 0. Here σ is the scale at threshold zero; the scale at u is σ + ξu.
- Sample mean excess
- ê(u) = Σ(xᵢ − u) over xᵢ > u, ÷ N_u
- Plot ê(u) against u to choose the threshold.
How to solve Peaks over Threshold and Generalised Pareto Distribution questions
Use this order for any POT or GPD question. State your assumptions as you go.
- 1Identify the threshold u and the data. Count n (all observations) and N_u (those above u).
- 2Form the excesses yᵢ = xᵢ − u for each xᵢ > u.
- 3Choose or confirm the model: GPD for the excesses, with parameters ξ and σ. If ξ and σ are not given, estimate them (maximum likelihood, or method of moments if the question asks).
- 4Write the formula you need in standard notation: tail probability, quantile or mean excess. Check its conditions, such as ξ ≠ 0 or ξ < 1.
- 5Substitute carefully. Remember to use N_u ÷ n for the probability of exceeding u.
- 6Compute the answer and check it is sensible: probabilities between 0 and 1, and quantiles above u.
- 7If asked about threshold choice, comment on the bias and variance trade-off and the evidence from the mean excess plot.
- 8State the answer with units (for example ₹ lakh) and mention key assumptions: independent, identically distributed exceedances and a threshold high enough for the GPD approximation.
Quickest way: Plug-in method for tail probabilities and quantiles
When to use it: Use when ξ, σ, u, N_u and n are given and you need a tail probability, quantile or mean excess.
- Write down u, ξ, σ, N_u and n in one line.
- Compute the exceedance probability at u: N_u ÷ n.
- For a probability, find 1 + ξ(x − u)/σ, raise it to the power −1/ξ, then multiply by N_u ÷ n.
- For a quantile, solve (N_u ÷ n)(1 + ξ(x − u)/σ)^(−1/ξ) = 1 − p for x. This gives the quantile formula.
- Check x > u. If not, the formula does not apply and you must use the body of the distribution.
Common mistakes in Peaks over Threshold and Generalised Pareto Distribution
Fitting the GPD to the raw losses instead of the excesses.
Students forget that the GPD describes X − u, not X.
Fix: Always subtract u first. Add u back when you convert to a quantile of X.
Using the GPD tail probability without multiplying by N_u ÷ n.
The GPD gives P(X − u > y | X > u), which is conditional.
Fix: Multiply by P(X > u), estimated as N_u ÷ n, to get the unconditional tail probability.
Choosing the threshold too low, or too high, without justification.
Students treat the threshold as arbitrary or pick the lowest value to keep more data.
Fix: State the trade-off: low u gives bias, high u gives high variance. Support the choice with a mean excess plot that is roughly linear above u.
Confusing the sign of ξ and tail heaviness.
The GEV and GPD both use ξ and the signs are easy to mix up.
Fix: Remember: ξ > 0 heavy tail, ξ = 0 exponential tail, ξ < 0 bounded upper endpoint at u − σ/ξ.
Stating that POT and block maxima are the same method.
Both belong to extreme value theory and both give ξ.
Fix: Block maxima: one maximum per block, GEV distribution. POT: all values above u, GPD for excesses.
Using the mean formula σ ÷ (1 − ξ) when ξ ≥ 1.
Students apply formulas without checking conditions.
Fix: The mean exists only if ξ < 1. Say so if ξ is 1 or more.
Worked examples
Example 1
A portfolio has 500 claims. A threshold of ₹10 lakh is exceeded by 50 claims. The excesses are modelled by a GPD with ξ = 0.25 and σ = ₹4 lakh. Estimate the probability that a claim exceeds ₹20 lakh.
Show the solution
- u = 10, N_u = 50, n = 500, so N_u ÷ n = 0.1.
- x − u = 20 − 10 = 10.
- 1 + ξ(x − u)/σ = 1 + 0.25 × 10 ÷ 4 = 1 + 0.625 = 1.625.
- Exponent −1/ξ = −4.
- 1.625² = 2.640625, and 1.625⁴ = 2.640625² = 6.97290...
- So 1.625^(−4) = 1 ÷ 6.97290 = 0.14341.
- P(X > 20) ≈ 0.1 × 0.14341 = 0.01434.
Answer: About 0.0143, or roughly 1.4%.
Example 2
Using the same data (u = ₹10 lakh, N_u = 50, n = 500, ξ = 0.25, σ = ₹4 lakh), estimate the 99.5% quantile of the claim size. Also explain why a threshold that is too low would be a problem.
Show the solution
- 1 − p = 0.005. (n ÷ N_u) × (1 − p) = 10 × 0.005 = 0.05.
- Raise to the power −ξ = −0.25: 0.05^(−0.25) = 20^(0.25).
- 20^0.5 = 4.47214, and 4.47214^0.5 = 2.11474. So 20^0.25 ≈ 2.1147.
- Subtract 1: 2.1147 − 1 = 1.1147.
- σ/ξ = 4 ÷ 0.25 = 16. Then 16 × 1.1147 = 17.835.
- x_p = u + 17.835 = 10 + 17.835 = 27.835.
- Check: 27.835 > 10, so the formula applies.
- Threshold comment: a threshold that is too low includes observations from the body of the distribution, where the GPD approximation does not hold. This biases the estimates of ξ and σ.
Answer: The 99.5% quantile is about ₹27.8 lakh. A low threshold causes bias in the fitted GPD.
Exam tips
- Write the GPD formula in standard notation and state its conditions (y ≥ 0, ξ ≠ 0, ξ < 1 for the mean). Marks are given for method.
- Always show N_u ÷ n explicitly. Examiners look for the conditional versus unconditional step.
- For threshold choice, give both sides of the bias–variance trade-off and name the mean excess plot. A one-word answer rarely earns full marks.
- In the computer-based paper, fit the GPD to the excesses after subtracting u. State the threshold, the number of exceedances and the fitted ξ and σ before you interpret.
- For block maxima versus POT questions, give one point on data use and one on the distribution used (GEV versus GPD).
Practice questions from Introduction to extreme value theory
- When choosing the threshold u in a POT analysis, what is the main trade-off?
- Block maxima of a loss variable are fitted with a generalised extreme value distribution with shape parameter ξ. For which value of ξ is the…
- Annual maximum flood losses at a site are modelled by a GEV distribution with μ = 100, σ = 20 and ξ = 0 (Gumbel), in Rs crore. What is the p…
- In the peaks-over-threshold method, an analyst chooses a threshold that is set too low. What is the most likely consequence?
- X1, ..., Xn are independent Exponential variables with rate 1, and Mn is their maximum. Which statement about the limiting behaviour of Mn -…
Peaks over Threshold and Generalised Pareto Distribution in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Peaks over Threshold and Generalised Pareto Distribution: frequently asked questions
What is the difference between block maxima and peaks over threshold?
Block maxima takes the largest value in each period and fits the GEV distribution. Peaks over threshold takes every value above a high threshold and fits the GPD to the excesses. POT usually uses the data more efficiently because it keeps all extreme observations.
How do I choose the threshold in peaks over threshold?
Pick the lowest threshold above which the mean excess plot looks roughly linear and the fitted ξ is stable. A lower threshold gives more data but more bias. A higher threshold gives less bias but more variance. State this trade-off in your answer.
What does the shape parameter ξ tell me?
It describes the tail. If ξ > 0 the tail is heavy and polynomial, like a Pareto. If ξ = 0 the tail is exponential. If ξ < 0 the distribution has a finite upper limit.
Why is the GPD used for exceedances?
A limit theorem shows that, for a wide class of distributions, the excess over a high threshold is approximately GPD as the threshold rises. This gives a standard model for tail risk without knowing the full distribution.