Actuarial Statistics · Bayesian inference: priors, posteriors, loss functions and credible intervals
Prior Distributions and Conjugate Priors for Actuarial Statistics
Updated 11 October 2026 · Fact-checked
A prior distribution describes what you believe about a parameter before seeing data. A conjugate prior is one that gives a posterior from the same family. To solve questions, write posterior ∝ likelihood × prior, keep only terms in the parameter, and match the kernel to a known distribution.
Understand Prior Distributions and Conjugate Priors
In Bayesian statistics the parameter θ is treated as random. The prior distribution f(θ) states your beliefs about θ before you see data. After you observe data x, you update the prior to the posterior f(θ | x) using Bayes' theorem.
An informative prior carries real knowledge, for example past claims experience or expert opinion. A non-informative prior (also called vague or diffuse) tries to let the data dominate. Examples are a uniform prior, or an improper prior such as f(θ) ∝ 1 on the real line. An improper prior does not integrate to 1, but it can still give a proper posterior. Check that it does.
A prior is conjugate to a likelihood if the posterior is in the same family as the prior. This is convenient because the update only changes the parameters. You do not need to do any integration. The common pairs are Beta-Binomial, Gamma-Poisson and Normal-Normal (known variance).
The key idea is the kernel. The posterior is proportional to likelihood × prior, so you can drop every factor that does not involve θ. What remains is the kernel. If the kernel looks like a known density, the normalising constant follows automatically, and you can name the posterior without integrating.
The prior parameters often act like pseudo-data. In Beta(α, β), α behaves like prior successes and β like prior failures. The posterior mean is a weighted average of the prior mean and the data estimate, which links directly to credibility theory.
Key rules to remember
- Bayes' theorem for densities
- f(θ | x) = f(x | θ) f(θ) ÷ ∫ f(x | θ) f(θ) dθ
- In practice write f(θ | x) ∝ f(x | θ) f(θ) and identify the kernel.
- Beta-Binomial
- Prior θ ~ Beta(α, β); data X | θ ~ Bin(n, θ); posterior θ | x ~ Beta(α + x, β + n − x)
- Prior mean = α ÷ (α + β). Posterior mean = (α + x) ÷ (α + β + n).
- Beta density kernel
- f(θ) ∝ θ^(α−1) (1 − θ)^(β−1), 0 < θ < 1
- Use this to recognise the Beta in a posterior.
- Gamma-Poisson
- Prior λ ~ Gamma(α, β) with density ∝ λ^(α−1) e^(−βλ); data X1..Xn | λ iid Poisson(λ); posterior λ | x ~ Gamma(α + Σxᵢ, β + n)
- Here β is the rate. Prior mean = α ÷ β. Posterior mean = (α + Σxᵢ) ÷ (β + n).
- Normal-Normal (known variance)
- Prior μ ~ N(μ0, τ²); X1..Xn | μ iid N(μ, σ²) with σ² known; posterior μ | x ~ N(μ1, τ1²), where 1 ÷ τ1² = 1 ÷ τ² + n ÷ σ² and μ1 = τ1² (μ0 ÷ τ² + n x̄ ÷ σ²)
- Precision (1 ÷ variance) adds. The posterior mean is a precision-weighted average of μ0 and x̄.
- Credibility form of posterior mean
- Posterior mean = Z × (data estimate) + (1 − Z) × (prior mean)
- For Gamma-Poisson, Z = n ÷ (β + n) with data estimate x̄. For Normal-Normal, Z = (n ÷ σ²) ÷ (n ÷ σ² + 1 ÷ τ²).
- Improper uniform prior
- f(θ) ∝ 1 (non-informative)
- Gives posterior ∝ likelihood. Beta(1, 1) is the proper uniform prior on (0, 1).
How to solve Prior Distributions and Conjugate Priors questions
Use this method for any conjugate prior question. It works even when the pair is not one you have memorised.
- 1Write down the likelihood for the data, given the parameter. For a sample, multiply the individual densities or probabilities.
- 2Write down the prior density of the parameter. Note the parameterisation, such as rate or scale for the Gamma.
- 3Multiply them: posterior ∝ likelihood × prior. Drop every factor that does not contain the parameter.
- 4Collect powers of the parameter and exponential terms. For the Normal, complete the square in the exponent.
- 5Match the kernel to a known family and read off the posterior parameters.
- 6Answer the actual question: posterior mean, variance, a probability, or an estimate under a stated loss function.
- 7State the result in words and check it is sensible. For example, the posterior mean should lie between the prior mean and the data estimate.
Quickest way: Add the data to the prior parameters
When to use it: Use when the pair is Beta-Binomial, Gamma-Poisson or Normal-Normal and the question only asks for posterior parameters or a posterior mean. Under time pressure, state the kernel in one line to earn method marks.
- Beta-Binomial: add the successes to α and the failures to β.
- Gamma-Poisson: add Σxᵢ to α and the number of observations to the rate β.
- Normal-Normal: add precisions, then take the precision-weighted mean.
- Compute the posterior mean using the credibility form if a Z value is helpful.
- Write one line showing posterior ∝ likelihood × prior and the kernel, then quote the result.
Common mistakes in Prior Distributions and Conjugate Priors
Using the Gamma scale parameter as the rate (or the other way round) in Gamma-Poisson.
Different books parameterise Gamma differently. Students memorise 'add n' without checking which parameter it goes to.
Fix: Write the density first. If it is ∝ λ^(α−1) e^(−βλ), β is a rate and you add n to it. If it is e^(−λ/θ), the scale is θ and 1 ÷ θ becomes 1 ÷ θ + n.
Adding the number of trials to β in Beta-Binomial instead of the number of failures.
Students copy the Gamma-Poisson pattern of adding n.
Fix: Derive the exponents: θ^(α+x−1) (1−θ)^(β+n−x−1). So β becomes β + n − x.
Adding variances instead of precisions in Normal-Normal.
Adding variances is correct for independent sums, so students apply it here by habit.
Fix: Posterior precision = prior precision + data precision. Then invert to get the variance.
Keeping constants and factors without the parameter and then failing to recognise the kernel.
Students are afraid of dropping terms and make the algebra messy.
Fix: Anything not involving θ goes into the constant of proportionality. Only the shape in θ matters.
Calling any flat prior harmless without checking the posterior is proper.
Students treat non-informative as always safe.
Fix: With an improper prior, check that the posterior integrates to a finite value. Also note that a flat prior is not flat after a transformation of the parameter.
Confusing the Poisson likelihood sum with the sample mean.
The data estimate appears as x̄ in the credibility form but as Σxᵢ in the parameter update.
Fix: Update with Σxᵢ and n. Convert to x̄ only when writing the credibility weights.
Worked examples
Example 1
An insurer models the probability θ that a claim is rejected as Beta(2, 8) a priori. In a sample of 20 claims, 6 are rejected. Find the posterior distribution, the prior mean and the posterior mean of θ.
Show the solution
- Likelihood: X | θ ~ Bin(20, θ), so f(x | θ) ∝ θ^6 (1 − θ)^14.
- Prior kernel: θ^(2−1) (1 − θ)^(8−1) = θ^1 (1 − θ)^7.
- Posterior ∝ θ^(6+1) (1 − θ)^(14+7) = θ^7 (1 − θ)^21.
- This is the kernel of Beta(8, 22). Check: α = 2 + 6 = 8 and β = 8 + 20 − 6 = 22.
- Prior mean = 2 ÷ (2 + 8) = 0.2.
- Posterior mean = 8 ÷ (8 + 22) = 8 ÷ 30 = 0.2667 (to 4 decimal places).
Answer: Posterior is Beta(8, 22). Prior mean = 0.2. Posterior mean = 0.2667.
Example 2
The number of claims per year on a policy is Poisson(λ). The prior for λ is Gamma(α = 3, β = 2), where β is the rate. Over 4 years the claim counts are 1, 0, 2 and 1. (a) Find the posterior distribution of λ. (b) Show that the posterior mean is a credibility estimate and find Z.
Show the solution
- Likelihood: f(x | λ) ∝ λ^(Σxᵢ) e^(−nλ). Here Σxᵢ = 1 + 0 + 2 + 1 = 4 and n = 4, so the likelihood ∝ λ^4 e^(−4λ).
- Prior kernel: λ^(3−1) e^(−2λ) = λ^2 e^(−2λ).
- Posterior ∝ λ^(4+2) e^(−(4+2)λ) = λ^6 e^(−6λ).
- This is the kernel of Gamma(7, 6). Check: α = 3 + 4 = 7 and β = 2 + 4 = 6.
- Posterior mean = 7 ÷ 6 = 1.1667.
- Prior mean = 3 ÷ 2 = 1.5. Sample mean x̄ = 4 ÷ 4 = 1.
- Z = n ÷ (β + n) = 4 ÷ 6 = 2/3.
- Check: Z x̄ + (1 − Z) × prior mean = (2/3)(1) + (1/3)(1.5) = 0.6667 + 0.5 = 1.1667. This matches.
Answer: (a) Posterior is Gamma(7, 6). (b) Posterior mean = 7/6 ≈ 1.1667 = Z x̄ + (1 − Z) × 1.5 with Z = 2/3.
Exam tips
- In written questions, always show the line posterior ∝ likelihood × prior and name the kernel. Method marks are given for this step even if the arithmetic goes wrong.
- Check the parameterisation of the Gamma, and of the Exponential, before you update. The IAI formulae book and the question wording set which one applies.
- Expect MCQs asking you to identify the posterior or its mean from a short data summary. Use the add-the-data shortcut and do the check that the posterior mean lies between the prior mean and the data estimate.
- For the Normal-Normal case, work with precisions. Write the weights explicitly so that you can link the answer to credibility if the question asks.
- Be ready to explain in words why a prior is informative or non-informative, and to comment on how much the data overrides the prior as n grows.
Practice questions from Bayesian inference: priors, posteriors, loss functions and credible intervals
- The prior for a proportion p is Beta(a, b) with prior mean 0.25 and a + b = 12. After observing 6 successes in 18 trials, what is the poster…
- A posterior distribution for a parameter θ is Normal with mean 40 and variance 25. What is the equal-tailed 95% Bayesian credible interval f…
- In the Bayesian framework, the posterior density of a parameter θ given data x is proportional to which of the following?
- A prior distribution for a parameter θ is combined with data x. Which statement correctly describes the Bayesian relationship between the po…
- The posterior of θ is discrete: P(θ=1)=0.2, P(θ=2)=0.5, P(θ=4)=0.3. Under all-or-nothing (0-1) loss, the Bayes estimate of θ is:
Prior Distributions and Conjugate Priors: frequently asked questions
What is a conjugate prior in Bayesian statistics?
A conjugate prior is a prior distribution that gives a posterior in the same family when combined with a given likelihood. For example, a Beta prior with a Binomial likelihood gives a Beta posterior. This means the update only changes the parameters, with no integration needed.
What is the difference between an informative and a non-informative prior?
An informative prior carries real knowledge about the parameter and can strongly shape the posterior. A non-informative prior is chosen to have minimal influence so that the data dominate. Non-informative priors can be improper, so you must check that the posterior is proper.
How do I identify the posterior distribution from the kernel?
Multiply likelihood by prior and drop any factor that does not contain the parameter. Compare what is left with the density of known distributions, such as θ^(a−1)(1−θ)^(b−1) for the Beta. If the shape matches, the posterior is that distribution with those parameters.
Do I need to memorise the conjugate pairs for the IAI exam?
It is wise to know Beta-Binomial, Gamma-Poisson and Normal-Normal, since they save time. You should also be able to derive the posterior from scratch using the kernel method, because a question may use a less familiar likelihood.
Why does the posterior mean behave like a credibility estimate?
In the standard conjugate cases the posterior mean is a weighted average of the prior mean and the data estimate. The weight on the data grows as the sample size grows. This is the link between Bayesian inference and credibility theory.