Skip to content

Actuarial Statistics · Bayesian inference: priors, posteriors, loss functions and credible intervals

Posterior Distributions and Bayesian Estimators: How to Solve

Updated 11 October 2026 · Fact-checked

A posterior distribution is your updated belief about a parameter after seeing data. Multiply the likelihood by the prior, keep only terms involving the parameter, and recognise the standard distribution. Then choose the estimate that suits the loss function: mean for quadratic loss, median for absolute loss, mode for all-or-nothing loss.

Understand Posterior Distributions and Bayesian Estimators

In Bayesian inference the parameter θ is treated as random. Before you see data, you describe it with a prior distribution f(θ). After you see data x, you update it to the posterior distribution f(θ | x).

The update rule is Bayes' theorem in proportional form: posterior ∝ likelihood × prior. The denominator f(x) does not depend on θ, so you ignore it. You only keep terms that contain θ, then match what is left to a known distribution. This gives the normalising constant for free.

With a conjugate prior, the posterior is in the same family as the prior. Poisson data with a gamma prior gives a gamma posterior. Normal data with known variance and a normal prior gives a normal posterior. Binomial data with a beta prior gives a beta posterior. These are the cases the exam expects you to do quickly.

The posterior is a whole distribution. To get one number you need a Bayesian point estimate, and the right one depends on the loss function. Quadratic loss (θ̂ − θ)² gives the posterior mean. Absolute loss |θ̂ − θ| gives the posterior median. All-or-nothing loss (zero if exactly right, one otherwise) gives the posterior mode. For a symmetric single-peaked posterior such as the normal, all three are equal.

The posterior mean is always a weighted average of the sample estimate and the prior mean. This is the link to credibility theory. More data means more weight on the data and less on the prior.

Key rules to remember

Posterior proportionality
f(θ | x) ∝ f(x | θ) × f(θ)
Drop every factor that does not contain θ. Then identify the distribution from what remains.
Poisson likelihood, gamma prior
X₁,…,Xₙ ~ Poisson(λ), λ ~ Gamma(α, β) ⇒ λ | x ~ Gamma(α + Σxᵢ, β + n)
Here β is the rate, so the gamma mean is α/β. Check which parametrisation the question uses.
Gamma mean, variance and mode
Mean = α/β; Variance = α/β²; Mode = (α − 1)/β for α ≥ 1
Mode formula applies when α ≥ 1. For α < 1 the density has no interior mode.
Binomial likelihood, beta prior
X ~ Bin(n, θ), θ ~ Beta(α, β) ⇒ θ | x ~ Beta(α + x, β + n − x)
Beta mean is α/(α + β). Beta mode is (α − 1)/(α + β − 2) when α, β > 1.
Normal mean, known variance, normal prior
X₁,…,Xₙ ~ N(μ, σ²), μ ~ N(μ₀, σ₀²) ⇒ μ | x ~ N(μ₁, σ₁²), where 1/σ₁² = n/σ² + 1/σ₀² and μ₁ = (n x̄/σ² + μ₀/σ₀²) × σ₁²
Precisions (1 ÷ variance) add. The posterior mean is a precision-weighted average of x̄ and μ₀.
Credibility form of posterior mean
E(θ | x) = Z × (sample estimate) + (1 − Z) × (prior mean)
Poisson-gamma: Z = n/(n + β), sample estimate x̄. Normal: Z = (n/σ²) ÷ (n/σ² + 1/σ₀²).
Bayes estimates by loss function
Quadratic loss → posterior mean; Absolute loss → posterior median; All-or-nothing loss → posterior mode
State the loss function in your answer. The estimate depends on it.

How to solve Posterior Distributions and Bayesian Estimators questions

Use this method for any posterior-and-estimate question. It works for standard conjugate models and for non-standard ones.

  1. 1Write the model: the distribution of the data given θ, and the prior for θ. Note the parametrisation (rate or scale).
  2. 2Write the likelihood for the whole sample, L(θ) = Π f(xᵢ | θ). Simplify it, keeping only powers and exponentials of θ.
  3. 3Multiply by the prior density. Collect all terms in θ into a single expression such as θ^a e^(−bθ).
  4. 4Match this to a known distribution and read off the new parameters. If it matches nothing standard, integrate to find the constant.
  5. 5Check against the conjugate result, for example Gamma(α + Σx, β + n). A quick check catches algebra errors.
  6. 6Identify the loss function and take the matching summary: mean, median or mode of the posterior.
  7. 7If asked, give a credible interval using posterior quantiles, or write the estimate in credibility form Z x̄ + (1 − Z) × prior mean.
  8. 8State the final answer with the posterior distribution named in full, including its parameters.

Quickest way: Conjugate update shortcut

When to use it: Use it when the model is Poisson-gamma, binomial-beta or normal-normal with known variance. This covers most exam questions.

  1. Recognise the pair. Do not re-derive unless the question says 'show that' or asks for a proof.
  2. Update the parameters. Poisson-gamma: add Σx to α and n to β. Binomial-beta: add successes to α and failures to β. Normal-normal: add precisions.
  3. Compute the summary the question asks for, using the standard mean, mode or variance formula.
  4. Cross-check with the credibility form Z x̄ + (1 − Z) × prior mean. If both agree, you are safe.
  5. If the question says 'show that', write the likelihood × prior line and match the kernel. That earns the method marks.

Common mistakes in Posterior Distributions and Bayesian Estimators

  • Using x̄ instead of Σx when updating a gamma prior.

    Students remember 'add the data' but forget that the shape parameter takes the total count, while the rate takes n.

    Fix: Write Gamma(α + Σx, β + n) every time. Sanity-check that posterior mean (α + Σx)/(β + n) lies between x̄ and α/β.

  • Mixing up rate and scale for the gamma distribution.

    Some sources write Gamma(α, θ) with mean αθ. The IAI formulae use rate λ, with mean α/λ.

    Fix: Read the question's density or the Tables. Write the mean of the prior to confirm which form is used before updating.

  • Adding variances instead of precisions in the normal-normal model.

    It feels natural to combine variances, but the update is additive in 1/variance.

    Fix: Convert to precisions first: n/σ² and 1/σ₀². Add them, then invert to get the posterior variance.

  • Giving the posterior mean when the question specifies absolute or all-or-nothing loss.

    The mean is the most familiar estimate, so students default to it without reading the loss function.

    Fix: Underline the loss function in the question. Quadratic → mean, absolute → median, all-or-nothing → mode.

  • Using the formula α/β for the mode of a gamma distribution.

    Students confuse the mean α/β with the mode (α − 1)/β.

    Fix: Remember the mode comes from differentiating θ^(α−1) e^(−βθ), which gives (α − 1)/β. It requires α ≥ 1.

  • Keeping constants and ignoring the proportionality sign, or dropping terms that do contain θ.

    Students are unsure which factors to discard.

    Fix: Discard only factors free of θ. Keep every power of θ and every exponential in θ, then match the kernel.

Worked examples

Example 1

The number of claims per year on a policy is Poisson(λ). The prior for λ is Gamma(α = 3, β = 2), with β a rate. Over 5 years the claim counts are 2, 0, 1, 3, 1. (a) Find the posterior distribution of λ. (b) Find the Bayes estimate under quadratic loss and under all-or-nothing loss. (c) Show the quadratic-loss estimate in credibility form.

Show the solution
  1. Model: Xᵢ | λ ~ Poisson(λ), prior f(λ) ∝ λ^(3−1) e^(−2λ).
  2. Likelihood: L(λ) ∝ λ^(Σx) e^(−nλ). Here Σx = 2 + 0 + 1 + 3 + 1 = 7 and n = 5, so L(λ) ∝ λ⁷ e^(−5λ).
  3. Posterior: f(λ | x) ∝ λ⁷ e^(−5λ) × λ² e^(−2λ) = λ⁹ e^(−7λ). This is the kernel of Gamma(10, 7).
  4. Quadratic loss gives the posterior mean: 10/7 = 1.4286.
  5. All-or-nothing loss gives the posterior mode: (10 − 1)/7 = 9/7 = 1.2857.
  6. Credibility form: x̄ = 7/5 = 1.4 and Z = n/(n + β) = 5/7. Prior mean = 3/2 = 1.5.
  7. Z x̄ + (1 − Z) × 1.5 = (5/7)(1.4) + (2/7)(1.5) = 1.0 + 0.4286 = 1.4286, which matches 10/7.

Answer: Posterior: λ | x ~ Gamma(10, 7). Quadratic-loss estimate = 10/7 ≈ 1.4286. All-or-nothing-loss estimate = 9/7 ≈ 1.2857. In credibility form, Z = 5/7.

Example 2

Claim sizes (in ₹ thousands) X₁,…,X₄ are independent N(μ, 16), with σ² = 16 known. The prior is μ ~ N(50, 25). The sample mean is x̄ = 58. (a) Find the posterior distribution of μ. (b) State the Bayes estimate under quadratic loss and under absolute loss. (c) Find a 95% equal-tailed credible interval.

Show the solution
  1. Data precision: n/σ² = 4/16 = 0.25. Prior precision: 1/σ₀² = 1/25 = 0.04.
  2. Posterior precision = 0.25 + 0.04 = 0.29. Posterior variance = 1/0.29 = 3.448.
  3. Posterior mean = (0.25 × 58 + 0.04 × 50) ÷ 0.29 = (14.5 + 2.0) ÷ 0.29 = 16.5 ÷ 0.29 = 56.897.
  4. Check with credibility: Z = 0.25/0.29 = 0.8621, so Z × 58 + (1 − Z) × 50 = 50.0 + 6.897 − ... computed directly: 0.8621 × 58 = 50.00 and 0.1379 × 50 = 6.897, total 56.897. This agrees.
  5. The posterior is normal, so it is symmetric. Mean = median = mode = 56.897. Both quadratic and absolute loss give 56.90.
  6. Posterior standard deviation = √3.448 = 1.857.
  7. 95% interval: 56.897 ± 1.96 × 1.857 = 56.897 ± 3.640, which gives (53.26, 60.54).

Answer: Posterior: μ | x ~ N(56.90, 3.448). The Bayes estimate under quadratic or absolute loss is 56.90 (₹56,900 as the claim-size mean). The 95% credible interval is (53.26, 60.54) in ₹ thousands.

Exam tips

  • Read the loss function before computing anything. Many marks depend on choosing mean, median or mode correctly.
  • If the question says 'show that the posterior is…', write likelihood × prior and identify the kernel. Quoting the conjugate result without working loses method marks.
  • Always state the posterior in full: name, and both parameters. Then give the estimate.
  • Check your posterior mean against the credibility form Z x̄ + (1 − Z) × prior mean. It takes ten seconds and catches most errors.
  • In computer-based Paper B, use the conjugate update in R, for example qgamma(c(0.025, 0.975), shape, rate) for a credible interval. Show the update formula in your working as well.

Practice questions from Bayesian inference: priors, posteriors, loss functions and credible intervals

Posterior Distributions and Bayesian Estimators: frequently asked questions

How do I find the posterior distribution for a Poisson-gamma model?

Multiply the Poisson likelihood λ^(Σx) e^(−nλ) by the gamma prior λ^(α−1) e^(−βλ). This gives λ^(α+Σx−1) e^(−(β+n)λ), which is Gamma(α + Σx, β + n). The shape takes the total count and the rate takes the number of observations.

What is the difference between posterior mean, median and mode as estimates?

Each one minimises expected loss for a different loss function. The mean is optimal for quadratic loss, the median for absolute loss and the mode for all-or-nothing loss. For a symmetric single-peaked posterior such as the normal, all three coincide.

Why does the posterior mean look like a credibility estimate?

The posterior mean is a weighted average of the sample mean and the prior mean. The weight on the data, Z, rises as n grows. In the Poisson-gamma model, Z = n/(n + β).

Do I need to derive the posterior from scratch in the exam?

Only if the question asks you to show or derive it. Otherwise you can quote the standard conjugate result, but state the model and the updated parameters clearly. When unsure, write the proportionality step. It is short and earns method marks.