Actuarial Statistics · Bayesian inference: priors, posteriors, loss functions and credible intervals
Bayes' Theorem and the Bayesian Framework: Prior, Likelihood and Posterior
Updated 11 October 2026 · Fact-checked
Bayesian inference treats a parameter as a random variable. You start with a prior distribution, multiply it by the likelihood of the observed data, and normalise to get the posterior: f(θ | x) ∝ f(θ) × L(x | θ). To solve questions, write prior and likelihood, drop constants, and recognise the posterior's form.
Understand Bayes' Theorem and Bayesian Framework
In classical (frequentist) statistics, a parameter θ is a fixed but unknown number. Probability describes only the data. You estimate θ with a point estimate or a confidence interval, and probability statements refer to repeated sampling.
In Bayesian statistics, θ is treated as a random variable. Your uncertainty about it, before seeing data, is written as a prior distribution f(θ). The prior can come from past experience, expert judgement or earlier data. For example, an insurer may hold a belief about a claim rate from similar portfolios.
The data x enter through the likelihood L(x | θ), which is the joint probability or density of the data given θ. Bayes' theorem combines the two: the posterior f(θ | x) = f(x | θ) f(θ) ÷ f(x), where f(x) is the integral (or sum) of f(x | θ) f(θ) over θ. This denominator does not depend on θ. It is only a normalising constant that makes the posterior integrate to 1.
So in practice you use: posterior ∝ prior × likelihood. Any factor not containing θ can be dropped. Then you look at what is left and match it to a known distribution, such as a gamma or beta, and read off its parameters.
The main differences from the classical approach: Bayesian inference uses a prior, gives a full distribution for θ, and allows direct statements such as 'the probability that θ lies in this interval is 0.95'. Classical inference uses only the data, and its probability statements are about the procedure. With a lot of data the likelihood dominates and the prior matters less. With little data the prior has more influence.
Key rules to remember
- Bayes' theorem for events
- P(A | B) = P(B | A) × P(A) ÷ P(B)
- Valid when P(B) > 0. P(B) is found by the law of total probability.
- Bayes' theorem for a parameter (continuous θ)
- f(θ | x) = f(x | θ) f(θ) ÷ ∫ f(x | θ) f(θ) dθ
- For a discrete θ, replace the integral with a sum over θ.
- Proportionality form
- f(θ | x) ∝ f(θ) × L(x | θ)
- Drop every factor that does not involve θ, then identify the distribution.
- Likelihood for independent data
- L(x | θ) = Π f(xᵢ | θ), i = 1 to n
- Assumes the observations are independent given θ.
- Gamma density kernel
- f(θ) ∝ θ^(α−1) e^(−βθ), θ > 0
- Gamma(α, β) with rate β. Mean α ÷ β.
- Beta density kernel
- f(θ) ∝ θ^(α−1) (1−θ)^(β−1), 0 < θ < 1
- Mean α ÷ (α + β).
How to solve Bayes' Theorem and Bayesian Framework questions
Use this routine for any question that asks you to find or describe a posterior distribution.
- 1Write down the prior f(θ) and state its range for θ.
- 2Write the likelihood L(x | θ) for the data. For independent observations, multiply the individual densities or probabilities.
- 3Multiply: posterior ∝ prior × likelihood. Collect powers of θ and exponentials together.
- 4Drop all constants and any terms not involving θ. Keep only the kernel in θ.
- 5Compare the kernel with a standard distribution (gamma, beta, normal). State the name and its parameters.
- 6If the kernel is not recognisable, find the normalising constant by integrating over θ.
- 7Answer what is asked: the posterior distribution, its mean, or a probability. Check the answer is sensible, for example the posterior mean lies between the prior mean and the data estimate.
Quickest way: Kernel matching
When to use it: Use it when the prior and likelihood belong to a standard pair, such as Poisson with gamma, binomial with beta, or exponential with gamma, and the question asks only for the posterior form.
- Write only the θ-dependent parts of the prior and the likelihood.
- Add the exponents of θ and combine the exponential terms.
- Read the new parameters straight from the combined kernel.
- State the distribution and its parameters without computing any constant.
Common mistakes in Bayes' Theorem and Bayesian Framework
Forgetting that ∝ hides a constant and treating the product as a proper density.
Students stop after multiplying prior by likelihood.
Fix: Always identify the kernel with a named distribution, or integrate to normalise, before quoting probabilities or means.
Keeping terms that do not involve θ inside the kernel.
Students copy the whole likelihood, including factorials and other constants.
Fix: Remove everything free of θ. Then the powers of θ and the exponent are easy to read.
Using only one observation in the likelihood when n observations are given.
Rushing, or not seeing that independence gives a product.
Fix: Write L = Π f(xᵢ | θ) and combine the powers, for example θ^(Σxᵢ) e^(−nθ) for a Poisson sample.
Reading the gamma parameters wrongly, such as mixing up rate and scale.
Different conventions exist for the gamma distribution.
Fix: State the convention you use. With rate β the kernel is θ^(α−1) e^(−βθ). Check it against the Tables.
Saying the Bayesian 95% interval means the same as a classical confidence interval.
Both give a range and use the same number.
Fix: In Bayesian inference, θ is random, so the probability that θ lies in the interval is 0.95 given the data. A classical interval refers to repeated samples.
Writing the posterior as P(x | θ) instead of P(θ | x).
Confusing the likelihood with the posterior.
Fix: The likelihood is the data given the parameter. The posterior is the parameter given the data.
Worked examples
Example 1
The number of claims in a year on a policy is Poisson with mean θ. The prior for θ is gamma with parameters α = 3 and rate β = 2. In one year, 5 claims are observed. Find the posterior distribution of θ and its mean.
Show the solution
- Prior: f(θ) ∝ θ^(3−1) e^(−2θ) = θ² e^(−2θ), for θ > 0.
- Likelihood: L(x | θ) = e^(−θ) θ⁵ ÷ 5!, which is ∝ θ⁵ e^(−θ).
- Posterior ∝ θ² e^(−2θ) × θ⁵ e^(−θ) = θ⁷ e^(−3θ).
- Match with gamma kernel θ^(α'−1) e^(−β'θ): α' − 1 = 7, so α' = 8, and β' = 3.
- Posterior mean = α' ÷ β' = 8 ÷ 3 = 2.667. (The prior mean was 3 ÷ 2 = 1.5 and the data value was 5, so 2.667 lies between them.)
Answer: The posterior is gamma with α = 8 and rate β = 3, with mean 8/3 ≈ 2.67.
Example 2
A parameter θ takes only two values: θ = 0.2 with prior probability 0.6, and θ = 0.5 with prior probability 0.4. A single trial is a Bernoulli trial with success probability θ, and one success is observed. Find the posterior probability that θ = 0.5.
Show the solution
- Prior: P(θ = 0.2) = 0.6 and P(θ = 0.5) = 0.4.
- Likelihood of one success: P(x = 1 | θ) = θ. So it is 0.2 for θ = 0.2 and 0.5 for θ = 0.5.
- Prior × likelihood: for θ = 0.2, 0.6 × 0.2 = 0.12. For θ = 0.5, 0.4 × 0.5 = 0.20.
- The normalising constant is the sum: 0.12 + 0.20 = 0.32.
- Posterior P(θ = 0.5 | x = 1) = 0.20 ÷ 0.32 = 0.625.
Answer: The posterior probability that θ = 0.5 is 0.625.
Exam tips
- In written questions, show prior, likelihood, product and kernel as separate lines. Marks are given for each stage.
- Always state the final distribution with its name and parameters. Do not stop at 'proportional to'.
- For discrete θ, build a small table of prior, likelihood, product and posterior. It reduces arithmetic slips.
- In MCQs, check whether the question wants the posterior mean, the posterior distribution or a probability. These give different numbers.
- Compare Bayesian and classical ideas in words when asked. Mention the prior, θ as random, and the meaning of the interval.
Practice questions from Bayesian inference: priors, posteriors, loss functions and credible intervals
- A claim-frequency parameter p (probability a policy has a claim) is given a Beta(2, 8) prior. In a sample of 20 policies, 5 have a claim. Wh…
- The prior for a proportion p is Beta(a, b) with prior mean 0.25 and a + b = 12. After observing 6 successes in 18 trials, what is the poster…
- A posterior distribution for a parameter θ is Normal with mean 40 and variance 25. What is the equal-tailed 95% Bayesian credible interval f…
- In the Bayesian framework, the posterior density of a parameter θ given data x is proportional to which of the following?
- A prior distribution for a parameter θ is combined with data x. Which statement correctly describes the Bayesian relationship between the po…
Bayes' Theorem and Bayesian Framework in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Bayes' Theorem and Bayesian Framework: frequently asked questions
What is the difference between Bayesian and classical inference?
Classical inference treats θ as a fixed unknown and uses only the data. Bayesian inference treats θ as random, with a prior that is updated by the data into a posterior. Bayesian methods give a full distribution for θ.
Why can I drop constants when finding the posterior?
The posterior must integrate to 1, so any constant free of θ is fixed by normalisation. If you recognise the kernel as a standard distribution, you already know its normalising constant.
What is the likelihood in Bayesian inference?
It is the probability or density of the observed data viewed as a function of θ. For independent observations it is the product of the individual terms.
Does the prior always matter?
It matters most when data are few. As the sample size grows, the likelihood tends to dominate, and the posterior is driven mainly by the data.