Actuarial Statistics · Random sampling and sampling distributions
Sampling Distributions and the Central Limit Theorem
Updated 11 October 2026 · Fact-checked
A sampling distribution is the probability distribution of a statistic, such as the sample mean, over repeated samples. For normal populations, the sample mean is exactly normal. For other populations with finite variance, the Central Limit Theorem says it is approximately normal when n is large. Standardise, then use normal tables.
Understand Sampling Distributions and Central Limit Theorem
A statistic is a function of sample data, such as the sample mean. Before you collect the data, the statistic is a random variable. Its probability distribution is called its sampling distribution. It tells you how the statistic varies from sample to sample.
Take X₁, ..., Xₙ as independent and identically distributed (iid) with mean μ and variance σ². The sample mean X̄ always has mean μ and variance σ² ÷ n. This holds for any distribution with finite variance. The sum has mean nμ and variance nσ². The standard deviation of X̄ is σ ÷ √n, called the standard error. It shrinks as n grows.
If the population is normal, X̄ is exactly normal: X̄ ~ N(μ, σ²/n), for every n. No approximation is needed.
If the population is not normal, the Central Limit Theorem (CLT) applies. For iid variables with finite mean μ and finite variance σ², (X̄ − μ) ÷ (σ ÷ √n) converges in distribution to N(0, 1) as n → ∞. So for large n, X̄ is approximately N(μ, σ²/n) and the sum Sₙ is approximately N(nμ, nσ²). How large n must be depends on skewness. A common rule of thumb is n around 30, but it is only a guide.
In actuarial work, the CLT is used for aggregate claims, where total claims are a sum of many individual claims. It is also the basis for large-sample confidence intervals.
Key rules to remember
- Mean and variance of the sample mean
- E(X̄) = μ ; Var(X̄) = σ² ÷ n
- Needs iid observations with finite variance. Holds for any distribution.
- Standard error
- SE(X̄) = σ ÷ √n
- Use s in place of σ only when σ is unknown, and then the t distribution is relevant for normal data.
- Mean and variance of a sum
- E(Sₙ) = nμ ; Var(Sₙ) = nσ²
- Sₙ = X₁ + ... + Xₙ. Variance adds only for independent variables.
- Normal population
- X̄ ~ N(μ, σ² ÷ n)
- Exact for any n when each Xᵢ ~ N(μ, σ²).
- Central Limit Theorem
- (X̄ − μ) ÷ (σ ÷ √n) → N(0, 1) as n → ∞
- Requires iid with finite variance. It is an approximation for finite n.
- CLT for a sum
- Z = (Sₙ − nμ) ÷ (σ √n) ≈ N(0, 1)
- Use for totals such as aggregate claims.
- Continuity correction
- P(Sₙ ≤ k) ≈ Φ((k + 0.5 − nμ) ÷ (σ√n))
- Use only when the sum is integer valued, for example binomial or Poisson.
How to solve Sampling Distributions and Central Limit Theorem questions
Use this method for any question that asks for a probability or a value involving a sample mean or a sum.
- 1Identify the random variable: a single observation, the sample mean X̄ or the sum Sₙ. Note n.
- 2Write down μ and σ² for one observation. If only the distribution is given, compute them from its parameters.
- 3Find the mean and variance of the statistic: μ and σ²/n for X̄, or nμ and nσ² for Sₙ.
- 4State the distribution. Say exact if the population is normal. Otherwise say approximate by the CLT, and note that n is large.
- 5Apply a continuity correction only if the variable is discrete and integer valued.
- 6Standardise: Z = (value − mean) ÷ standard deviation of the statistic.
- 7Read Φ from the tables, using symmetry for negative values, and give the probability.
- 8Check the answer is sensible and state any assumption (independence, identical distribution).
Quickest way: Standardise the statistic in one line
When to use it: Use for multiple-choice questions and for fast checks in written answers, when n is large or the population is normal.
- Write the mean and standard deviation of the statistic directly: μ and σ/√n for the mean, or nμ and σ√n for the sum.
- Compute z = (value − mean) ÷ sd in one step.
- Use Φ(z), or 1 − Φ(z) for an upper tail.
- For a sum of integer variables, move the boundary by 0.5 before computing z.
- State the CLT as the reason in one sentence.
Common mistakes in Sampling Distributions and Central Limit Theorem
Using σ as the standard deviation of X̄ instead of σ/√n.
Students standardise as if they were dealing with one observation.
Fix: Underline the word mean or sum in the question. For X̄, always divide σ by √n.
Writing Var(Sₙ) = n²σ² or Var(X̄) = σ².
Confusing the scaling rule Var(aX) = a²Var(X) with adding independent variables.
Fix: Sum of n independent variables: variance nσ². Then X̄ = Sₙ/n, so the variance is nσ²/n² = σ²/n.
Saying the CLT makes the population or a single observation normal.
Misreading what is converging.
Fix: The CLT is about the distribution of the sample mean or sum. The data stay as they are.
Treating the CLT as exact, or applying it to a small n from a skewed distribution without comment.
The result is memorised as a rule without conditions.
Fix: Say approximately. Mention that accuracy depends on n and on skewness. If the population is normal, state that the result is exact.
Forgetting the continuity correction for discrete sums, or using it for a continuous variable.
Students do not check whether the variable is integer valued.
Fix: Correct by 0.5 only for integer-valued variables. For P(S ≥ k) use k − 0.5.
Applying the CLT when the variance is infinite or the observations are not independent.
Conditions are skipped.
Fix: Check iid and finite variance first. Heavy-tailed claim distributions may not satisfy the conditions.
Worked examples
Example 1
The claim amounts on a policy are iid with mean ₹8,000 and standard deviation ₹3,000. A portfolio has 400 such policies, each with one claim. Using the CLT, find the approximate probability that the total claims exceed ₹33,00,000.
Show the solution
- Let S = total claims = sum of 400 iid variables. μ = 8,000, σ = 3,000, n = 400.
- E(S) = 400 × 8,000 = 32,00,000.
- Var(S) = 400 × 3,000² = 400 × 90,00,000 = 360,00,00,000. SD(S) = 3,000 × √400 = 3,000 × 20 = 60,000.
- By the CLT, S is approximately N(32,00,000, 60,000²). No continuity correction, as the claims are continuous amounts.
- z = (33,00,000 − 32,00,000) ÷ 60,000 = 1,00,000 ÷ 60,000 = 1.6667.
- P(S > 33,00,000) ≈ 1 − Φ(1.67) = 1 − 0.9525 = 0.0475.
Answer: About 0.048 (approximately 4.8%).
Example 2
The number of claims per day is Poisson with mean 4, and days are independent. Using the CLT with continuity correction, find the approximate probability that there are at most 110 claims in 30 days.
Show the solution
- Let S = total claims in 30 days. A sum of independent Poisson variables is Poisson, so S ~ Poisson(120).
- E(S) = 120 and Var(S) = 120. SD = √120 = 10.954.
- S is integer valued, so P(S ≤ 110) ≈ P(Normal ≤ 110.5).
- z = (110.5 − 120) ÷ 10.954 = −9.5 ÷ 10.954 = −0.867.
- P ≈ Φ(−0.867) = 1 − Φ(0.867). From tables Φ(0.87) = 0.8078, so the probability is about 0.1922; interpolating gives about 0.193.
Answer: About 0.19.
Exam tips
- Start every answer by writing the mean and variance of the statistic. Marks are given for this even if the final number is wrong.
- Say whether the result is exact (normal population) or approximate (CLT). Examiners look for this statement.
- Check whether the question uses the variance or the standard deviation before you take the square root.
- For discrete sums such as Poisson or binomial totals, apply the continuity correction and show the 0.5 explicitly.
- In the computer-based paper, you can check a CLT answer by simulating many samples and comparing the histogram of means with the normal curve.
Practice questions from Random sampling and sampling distributions
- A sample of n = 3 is taken from an Exponential distribution with mean 10 (rate 0.1). What is the expected value of the sample minimum?
- A random sample X1, ..., Xn is drawn from a population with mean 50 and variance 64. Which of the following is the standard deviation of the…
- The ordered sample of nine premium amounts (in ₹ thousand) is 12, 15, 18, 21, 25, 30, 34, 41, 60. Using the definition in which the p-quanti…
- A sample of 16 claim amounts from a normal population has mean 50 and sample standard deviation 8 (divisor n-1). What is the t statistic for…
- A random sample of 5 values from a population gives 4, 7, 9, 12, 8. What is the unbiased estimate of the population variance?
Sampling Distributions and Central Limit Theorem in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Sampling Distributions and Central Limit Theorem: frequently asked questions
What is the sampling distribution of the sample mean?
It is the distribution of X̄ over all possible samples of size n from the population. It has mean μ and variance σ²/n. It is exactly normal if the population is normal, and approximately normal for large n otherwise.
How large must n be for the CLT to work?
There is no fixed value. A rule of thumb is about 30, but a strongly skewed population may need far more. A symmetric population may need far fewer.
How do I apply the CLT to a sum of iid random variables?
Find nμ and nσ² for the sum. Then treat the sum as N(nμ, nσ²) and standardise using the standard deviation σ√n. Use a continuity correction if the variables are integer valued.
What is the difference between standard deviation and standard error?
Standard deviation describes the spread of single observations. Standard error is the standard deviation of a statistic, such as σ/√n for the sample mean.