Actuarial Statistics · Confidence intervals and prediction intervals
Confidence Intervals for Proportions and Large Samples
Updated 11 October 2026 · Fact-checked
A large-sample confidence interval uses a normal approximation: estimate ± z × standard error. For a proportion, the standard error is √(p̂(1 − p̂)/n). For an MLE, it is 1/√I(θ̂), where I is the Fisher information. Find the estimate, find its variance, plug in the estimate, then multiply by the z value.
Understand Confidence Intervals for Proportions and Large Samples
Exact confidence intervals need a pivotal quantity with a known distribution. For a normal mean, that is the t distribution. For a binomial proportion or a Poisson mean, no such simple pivot exists. So we use a large-sample approximation.
The idea is simple. If an estimator is approximately normal with a known mean and variance, then (estimator − parameter) ÷ standard error is approximately N(0,1). You can then read off a z value and build the interval in the usual way.
For a proportion, the number of successes X is Binomial(n, p). By the central limit theorem, p̂ = X/n is approximately N(p, p(1 − p)/n) when n is large. The variance contains the unknown p, so you replace p by p̂ in the standard error. This is why the interval is called approximate.
For a general parameter, use the asymptotic result for maximum likelihood estimators. For large n, θ̂ is approximately N(θ, 1/I(θ)). Here I(θ) is the Fisher information from the whole sample. It equals −E[d²l/dθ²], where l is the log-likelihood. The same variance is the Cramér-Rao lower bound, so the MLE is asymptotically efficient. Replace θ by θ̂ in I and you get the interval.
For two independent proportions, the variances of p̂1 and p̂2 add. The difference p̂1 − p̂2 is approximately normal with that summed variance. If the interval for p1 − p2 contains 0, the data do not show a difference at that confidence level. The approximation is poor when n is small or p is close to 0 or 1. Say so in your answer.
Key rules to remember
- Interval for one proportion
- p̂ ± z × √(p̂(1 − p̂) ÷ n), where p̂ = x ÷ n
- z is the 1 − α/2 point of N(0,1). For 95%, z = 1.96. For 90%, z = 1.645. For 99%, z = 2.576. Needs large n.
- Interval for difference of two proportions
- (p̂1 − p̂2) ± z × √(p̂1(1 − p̂1) ÷ n1 + p̂2(1 − p̂2) ÷ n2)
- Samples must be independent. The variances add, even though you are taking a difference.
- Fisher information
- I(θ) = −E[d²l/dθ²] = E[(dl/dθ)²]
- For n independent observations, I(θ) = n × I₁(θ), where I₁ is the information from one observation.
- Asymptotic distribution of the MLE
- θ̂ ≈ N(θ, 1 ÷ I(θ)) for large n
- Holds under regularity conditions. The variance equals the Cramér-Rao lower bound.
- Asymptotic interval from the MLE
- θ̂ ± z ÷ √I(θ̂)
- Evaluate the information at the estimate θ̂.
- Poisson mean
- λ̂ = x̄, I(λ) = n ÷ λ, interval: x̄ ± z × √(x̄ ÷ n)
- For a Poisson count over n exposure units, use the total count divided by n.
- Function of a parameter (delta method)
- g(θ̂) ≈ N(g(θ), [g′(θ)]² ÷ I(θ))
- Use it for an interval for a function of θ. Requires g′(θ) ≠ 0.
How to solve Confidence Intervals for Proportions and Large Samples questions
Use this method for any large-sample interval question, whether the parameter is a proportion, a Poisson mean or a general MLE.
- 1Name the parameter and the model, for example X ~ Binomial(n, p) or Poisson(λ). State that you are using a large-sample normal approximation.
- 2Write the estimate. For a proportion, p̂ = x/n. For other cases, find the MLE by maximising the log-likelihood.
- 3Find the variance of the estimator. For a proportion, use p(1 − p)/n. For an MLE, use 1/I(θ), where I(θ) = −E[d²l/dθ²].
- 4Replace the unknown parameter by its estimate in the variance. Take the square root to get the standard error.
- 5Pick z from the confidence level: 1.645 for 90%, 1.96 for 95%, 2.576 for 99%.
- 6Compute estimate ± z × standard error. Keep at least four decimal places until the final answer.
- 7Check the answer is sensible. A proportion interval should lie within 0 and 1. For a difference, state whether 0 lies inside the interval.
- 8Write a one-line conclusion in words, and note any caveat about the sample size.
Quickest way: Estimate, standard error, z
When to use it: Use this for multiple-choice questions and for the final calculation in written questions, once the model is clear.
- Write estimate ± z × SE. Fix z first.
- For proportions, compute p̂(1 − p̂) before dividing by n. Divide, then square root.
- For a Poisson mean, SE = √(x̄ ÷ n). You do not need the information formula.
- For a difference, compute both variances separately, add them, then take the root.
- Estimate the answer first. The margin should be about 2 standard errors at 95%, so check that your final number is close to that.
Common mistakes in Confidence Intervals for Proportions and Large Samples
Subtracting the variances when finding the interval for p1 − p2.
Students link a difference of estimates to a difference of variances.
Fix: For independent samples, Var(p̂1 − p̂2) = Var(p̂1) + Var(p̂2). Always add.
Forgetting to divide by n inside the square root, or using p̂(1 − p̂) as the variance of p̂.
This confuses the variance of one observation with that of the mean.
Fix: The variance of p̂ is p(1 − p)/n. Write it out with the n before you compute.
Using the information from one observation instead of the whole sample.
Students work out I₁(θ) and forget the factor n.
Fix: Use I(θ) = n × I₁(θ) for n independent observations. Check that the standard error shrinks as n grows.
Finding the second derivative but not taking the negative sign or the expectation.
Rushing through the differentiation.
Fix: Fisher information is −E[d²l/dθ²]. It must be positive. For the Poisson, d²l/dλ² = −Σx/λ², and the expectation gives −n/λ.
Using the t distribution for a proportion or Poisson interval.
Students link every confidence interval with t.
Fix: The t distribution comes from a normal sample with unknown variance. Here the variance is estimated from the model, so use z.
Applying the normal approximation when n is small or p̂ is near 0 or 1, and giving limits outside 0 to 1.
Students apply the formula without checking it is valid.
Fix: Check that n p̂ and n(1 − p̂) are both reasonably large. If the limits fall outside [0, 1], say the approximation is poor.
Worked examples
Example 1
In region A, 90 of 300 sampled policyholders lapsed within a year. In region B, 60 of 250 lapsed. Assuming independent samples, find an approximate 95% confidence interval for the difference in lapse probabilities pA − pB. Does the interval suggest a real difference?
Show the solution
- Model: XA ~ Binomial(300, pA) and XB ~ Binomial(250, pB), independent. n is large, so use the normal approximation.
- Estimates: p̂A = 90/300 = 0.30 and p̂B = 60/250 = 0.24. The difference is 0.06.
- Variance of p̂A: 0.30 × 0.70 ÷ 300 = 0.21 ÷ 300 = 0.000700.
- Variance of p̂B: 0.24 × 0.76 ÷ 250 = 0.1824 ÷ 250 = 0.0007296.
- Add the variances: 0.000700 + 0.0007296 = 0.0014296. SE = √0.0014296 = 0.03781.
- Margin: 1.96 × 0.03781 = 0.07411.
- Interval: 0.06 ± 0.07411 gives (−0.0141, 0.1341).
Answer: The approximate 95% interval for pA − pB is (−0.014, 0.134). It contains 0, so the data do not show a difference in lapse rates at the 95% level.
Example 2
The number of claims on each of 80 identical policies in one year is modelled as independent Poisson(λ). The total number of claims is 120. (a) Find the MLE of λ and the Fisher information. (b) Find an approximate 95% confidence interval for λ.
Show the solution
- Log-likelihood: l(λ) = −nλ + (Σx) ln λ − Σ ln(x!).
- First derivative: dl/dλ = −n + Σx ÷ λ. Setting it to zero gives λ̂ = Σx ÷ n = 120 ÷ 80 = 1.5.
- Second derivative: d²l/dλ² = −Σx ÷ λ².
- Take expectations, using E[Σx] = nλ: E[d²l/dλ²] = −nλ ÷ λ² = −n ÷ λ. So I(λ) = n ÷ λ.
- Evaluate at the estimate: I(λ̂) = 80 ÷ 1.5 = 53.333.
- Standard error: 1 ÷ √53.333 = √0.01875 = 0.13693.
- Margin: 1.96 × 0.13693 = 0.2684.
- Interval: 1.5 ± 0.2684 gives (1.2316, 1.7684).
Answer: λ̂ = 1.5 and I(λ) = n/λ = 53.33 at the estimate. The approximate 95% confidence interval for λ is (1.232, 1.768).
Exam tips
- Always state that the interval is approximate and rests on a large sample. Examiners give a mark for the assumption.
- If asked to derive the interval from an MLE, show the log-likelihood, both derivatives and the expectation. Marks are given for each step.
- Check that the information is positive and that the standard error falls as n rises. This catches sign and n errors quickly.
- For a difference of proportions, say whether 0 lies inside the interval and what that means. Interpretation questions are common.
- In the computer-based paper, show the formula and the inputs before the result. For example, write phat, SE and the qnorm value as separate lines in R.
Practice questions from Confidence intervals and prediction intervals
- Two branches of an Indian life insurer sample policy renewals. Branch A: 150 of 200 renewed. Branch B: 120 of 200 renewed. Using the normal …
- A random sample of 25 claim sizes (in Rs thousand) from a normal distribution has sample mean 120 and sample standard deviation 10. Using t(…
- For a sample of n from a normal distribution with unknown mean and variance, which statement about the 95% confidence interval for μ based o…
- A 95% confidence interval for a normal mean with known variance has width 6 using a sample of size 36. Keeping the same confidence level and…
- For a 95% confidence interval for the ratio of two normal population variances sigma1^2/sigma2^2, based on s1^2 = 20 (n1 = 9) and s2^2 = 10 …
Confidence Intervals for Proportions and Large Samples in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Confidence Intervals for Proportions and Large Samples: frequently asked questions
When can I use the normal approximation for a proportion?
Use it when n is large enough that both n p̂ and n(1 − p̂) are reasonably big. If p̂ is near 0 or 1, or n is small, the interval can be poor and may fall outside 0 to 1. State this caveat in your answer.
Why do I use p̂ in the standard error instead of p?
The true p is unknown, so the variance p(1 − p)/n cannot be evaluated. Replacing p by p̂ gives a close approximation for large n. This is why the interval is only approximate.
How do I get a confidence interval from Fisher information?
Find the MLE θ̂. Work out I(θ) = −E[d²l/dθ²] for the whole sample. Then use θ̂ ± z ÷ √I(θ̂), because θ̂ is approximately N(θ, 1/I(θ)) for large n.
Do I add or subtract variances for a difference of proportions?
Add them. For independent samples, Var(p̂1 − p̂2) = Var(p̂1) + Var(p̂2). Only the estimates are subtracted.
What z value do I use for a 95% interval?
Use 1.96 for 95%, because 2.5% lies in each tail of N(0,1). Use 1.645 for 90% and 2.576 for 99%.