IAI Actuarial Core Principles · Actuarial Statistics
Confidence Intervals and Prediction Intervals Explained
A confidence interval is a range, built from sample data, that captures an unknown parameter in a stated proportion of repeated samples. You find a pivotal quantity, take its distribution, and invert the probability statement. A prediction interval does the same for a future observation, so it is wider.
What this chapter covers
This chapter is about turning a sample into a range of plausible values. A confidence interval gives a range for a fixed but unknown parameter such as a mean, a variance or a proportion. A prediction interval gives a range for a single future observation. Both rest on one idea: a pivotal quantity, a function of the data and the parameter whose distribution does not depend on the parameter.
You first learn the general method, then apply it case by case. For a normal mean with unknown variance you use the t distribution. For a normal variance you use the chi-squared distribution. For two samples you compare means, variances or proportions, and you must decide whether variances can be pooled. For proportions and large samples you use the normal approximation from the central limit theorem. Prediction intervals then add the extra variance of the new observation.
This chapter links directly to the rest of CS1. It uses the sampling distributions from the random variables and distributions section. It sits beside hypothesis testing, because a confidence interval and a two-sided test use the same pivot. It feeds into regression, where you build intervals for slope coefficients, mean responses and predicted responses. Bayesian credible intervals later give you a contrast in interpretation. It also helps in Paper B, where you compute intervals in R.
Statistical inference carries a large share of the CS1 syllabus, and intervals appear in both the multiple-choice section and the written questions. The written questions reward method: naming the pivot, stating assumptions, giving the quantiles and writing the final interval. These steps are mechanical once practised, so the marks are reliable. The same skills carry into regression and into the computer-based paper, so the effort pays back more than once.
Confidence intervals and prediction intervals: topics in the order to study them
- 1Confidence Interval Basics and Pivotal QuantitiesEvery later interval is built with the pivot method, and the interpretation of confidence level has to be right before you use it.
- 2Confidence Intervals for Normal Mean and VarianceThese are the core worked cases, using the standard normal, t and chi-squared distributions, and they show the pivot method in full.
- 3Confidence Intervals for Proportions and Large SamplesThis applies the central limit theorem to the same pivot idea and is a simpler extension before you handle two samples.
- 4Confidence Intervals for Two SamplesYou need single-sample intervals first, because two-sample cases combine them and add decisions about pooling and paired data.
- 5Prediction IntervalsThis comes last because it reuses the mean interval and adds the variance of a new observation, which is a common source of confusion.
How to prepare Confidence intervals and prediction intervals
Aim to understand one method and apply it repeatedly, rather than memorise a list of formulas. Short daily sessions work well if you study alongside work.
- Write the pivot method in your own words: find a pivot, write P(a < pivot < b) = 1 − α, then rearrange for the parameter.
- Learn which distribution goes with which case: known variance, unknown variance, variance itself, large samples. Make a one-page table of conditions and pivots.
- Practise reading t and chi-squared quantiles from the Tables until you can do it without hesitation, including the upper and lower points for variance intervals.
- Do two-sample questions by first deciding the setup: independent or paired, equal variances or not, normal or large sample. State that choice before calculating.
- Practise prediction intervals next to confidence intervals for the same data, and write down why one is wider.
- Do past-paper written questions under time, then check that you stated assumptions, quantiles and a conclusion in words.
- Repeat the key intervals in R for Paper B, and check your hand answers against the output.
Common mistakes in Confidence intervals and prediction intervals
Saying there is a 95% probability that the parameter lies in the calculated interval.
Fix: Say that the method gives intervals that contain the parameter in 95% of repeated samples. Use this wording in written answers.
Using the z value when the variance is estimated from a small sample.
Fix: Ask first whether σ is given. If you use s with a normal sample, use the t distribution with n−1 degrees of freedom.
Taking the wrong chi-squared quantiles for a variance interval, or applying symmetric ± logic.
Fix: Divide (n−1)s² by the upper quantile to get the lower limit, and by the lower quantile to get the upper limit. Check that the lower limit is smaller.
Pooling variances without checking the assumption in two-sample problems.
Fix: State whether the variances are assumed equal, and justify it from the question or from an F comparison. Treat paired data as a one-sample problem on differences.
Using the confidence interval for the mean when the question asks for a prediction interval.
Fix: Underline whether the question concerns a parameter or a future observation. For a future value, include the √(1 + 1/n) factor.
Omitting assumptions and a concluding sentence in written answers.
Fix: State the model, the pivot and the assumptions, show the quantiles, then interpret the interval in the context of the question.
Last-day revision: Confidence intervals and prediction intervals
- A pivotal quantity is a function of the data and the parameter whose distribution is completely known and free of the parameter.
- A 95% confidence interval means 95% of intervals built this way would contain the true parameter. The parameter itself is fixed.
- Normal mean, variance known: x̄ ± z × σ ÷ √n.
- Normal mean, variance unknown: x̄ ± t(n−1) × s ÷ √n.
- Normal variance: (n−1)s² ÷ χ²(upper) to (n−1)s² ÷ χ²(lower), with n−1 degrees of freedom.
- The variance interval is not symmetric about s², and you take the square root of its limits to get one for σ.
- Two independent normal means with equal variances use a pooled variance and n₁ + n₂ − 2 degrees of freedom.
- Paired data: take differences and use a one-sample t interval on them.
- The ratio of two variances uses the F distribution, with s₁²/s₂² as the base.
- Large-sample proportion: p̂ ± z × √(p̂(1 − p̂) ÷ n).
- Prediction interval for a new normal observation: x̄ ± t × s × √(1 + 1/n).
- A prediction interval is always wider than the confidence interval for the mean, and it does not shrink to zero as n grows.
Confidence intervals and prediction intervals practice questions
- In a sample of 400 policyholders of a Pune-based insurer, 80 lapsed their policy within a year. Using the normal approximation, what is the …
- X1,...,Xn are independent N(mu, sigma^2) with both parameters unknown. Which of the following is a pivotal quantity for mu that can be used …
- Two branches of an Indian life insurer sample policy renewals. Branch A: 150 of 200 renewed. Branch B: 120 of 200 renewed. Using the normal …
- A random sample of 25 claim sizes (in Rs thousand) from a normal distribution has sample mean 120 and sample standard deviation 10. Using t(…
- For a sample of n from a normal distribution with unknown mean and variance, which statement about the 95% confidence interval for μ based o…
- A 95% confidence interval for a normal mean with known variance has width 6 using a sample of size 36. Keeping the same confidence level and…
- A 95% confidence interval for a proportion based on a sample of size n is (0.30, 0.40). The same sample proportion is used, but the sample s…
- For a 95% confidence interval for the ratio of two normal population variances sigma1^2/sigma2^2, based on s1^2 = 20 (n1 = 9) and s2^2 = 10 …
Confidence intervals and prediction intervals in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Confidence intervals and prediction intervals: frequently asked questions
What is the difference between a confidence interval and a prediction interval?
A confidence interval estimates a fixed parameter such as the population mean. A prediction interval estimates where a single new observation will fall. The prediction interval is wider because it includes the variability of the new observation as well as the uncertainty in the estimated mean.
When do I use t instead of z for a confidence interval?
Use t when the data are normal and the population variance is unknown, so you use the sample standard deviation. The degrees of freedom are n−1 for a single sample. For large samples the t and z values are close, but a normal-data question usually expects t.
How do I decide between pooled and unpooled two-sample intervals?
Use the pooled interval only when it is reasonable to assume that the two population variances are equal. The question may state this or you may check it with an F test or by comparing sample variances. If the assumption is not supported, use the unpooled approach.
Do I need to do confidence intervals in R for Paper B?
Paper B is a computer-based exam and can test inference in R, so you should be able to produce intervals from data. Practise functions such as t.test and then confirm that you can read the interval from the output. Also be able to compute an interval by hand from summary statistics.