IAI Actuarial Core Principles · Actuarial Statistics
Hypothesis Testing and Goodness of Fit Explained
Hypothesis testing checks whether sample data are consistent with a stated claim about a population. You set H₀ and H₁, pick a test statistic with a known distribution under H₀, find a p-value or critical value, and conclude. Goodness of fit tests ask whether data follow a specified distribution or whether two factors are independent.
What this chapter covers
This chapter covers how to turn a question about data into a formal decision. You start with the framework: null and alternative hypotheses, test statistic, significance level, p-value, critical region, and Type I and Type II errors. You then apply it to means, variances, proportions and likelihood ratio tests.
The second half moves to fit. The chi-square goodness of fit test compares observed and expected counts. Contingency tables extend that idea to independence. Kolmogorov-Smirnov, runs and similar tests check fit in other ways. Non-parametric and permutation tests are used when normality or other assumptions are doubtful.
The chapter links to the rest of the paper in several ways. It builds on random variables, sampling distributions and estimation. It supports regression, where you test coefficients and model fit. In CS2 it is used to check whether a claims distribution, time series residuals or survival model fits the data. Treat it as a toolkit you will reuse across the Actuarial Statistics module.
In the CS1 syllabus, statistical inference carries a 25% weighting, and testing and fit methods also appear inside regression, Bayesian and CS2 modelling questions. Both the multiple-choice section and the written section reward this chapter. MCQs test whether you can pick the right test and read a result. Written questions ask you to state hypotheses, show the working, give the conclusion in context and comment on assumptions. Each step earns marks, so a method you know well is a reliable source of them. Paper B, the computer-based exam, also expects you to run tests in R and interpret the output.
Hypothesis testing and goodness of fit: topics in the order to study them
- 1Hypothesis Testing FrameworkEvery other topic uses its language: hypotheses, test statistic, p-value, critical region and error types.
- 2Tests for Means and VariancesThese are the standard normal, t and chi-square tests, and they show the framework in its simplest form.
- 3Tests for Proportions and Likelihood Ratio TestsProportions use approximations you must justify, and the likelihood ratio test gives a general method that builds on estimation.
- 4Chi-Square Goodness of Fit TestIt introduces observed versus expected counts, degrees of freedom and the rule on merging small expected frequencies.
- 5Contingency Tables and Tests of IndependenceIt reuses the chi-square statistic, so learn it straight after goodness of fit, with expected counts from row and column totals.
- 6Kolmogorov-Smirnov, Runs and Other Fit TestsThese tests check fit differently, using the empirical distribution function and the order of observations, so they make sense once chi-square is clear.
- 7Non-Parametric and Permutation TestsStudy these last, as they replace distributional assumptions and are best understood after you know what the parametric tests assume.
How to prepare Hypothesis testing and goodness of fit
Aim to recognise the right test from the question, then carry it out in a fixed sequence. Practise both by hand and in R.
- Learn the framework first. Be able to define H₀, H₁, test statistic, p-value, significance level, Type I error and Type II error in your own words.
- Build a one-page table of tests. For each, note when it applies, the statistic, its distribution under H₀ and the degrees of freedom.
- Use one written layout every time: hypotheses, statistic, distribution under H₀, calculation, p-value or critical value, conclusion in context.
- Practise the chi-square tests with real counts. Work out expected frequencies, merge cells where needed, and count degrees of freedom after subtracting estimated parameters.
- Attempt past-style MCQs on choosing a test and reading p-values, and time yourself.
- Repeat key tests in R. Run them on data, read the output, and write a short interpretation including assumptions.
- Finish with mixed questions and review each wrong answer by asking whether the error was test choice, calculation or interpretation.
Common mistakes in Hypothesis testing and goodness of fit
Saying H₀ is accepted or proved true.
Fix: Write 'there is insufficient evidence to reject H₀ at the 5% level' and then state what that means in context.
Using the wrong degrees of freedom in chi-square tests.
Fix: Count cells after merging, then subtract 1 and the number of parameters estimated from the data.
Skipping the merging of cells with small expected frequencies.
Fix: Check every expected count first and combine adjacent categories where needed before calculating.
Choosing the wrong test for a mean or a variance.
Fix: Before calculating, state what is known, what is assumed and why the chosen distribution applies.
Using the wrong tail or doubling the p-value incorrectly.
Fix: Write H₁ first. A two-sided H₁ needs both tails. A one-sided H₁ needs only the relevant tail.
Giving a conclusion without context or assumptions.
Fix: Finish with a sentence about the actual claim, and mention any assumptions that could weaken the result.
Last-day revision: Hypothesis testing and goodness of fit
- H₀ is the default claim. You reject it or fail to reject it; you never prove it true.
- The p-value is the probability, assuming H₀ is true, of a result at least as extreme as observed.
- Reject H₀ when the p-value is less than or equal to the significance level.
- Type I error is rejecting a true H₀. Its probability is the significance level.
- Type II error is not rejecting a false H₀. Power = 1 − P(Type II error).
- Use a t test for a mean when σ is unknown and the data are normal. Use the normal when σ is known.
- The variance test uses (n − 1)S² ÷ σ₀², which is χ² with n − 1 degrees of freedom under normality.
- Chi-square statistic = Σ (O − E)² ÷ E. Merge cells with small expected counts, commonly below 5.
- Goodness of fit degrees of freedom = number of cells − 1 − number of estimated parameters.
- For an r × c contingency table, E = row total × column total ÷ grand total, with (r − 1)(c − 1) degrees of freedom.
- The Kolmogorov-Smirnov statistic is the largest gap between the empirical and hypothesised distribution functions.
- Permutation tests build the null distribution by reshuffling labels, so they need few distributional assumptions.
Hypothesis testing and goodness of fit practice questions
- In a chi-squared test of independence on a 3×3 table, several expected frequencies are below 5. Which action is the standard remedy?
- A permutation test compares mean claim settlement times of two groups, 3 observations in group X and 4 in group Y (7 in total, no ties). The…
- A sign test is applied to 12 paired differences (after-before) in monthly premium collections for a branch. Of the 12, none is zero; 9 are p…
- In a 2×2 table of 100 observations the chi-squared statistic for independence is calculated as 4.20. At the 5% significance level, which con…
- A sample of 10 claim amounts from a normal population has sample variance 18. To test H0: sigma^2 = 12 against H1: sigma^2 > 12, which stati…
- Which statement about the Kolmogorov-Smirnov test for a fully specified continuous distribution is correct?
- A test of H0: μ = 50 against H1: μ > 50 is carried out at the 5% significance level. Which statement correctly describes the probability of …
- A die is rolled 120 times to test whether it is fair. The observed frequencies for faces 1 to 6 are 15, 25, 20, 18, 22 and 20. What is the v…
Hypothesis testing and goodness of fit in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Hypothesis testing and goodness of fit: frequently asked questions
Which tests from this chapter are most useful for CS1 and CS2?
The framework, t and chi-square tests, the chi-square goodness of fit and independence tests, and likelihood ratio tests are used most widely. They also support work in regression, time series and survival models. Learn these thoroughly before moving to the less common fit tests.
How do I decide between a p-value and a critical value?
Both give the same decision at the same significance level. Use the p-value when software gives it or tables allow a good estimate. Use the critical value when you must work from tables by hand.
When should I use a non-parametric or permutation test?
Use them when the assumptions of a parametric test, such as normality, are doubtful or the sample is small. They make fewer assumptions about the distribution. State the reason for choosing them in your answer.
How should I prepare for the R part of this chapter?
Run each test on small data sets and read every part of the output: statistic, degrees of freedom and p-value. Then write a short interpretation in words. Practise this until you can do it without looking up the function names.