Actuarial Statistics · Hypothesis testing and goodness of fit
Contingency Tables and Chi-Square Tests of Independence
Updated 11 October 2026 · Fact-checked
A contingency table counts items by two classification factors. The chi-square test of independence checks whether the factors are unrelated. Find each expected count as row total × column total ÷ grand total. Compute Σ (O − E)² ÷ E. Compare with chi-square on (r − 1)(c − 1) degrees of freedom.
Understand Contingency Tables and Tests of Independence
A contingency table shows counts of items classified by two factors. For example, policyholders classified by age band (rows) and claim or no claim (columns). Each cell holds the number of items with that combination.
The question is whether the two factors are independent. If they are, the proportion in each column is the same in every row. Knowing the age band would tell you nothing about claim behaviour.
The null hypothesis H0 is that the factors are independent. The alternative H1 is that they are associated. Under H0, P(row i and column j) = P(row i) × P(column j). Multiply by the grand total n and estimate the probabilities from the margins. This gives the expected count E = (row total × column total) ÷ n.
The test statistic compares observed counts O with expected counts E: X² = Σ (O − E)² ÷ E. A large value means the data sit far from independence. Under H0, X² is approximately chi-square distributed. You reject H0 only for large values, so the test is upper-tail.
The chi-square approximation works best when expected counts are not too small. A common rule of thumb is that each E should be at least 5. If not, combine adjacent categories where that makes sense.
Key rules to remember
- Expected count
- E(i,j) = (row i total × column j total) ÷ n
- n is the grand total. Row and column totals of the E table must equal those of the observed table.
- Test statistic
- X² = Σ (O − E)² ÷ E
- Sum over all r × c cells. Equivalent form: Σ O² ÷ E − n.
- Degrees of freedom
- ν = (r − 1)(c − 1)
- For an r by c table when expected counts are estimated from the margins. Nothing else is estimated.
- Decision rule
- Reject H0 if X² > upper critical value of χ²(ν) at the chosen significance level
- Upper-tail test. Equivalent: reject if the p-value is below the significance level.
- Hypotheses
- H0: the two factors are independent; H1: they are not independent
- Write them in words relating to the context.
How to solve Contingency Tables and Tests of Independence questions
Use this method for any independence test on a contingency table.
- 1State H0 (the two factors are independent) and H1 (they are associated), in the words of the question.
- 2Find row totals, column totals and the grand total n.
- 3Calculate each expected count as row total × column total ÷ n. Check that the expected counts add to the same totals.
- 4Check expected counts. If any are small (below about 5), combine categories and recount r or c.
- 5Calculate each (O − E)² ÷ E and sum them to get X².
- 6Find ν = (r − 1)(c − 1) using the final table size.
- 7Compare X² with the upper chi-square critical value at the stated level, and decide.
- 8Conclude in context. Where useful, comment on which cells contribute most to X².
Quickest way: Margins first, then cell contributions
When to use it: Use under time pressure on tables up to about 3 by 3, especially with calculator limits.
- Write the totals around the table once.
- Compute E for each cell and write it in brackets beside O.
- Compute (O − E)² ÷ E cell by cell in a single pass. Keep 2 or 3 decimals.
- Add the contributions. Note which cells are largest, as they explain the result.
- Use ν = (r − 1)(c − 1), read the table value, and write a one-line conclusion in context.
Common mistakes in Contingency Tables and Tests of Independence
Using degrees of freedom r × c − 1 or r × c.
Confusion with the goodness-of-fit test for a single set of categories.
Fix: For independence, estimating the margins removes r − 1 and c − 1 free values. Use (r − 1)(c − 1).
Calculating expected counts with the wrong totals, such as dividing by a row total.
The formula is memorised without the idea that it is n × P(row) × P(column).
Fix: Always use row total × column total ÷ grand total. Check that each E row and column adds to the observed totals.
Running the test on percentages or proportions instead of counts.
The table in the question shows percentages, or the student converts for convenience.
Fix: The statistic needs actual frequencies. Convert percentages back to counts using the sample size.
Ignoring small expected counts.
The calculation runs through without checking the E values.
Fix: Check E values after computing them. Combine sensible neighbouring categories, then recalculate ν from the reduced table.
Using a two-sided critical value or a lower tail.
Habit from z and t tests.
Fix: Large X² indicates departure from independence, so reject only in the upper tail.
Concluding that one factor causes the other.
Rejecting independence feels like proof of an effect.
Fix: State only that there is evidence of association between the factors at the chosen level.
Worked examples
Example 1
A sample of 200 policyholders is classified by vehicle type and whether a claim was made. Two-wheeler: 30 claim, 70 no claim. Car: 20 claim, 80 no claim. Test at the 5% level whether claim status is independent of vehicle type. The upper 5% point of χ² with 1 degree of freedom is 3.841.
Show the solution
- H0: claim status is independent of vehicle type. H1: they are associated.
- Row totals: two-wheeler 100, car 100. Column totals: claim 50, no claim 150. n = 200.
- E(two-wheeler, claim) = 100 × 50 ÷ 200 = 25. E(two-wheeler, no claim) = 100 × 150 ÷ 200 = 75. E(car, claim) = 25. E(car, no claim) = 75.
- Contributions: (30 − 25)² ÷ 25 = 1; (70 − 75)² ÷ 75 = 0.3333; (20 − 25)² ÷ 25 = 1; (80 − 75)² ÷ 75 = 0.3333.
- X² = 1 + 0.3333 + 1 + 0.3333 = 2.6667.
- ν = (2 − 1)(2 − 1) = 1. Critical value is 3.841.
- 2.667 < 3.841, so do not reject H0.
Answer: X² = 2.667 with 1 degree of freedom. This is below 3.841, so there is no significant evidence at the 5% level that claim status depends on vehicle type.
Example 2
Of 300 customers, each is classified by age group (Young, Middle, Old) and preferred premium payment mode (Annual, Monthly). Observed: Young: 40 Annual, 80 Monthly. Middle: 60 Annual, 40 Monthly. Old: 50 Annual, 30 Monthly. Test independence at the 5% level. The upper 5% point of χ² with 2 degrees of freedom is 5.991.
Show the solution
- H0: payment mode is independent of age group. H1: they are associated.
- Row totals: Young 120, Middle 100, Old 80. Column totals: Annual 150, Monthly 150. n = 350? Check: 120 + 100 + 80 = 300, so n = 300. Column totals: Annual 40 + 60 + 50 = 150; Monthly 80 + 40 + 30 = 150. Total 300.
- Expected Annual: Young 120 × 150 ÷ 300 = 60; Middle 100 × 150 ÷ 300 = 50; Old 80 × 150 ÷ 300 = 40. Expected Monthly are the same: 60, 50, 40.
- All expected counts are at least 5, so no combining is needed.
- Contributions Annual: (40 − 60)² ÷ 60 = 6.6667; (60 − 50)² ÷ 50 = 2; (50 − 40)² ÷ 40 = 2.5.
- Contributions Monthly: (80 − 60)² ÷ 60 = 6.6667; (40 − 50)² ÷ 50 = 2; (30 − 40)² ÷ 40 = 2.5.
- X² = 6.6667 + 2 + 2.5 + 6.6667 + 2 + 2.5 = 22.3333.
- ν = (3 − 1)(2 − 1) = 2. Critical value is 5.991. Since 22.33 > 5.991, reject H0.
- The Young group contributes most, preferring Monthly more than expected.
Answer: X² = 22.33 on 2 degrees of freedom, above 5.991. Reject H0: there is strong evidence at the 5% level that payment mode is associated with age group.
Exam tips
- Show the table of expected counts. Marks are usually given for the E values, the statistic and ν separately.
- State H0 and H1 in context, not just as symbols, and end with a conclusion in context.
- Check that the expected totals match the observed totals. It catches arithmetic slips quickly.
- If a question gives small expected counts, combine categories and say why. Then recalculate the degrees of freedom.
- In computer-based papers, show the R or Excel function used for the test, and still report X², ν and the decision.
Practice questions from Hypothesis testing and goodness of fit
- In a 2×2 table of 100 observations the chi-squared statistic for independence is calculated as 4.20. At the 5% significance level, which con…
- A sample of 10 claim amounts from a normal population has sample variance 18. To test H0: sigma^2 = 12 against H1: sigma^2 > 12, which stati…
- Which statement about the Kolmogorov-Smirnov test for a fully specified continuous distribution is correct?
- A test of H0: μ = 50 against H1: μ > 50 is carried out at the 5% significance level. Which statement correctly describes the probability of …
- A die is rolled 120 times to test whether it is fair. The observed frequencies for faces 1 to 6 are 15, 25, 20, 18, 22 and 20. What is the v…
Contingency Tables and Tests of Independence: frequently asked questions
How do I calculate expected frequencies in a contingency table?
Multiply the row total by the column total and divide by the grand total. Do this for every cell. The expected counts must add up to the same row and column totals as the observed table.
What are the degrees of freedom for an r by c contingency table?
They are (r − 1)(c − 1). For a 3 by 4 table this is 2 × 3 = 6. This applies when the expected counts are estimated from the table's own margins.
What does it mean if I reject the null hypothesis?
It means the data give evidence that the two factors are associated. It does not show that one causes the other. Look at the cells with the largest contributions to see where the association lies.
How is this different from the chi-square goodness of fit test?
Goodness of fit compares one set of observed counts with a specified or fitted distribution. The independence test compares a two-way table with the pattern implied by independence, and uses (r − 1)(c − 1) degrees of freedom.