Actuarial Statistics · Hypothesis testing and goodness of fit
Tests for Proportions and Likelihood Ratio Tests
Updated 11 October 2026 · Fact-checked
A test for a proportion checks whether a binomial probability equals a stated value, using an exact binomial or normal-approximation test. The Neyman-Pearson lemma gives the most powerful test for a simple null against a simple alternative. A likelihood ratio test compares maximised likelihoods, and −2 ln λ is approximately chi-square for large samples.
Understand Tests for Proportions and Likelihood Ratio Tests
A hypothesis test decides between a null hypothesis H0 and an alternative H1 using data. Two errors matter. The size (significance level) α is P(reject H0 | H0 true). The power is P(reject H0 | H1 true). A good test has a fixed small size and as much power as possible.
For a proportion, the number of successes X in n independent trials is Bin(n, p). To test H0: p = p0, you can use the exact binomial distribution of X under H0. For large n, with np0 and n(1 − p0) both reasonably large, you can use the normal approximation. The standard error under H0 uses p0, not the sample proportion, because you are assuming H0 is true.
The Neyman-Pearson lemma handles the simplest case: H0: θ = θ0 against H1: θ = θ1, both simple. The most powerful test of a given size rejects H0 when the likelihood ratio L(θ0) ÷ L(θ1) is small, that is, when the data are much more likely under H1. You then turn this into a condition on a simple statistic such as the sample mean or the count of successes. If the same rejection region works for every θ1 in a one-sided alternative, the test is uniformly most powerful for that alternative.
The likelihood ratio test (LRT) extends this to composite hypotheses. Let λ = max L under H0 ÷ max L over the full parameter space. Small λ is evidence against H0. For large samples and under regularity conditions, −2 ln λ is approximately chi-square with degrees of freedom equal to the number of free parameters in the full model minus the number under H0. The Wald test is a close cousin. It uses the squared distance between the MLE and the null value, divided by the estimated variance of the MLE. It also has an approximate chi-square distribution. The LRT needs the likelihood fitted under both hypotheses. The Wald test needs only the unrestricted MLE and its variance.
Key rules to remember
- Size and power
- α = P(reject H0 | H0 true); power = P(reject H0 | H1 true) = 1 − β
- β is the probability of a Type II error.
- Test statistic for a proportion (large n)
- z = (p̂ − p0) ÷ √(p0(1 − p0) ÷ n), where p̂ = x ÷ n
- Approximately N(0,1) under H0. Use p0 in the standard error. Use a continuity correction if you need to approximate the exact binomial tail.
- Exact binomial test
- Under H0, X ~ Bin(n, p0); p-value = P(X ≥ x) for H1: p > p0
- Use the lower tail for H1: p < p0. Because X is discrete, the exact size may be below the nominal α.
- Neyman-Pearson lemma
- Reject H0 if L(θ0; x) ÷ L(θ1; x) ≤ k, with k chosen so that the size equals α
- Applies to simple H0 against simple H1. For continuous data this gives the most powerful test of size α.
- Likelihood ratio statistic
- λ = L(θ̂0) ÷ L(θ̂); test statistic = −2 ln λ = 2[ln L(θ̂) − ln L(θ̂0)]
- θ̂0 is the MLE under H0 and θ̂ is the unrestricted MLE. Reject H0 for large −2 ln λ.
- Asymptotic distribution of the LRT
- −2 ln λ ≈ χ²(r), r = (free parameters in full model) − (free parameters under H0)
- Needs nested models and a large sample. Reject at level α if the statistic exceeds the upper α point of χ²(r).
- Binomial LRT for H0: p = p0
- −2 ln λ = 2[x ln(x ÷ (n p0)) + (n − x) ln((n − x) ÷ (n(1 − p0)))]
- Degrees of freedom = 1. Close to z² when n is large.
- Wald statistic (one parameter)
- W = (θ̂ − θ0)² ÷ Var̂(θ̂) ≈ χ²(1)
- Var̂ is usually from the inverse of the information at the MLE. Equivalent to a squared z-test.
How to solve Tests for Proportions and Likelihood Ratio Tests questions
Use this order for any question on proportions, Neyman-Pearson or likelihood ratio tests.
- 1Write H0 and H1 in terms of the parameter. State whether each is simple or composite and whether the test is one-sided or two-sided.
- 2Write down the likelihood (or the distribution of the statistic) and the log-likelihood if needed.
- 3For Neyman-Pearson, form L(θ0) ÷ L(θ1). Simplify and rearrange the inequality into a condition on a statistic, such as x̄ > c or x ≥ c. Watch the direction of the inequality when you divide or take logs.
- 4Find the critical value c from the size α, using the distribution of the statistic under H0. For discrete data, use the largest region with size not exceeding α, or state the exact size.
- 5For a likelihood ratio test, find the MLE under H0 and the unrestricted MLE. Compute −2 ln λ and find the degrees of freedom as the difference in free parameters.
- 6Compare with the critical value or find the p-value. State the decision in words, linked to the context.
- 7For power questions, compute the probability of falling in the rejection region under the stated H1 value.
Quickest way: Binomial proportion in two minutes
When to use it: Use when you are given x successes in n trials and a single null value p0, and the question asks for a test or a statistic.
- Compute p̂ = x ÷ n and check np0 and n(1 − p0) are both large enough for a normal approximation.
- Compute z = (p̂ − p0) ÷ √(p0(1 − p0) ÷ n). Compare with 1.645 (one-sided 5%) or 1.96 (two-sided 5%).
- If asked for the LRT, use the closed form 2[x ln(x ÷ (n p0)) + (n − x) ln((n − x) ÷ (n(1 − p0)))] and compare with χ²(1). The 5% point is 3.841.
- As a check, z² should be close to the LRT value. A big gap means an arithmetic slip or a small sample.
Common mistakes in Tests for Proportions and Likelihood Ratio Tests
Using the wrong direction in the Neyman-Pearson ratio, rejecting when L(θ0) ÷ L(θ1) is large.
Students remember 'ratio' but not which likelihood is on top.
Fix: Reject H0 when the data are more likely under H1. So reject when L(θ0) ÷ L(θ1) is small, or L(θ1) ÷ L(θ0) is large. Sanity-check with the sign of θ1 − θ0.
Using p̂ in the standard error for a hypothesis test on a proportion.
It is copied from the confidence interval formula.
Fix: For a test, the standard error is √(p0(1 − p0) ÷ n), because the distribution is computed under H0. Use p̂ in the standard error only for confidence intervals.
Taking the wrong degrees of freedom for the LRT.
Students use the sample size or the number of parameters in one model only.
Fix: Count free parameters in the full model and under H0, then subtract. Testing one parameter against one fixed value gives 1 degree of freedom.
Applying the chi-square approximation to models that are not nested, or with a very small sample.
The result is memorised without its conditions.
Fix: State that H0 is a special case of the full model and that n is large. If either condition fails, say the approximation may be poor.
Forgetting the discreteness of the binomial and quoting an exact size of 5%.
Continuous-data habits carry over.
Fix: Find the critical value c so that P(X ≥ c | H0) ≤ α, and quote the actual size obtained. Then state that the test is conservative.
Treating the Wald test and the LRT as identical.
Both give approximate chi-square statistics.
Fix: They agree asymptotically but not in finite samples. The LRT compares log-likelihoods at two fits. The Wald test uses only the unrestricted MLE and its estimated variance.
Worked examples
Example 1
A sample of n = 25 observations is taken from a normal distribution with known standard deviation 10. Use the Neyman-Pearson lemma to find the most powerful test of size 5% of H0: μ = 50 against H1: μ = 55. Find the power of the test.
Show the solution
- The likelihood ratio L(50) ÷ L(55) is a decreasing function of x̄ when μ1 > μ0. The ratio is small when x̄ is large. So the most powerful test rejects H0 when x̄ > c.
- Under H0, X̄ ~ N(50, 10² ÷ 25) = N(50, 2²), so the standard deviation of X̄ is 2.
- For size 5%, P(X̄ > c | μ = 50) = 0.05. So (c − 50) ÷ 2 = 1.645, giving c = 53.29.
- Under H1, X̄ ~ N(55, 2²). Power = P(X̄ > 53.29) = P(Z > (53.29 − 55) ÷ 2) = P(Z > −0.855).
- P(Z > −0.855) = Φ(0.855) ≈ 0.804.
Answer: Reject H0 if x̄ > 53.29. The power is about 0.804. The rejection region does not depend on the particular μ1 > 50, so this test is also uniformly most powerful against H1: μ > 50.
Example 2
In a sample of n = 100 policies, x = 60 lapsed in the first year. Test H0: p = 0.5 against H1: p ≠ 0.5 at the 5% level using a likelihood ratio test.
Show the solution
- The unrestricted MLE is p̂ = 60 ÷ 100 = 0.6. Under H0, p = 0.5, so n p0 = 50 and n(1 − p0) = 50.
- −2 ln λ = 2[60 ln(60 ÷ 50) + 40 ln(40 ÷ 50)].
- ln(1.2) = 0.18232 and ln(0.8) = −0.22314.
- 60 × 0.18232 = 10.939 and 40 × (−0.22314) = −8.926. The sum is 2.014.
- −2 ln λ = 2 × 2.014 = 4.03.
- The full model has 1 free parameter and H0 has 0, so the degrees of freedom are 1. The 5% point of χ²(1) is 3.841.
- As a check, z = (0.6 − 0.5) ÷ √(0.25 ÷ 100) = 2, so z² = 4, close to 4.03.
Answer: The statistic is about 4.03, which exceeds 3.841, so reject H0 at the 5% level. There is evidence that the first-year lapse proportion is different from 0.5.
Exam tips
- Show the Neyman-Pearson inequality being rearranged line by line. Marks go for the direction of the inequality and for naming the statistic.
- Always state the distribution of your test statistic under H0 before using a critical value.
- For LRT questions, write the log-likelihood at both MLEs first. Then compute −2 ln λ and state the degrees of freedom with the reason.
- Keep four decimal places in logs. A small rounding slip can change the decision when the statistic is close to the critical value.
- In the computer-based paper, do the same steps: state hypotheses, show the formula and working, and report the statistic, the critical value or p-value and the conclusion.
Practice questions from Hypothesis testing and goodness of fit
- In a chi-squared test of independence on a 3×3 table, several expected frequencies are below 5. Which action is the standard remedy?
- A sign test is applied to 12 paired differences (after-before) in monthly premium collections for a branch. Of the 12, none is zero; 9 are p…
- In a 2×2 table of 100 observations the chi-squared statistic for independence is calculated as 4.20. At the 5% significance level, which con…
- A Kolmogorov-Smirnov test compares an empirical distribution function with a fully specified continuous distribution. Which statement about …
- A sample of 10 claim amounts from a normal population has sample variance 18. To test H0: sigma^2 = 12 against H1: sigma^2 > 12, which stati…
Tests for Proportions and Likelihood Ratio Tests: frequently asked questions
What does the Neyman-Pearson lemma say?
For a simple null against a simple alternative, the most powerful test of a given size rejects H0 when the likelihood ratio L(θ0) ÷ L(θ1) is at most some constant k. You choose k so the size equals α. You then convert the condition into one on a sample statistic.
How many degrees of freedom does a likelihood ratio test have?
It is the number of free parameters in the full model minus the number free under H0. For testing one parameter against a single value it is 1. The result is asymptotic and requires nested models.
What is the difference between a Wald test and a likelihood ratio test?
The Wald test uses the squared difference between the MLE and the null value, divided by its estimated variance. The LRT compares the maximised log-likelihoods under H0 and the full model. They agree for large samples but can differ for small ones.
Do I use p̂ or p0 in the standard error when testing a proportion?
Use p0. A hypothesis test works out what happens if H0 is true, so the variance is p0(1 − p0) ÷ n. The sample proportion p̂ is used in confidence intervals.