Actuarial Statistics · Hypothesis testing and goodness of fit
Hypothesis Testing Framework: Errors, p-values and Power
Updated 11 October 2026 · Fact-checked
A hypothesis test decides whether sample data give enough evidence to reject a null hypothesis H₀ in favour of an alternative H₁. You compute a test statistic, compare it with a critical value or find a p-value, and reject H₀ if the p-value is at most the significance level α.
Understand Hypothesis Testing Framework
A hypothesis test is a rule for choosing between two claims about a population parameter, using sample data. The null hypothesis H₀ is the default claim, such as μ = 50. The alternative hypothesis H₁ is what you accept if the data contradict H₀, such as μ > 50 or μ ≠ 50.
You measure the gap between the data and H₀ with a test statistic. Its distribution is known when H₀ is true. If the statistic lands somewhere very unlikely under H₀, you doubt H₀. The set of statistic values that lead to rejection is the critical region (rejection region). Its boundary is the critical value.
The significance level α is the probability of rejecting H₀ when H₀ is true. You choose it before seeing the data, often 5% or 1%. The p-value is the probability, assuming H₀ is true, of getting a test statistic at least as extreme as the one observed. Small p-value means strong evidence against H₀. You reject H₀ if p-value ≤ α.
Two errors are possible. A Type I error is rejecting H₀ when it is true. Its probability is α. A Type II error is not rejecting H₀ when it is false. Its probability is β. The power of a test is 1 − β: the probability of rejecting H₀ when the true parameter is a specific value in H₁. Power depends on the true value, so it is a function of the parameter.
Failing to reject H₀ is not proof that H₀ is true. It only means the data did not give enough evidence against it. For a fixed sample size, lowering α makes β larger. A larger sample size can reduce both.
Key rules to remember
- Hypotheses
- H₀: θ = θ₀ against H₁: θ ≠ θ₀ (two-sided), θ > θ₀ or θ < θ₀ (one-sided)
- H₀ is a simple statement with equality. H₁ direction comes from the question wording, not from the data.
- Significance level
- α = P(reject H₀ | H₀ true)
- This is the Type I error probability. Set it before looking at the data.
- Type II error
- β(θ) = P(do not reject H₀ | true parameter is θ, θ in H₁)
- Depends on the true value θ.
- Power
- Power(θ) = 1 − β(θ) = P(reject H₀ | θ)
- At θ = θ₀ the power function equals α.
- p-value (one-sided upper)
- p = P(T ≥ t_obs | H₀)
- For lower-tailed use P(T ≤ t_obs). For two-sided with a symmetric distribution, p = 2 × P(T ≥ |t_obs|).
- Decision rule
- Reject H₀ if p ≤ α, or if t_obs lies in the critical region
- Both methods give the same decision for the same test.
- z statistic for a mean, known σ
- Z = (X̄ − μ₀) ÷ (σ ÷ √n) ~ N(0,1) under H₀
- Use the t statistic with n − 1 degrees of freedom when σ is estimated by S and the data are normal.
How to solve Hypothesis Testing Framework questions
Use the same structure for any hypothesis test question. Marks are given for each stage, so write them all down.
- 1State H₀ and H₁ in terms of the parameter. Decide from the wording whether H₁ is one-sided or two-sided.
- 2State the test statistic and its distribution under H₀, with any assumptions (normality, known variance, independence).
- 3Calculate the observed value of the test statistic from the data.
- 4Find the critical value for the chosen α, or calculate the p-value. For two-sided tests, split α between both tails.
- 5Compare and decide: reject H₀ if the statistic is in the critical region or p ≤ α.
- 6Write the conclusion in context, in words. Say 'there is sufficient evidence' or 'insufficient evidence', not 'H₀ is proved'.
- 7For error or power questions, find the critical region in terms of the statistic, then compute probabilities under the stated true parameter value.
Quickest way: Critical value in terms of the sample mean
When to use it: When a question asks for the critical region, Type II error or power for a normal mean with known variance.
- Write the decision rule as X̄ > c (or < c, or both tails) for the test.
- Find c from H₀: c = μ₀ + z × σ ÷ √n, using z from the normal table for α (or α ÷ 2 per tail).
- For power at a true mean μ₁, standardise c using μ₁: P(X̄ > c | μ₁) = P(Z > (c − μ₁) ÷ (σ ÷ √n)).
- Then β = 1 − power. Check your sign: power should rise as μ₁ moves further from μ₀ in the direction of H₁.
Common mistakes in Hypothesis Testing Framework
Saying 'accept H₀' when you fail to reject it.
It sounds like the natural opposite of rejecting.
Fix: Write 'do not reject H₀' and add that there is insufficient evidence against it.
Treating the p-value as the probability that H₀ is true.
The wording 'probability' is easy to misapply.
Fix: Define it as the probability, assuming H₀ is true, of a result at least as extreme as observed.
Using a one-sided critical value for a two-sided test, or forgetting to double the tail probability.
Students skip checking whether H₁ uses ≠.
Fix: If H₁ has ≠, use α ÷ 2 in each tail, or p = 2 × the tail probability.
Mixing up Type I and Type II errors, or computing β using H₀'s mean.
The names are similar and the conditioning is easy to forget.
Fix: Type I: reject when H₀ true, probability α. Type II: do not reject when H₁ true. Compute β using the true value in H₁.
Choosing H₁ after looking at the data.
The sample mean seems to point in a direction.
Fix: Set the direction from the question's wording before calculating anything.
Using z when the variance is unknown and the sample is small.
Students copy the known-variance formula.
Fix: If σ is estimated by S from a normal sample, use the t distribution with n − 1 degrees of freedom.
Worked examples
Example 1
A claims manager believes mean settlement time is 20 days. A sample of 36 claims has mean 22 days. Assume settlement times are normal with known standard deviation 6 days. Test H₀: μ = 20 against H₁: μ > 20 at the 5% level, and find the p-value.
Show the solution
- H₀: μ = 20, H₁: μ > 20 (one-sided upper).
- Under H₀, Z = (X̄ − 20) ÷ (6 ÷ √36) ~ N(0,1).
- Standard error = 6 ÷ 6 = 1.
- Observed z = (22 − 20) ÷ 1 = 2.
- Critical value at 5% one-sided is 1.645. Since 2 > 1.645, z is in the critical region.
- p-value = P(Z ≥ 2) = 1 − 0.9772 = 0.0228, which is less than 0.05.
Answer: Reject H₀ at the 5% level. The p-value is about 0.0228. There is sufficient evidence that mean settlement time exceeds 20 days.
Example 2
For X ~ N(μ, 16) with n = 16, test H₀: μ = 50 against H₁: μ > 50 using the rule 'reject H₀ if X̄ > c' with α = 0.05. Find c, and find the power of the test when the true mean is 52.
Show the solution
- Standard error = √16 ÷ √16 = 4 ÷ 4 = 1.
- Under H₀, X̄ ~ N(50, 1). For α = 0.05, c = 50 + 1.645 × 1 = 51.645.
- Power at μ = 52: X̄ ~ N(52, 1).
- P(X̄ > 51.645) = P(Z > (51.645 − 52) ÷ 1) = P(Z > −0.355).
- P(Z > −0.355) = Φ(0.355) ≈ 0.6387.
- So β ≈ 1 − 0.6387 = 0.3613.
Answer: Critical value c = 51.645. Power at μ = 52 is about 0.639, so the Type II error probability is about 0.361.
Exam tips
- Always write H₀ and H₁ in terms of the parameter, with the test statistic and its null distribution. These are easy marks.
- Finish with a conclusion in context. Many students lose the last mark by stopping at 'reject H₀'.
- For power questions, standardise using the true mean from H₁, not the null mean.
- In multiple-choice questions, check whether the test is one-sided or two-sided before reading the table.
- In the computer-based paper, quote the p-value from R output and link it to α before concluding.
Practice questions from Hypothesis testing and goodness of fit
- A Kolmogorov-Smirnov test compares an empirical distribution function with a fully specified continuous distribution. Which statement about …
- A test of H0: μ = 50 against H1: μ > 50 is carried out at the 5% significance level. Which statement correctly describes the probability of …
- A die is rolled 120 times to test whether it is fair. The observed frequencies for faces 1 to 6 are 15, 25, 20, 18, 22 and 20. What is the v…
- In a chi-square goodness of fit test, an expected frequency in one tail cell is only 1.2 while all others exceed 5. What is the standard rec…
- A sample of n = 25 observations from a normal distribution with known standard deviation 10 has sample mean 54. Test H0: μ = 50 against H1: …
Hypothesis Testing Framework in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Hypothesis Testing Framework: frequently asked questions
What is the difference between Type I and Type II error?
A Type I error is rejecting H₀ when it is true, with probability α. A Type II error is failing to reject H₀ when it is false, with probability β. Power is 1 − β.
How do I calculate a p-value in a hypothesis test?
Find the observed test statistic, then calculate the probability under H₀ of a value at least as extreme. For a two-sided test with a symmetric distribution, double the tail probability. Reject H₀ if p ≤ α.
What is the critical region and how does it relate to the significance level?
The critical region is the set of test statistic values that lead to rejection of H₀. It is chosen so that its probability under H₀ equals α. A smaller α gives a smaller critical region.
Does a large sample size change power?
Yes. For a fixed α and a true parameter in H₁, a larger sample reduces the standard error and increases power. Type II error falls as a result.