Skip to content

IAI Actuarial Core Principles · Actuarial Statistics

Hypothesis testing and goodness of fit: formula sheet

Full chapter guide

Key formulas

Hypotheses
H₀: θ = θ₀ against H₁: θ ≠ θ₀ (two-sided), θ > θ₀ or θ < θ₀ (one-sided)
H₀ is a simple statement with equality. H₁ direction comes from the question wording, not from the data.
Significance level
α = P(reject H₀ | H₀ true)
This is the Type I error probability. Set it before looking at the data.
Type II error
β(θ) = P(do not reject H₀ | true parameter is θ, θ in H₁)
Depends on the true value θ.
Power
Power(θ) = 1 − β(θ) = P(reject H₀ | θ)
At θ = θ₀ the power function equals α.
p-value (one-sided upper)
p = P(T ≥ t_obs | H₀)
For lower-tailed use P(T ≤ t_obs). For two-sided with a symmetric distribution, p = 2 × P(T ≥ |t_obs|).
Decision rule
Reject H₀ if p ≤ α, or if t_obs lies in the critical region
Both methods give the same decision for the same test.
z statistic for a mean, known σ
Z = (X̄ − μ₀) ÷ (σ ÷ √n) ~ N(0,1) under H₀
Use the t statistic with n − 1 degrees of freedom when σ is estimated by S and the data are normal.
One sample mean, σ known
Z = (X̄ − μ0) ÷ (σ ÷ √n) ~ N(0, 1)
Use for normal data with known σ. For large n with unknown σ, a z approximation is common but state it.
One sample mean, σ unknown
T = (X̄ − μ0) ÷ (S ÷ √n) ~ t(n − 1)
S² = Σ(xi − x̄)² ÷ (n − 1). Degrees of freedom are n − 1.
One sample variance
χ² = (n − 1)S² ÷ σ0² ~ χ²(n − 1)
Needs normal data. Upper-tail test for H1: σ² > σ0², lower-tail for H1: σ² < σ0².
Two independent means, equal variances
T = (X̄ − Ȳ − δ0) ÷ (Sp × √(1/n + 1/m)) ~ t(n + m − 2)
δ0 is usually 0. Pooled variance Sp² = [(n − 1)SX² + (m − 1)SY²] ÷ (n + m − 2).
Two independent means, variances known
Z = (X̄ − Ȳ − δ0) ÷ √(σX²/n + σY²/m) ~ N(0, 1)
Use when both population variances are given.
Two independent means, variances unequal
T = (X̄ − Ȳ − δ0) ÷ √(SX²/n + SY²/m)
Welch approach: the distribution is only approximately t. Use it only if the question tells you to or the F-test rejects equal variances.
F-test for variance ratio
F = (SX² ÷ σX²) ÷ (SY² ÷ σY²) ~ F(n − 1, m − 1)
Under H0: σX² = σY², F = SX² ÷ SY². The first degrees of freedom belong to the numerator.
Lower F critical value
F(1 − α; a, b) = 1 ÷ F(α; b, a)
Tables often give only upper points. Use this to get the lower point. Here F(α; a, b) is the upper α point.
Paired t-test
Di = Xi − Yi; T = (D̄ − δ0) ÷ (SD ÷ √n) ~ t(n − 1)
n is the number of pairs. Differences must be roughly normal.
Size and power
α = P(reject H0 | H0 true); power = P(reject H0 | H1 true) = 1 − β
β is the probability of a Type II error.
Test statistic for a proportion (large n)
z = (p̂ − p0) ÷ √(p0(1 − p0) ÷ n), where p̂ = x ÷ n
Approximately N(0,1) under H0. Use p0 in the standard error. Use a continuity correction if you need to approximate the exact binomial tail.
Exact binomial test
Under H0, X ~ Bin(n, p0); p-value = P(X ≥ x) for H1: p > p0
Use the lower tail for H1: p < p0. Because X is discrete, the exact size may be below the nominal α.
Neyman-Pearson lemma
Reject H0 if L(θ0; x) ÷ L(θ1; x) ≤ k, with k chosen so that the size equals α
Applies to simple H0 against simple H1. For continuous data this gives the most powerful test of size α.
Likelihood ratio statistic
λ = L(θ̂0) ÷ L(θ̂); test statistic = −2 ln λ = 2[ln L(θ̂) − ln L(θ̂0)]
θ̂0 is the MLE under H0 and θ̂ is the unrestricted MLE. Reject H0 for large −2 ln λ.
Asymptotic distribution of the LRT
−2 ln λ ≈ χ²(r), r = (free parameters in full model) − (free parameters under H0)
Needs nested models and a large sample. Reject at level α if the statistic exceeds the upper α point of χ²(r).
Binomial LRT for H0: p = p0
−2 ln λ = 2[x ln(x ÷ (n p0)) + (n − x) ln((n − x) ÷ (n(1 − p0)))]
Degrees of freedom = 1. Close to z² when n is large.
Wald statistic (one parameter)
W = (θ̂ − θ0)² ÷ Var̂(θ̂) ≈ χ²(1)
Var̂ is usually from the inverse of the information at the MLE. Equivalent to a squared z-test.
Expected frequency
E = n × p
n is the total number of observations. p is the probability of the cell under H0. Check that ΣE = ΣO = n.
Test statistic
X² = Σ (O − E)² ÷ E
Sum over all cells after combining. Large values indicate a poor fit.
Alternative form
X² = Σ (O² ÷ E) − n
Useful for quick calculation. It gives the same value if you do not round too early.
Degrees of freedom
ν = k − 1 − m
k is the number of cells after combining. m is the number of parameters estimated from the data. Use m = 0 if the model is fully specified.
Decision rule
Reject H0 if X² > χ²(ν) upper critical value at level α
Equivalently, reject if the p-value P(χ²(ν) > X²) is less than α.
Cell size rule of thumb
Combine adjacent cells until E ≥ 5
A guideline for the chi-square approximation. Combine cells that are adjacent or logically similar, such as tail cells.
Expected count
E(i,j) = (row i total × column j total) ÷ n
n is the grand total. Row and column totals of the E table must equal those of the observed table.
Test statistic
X² = Σ (O − E)² ÷ E
Sum over all r × c cells. Equivalent form: Σ O² ÷ E − n.
Degrees of freedom
ν = (r − 1)(c − 1)
For an r by c table when expected counts are estimated from the margins. Nothing else is estimated.
Decision rule
Reject H0 if X² > upper critical value of χ²(ν) at the chosen significance level
Upper-tail test. Equivalent: reject if the p-value is below the significance level.
Hypotheses
H0: the two factors are independent; H1: they are not independent
Write them in words relating to the context.
Empirical distribution function
F_n(x) = (number of sample values ≤ x) ÷ n
A step function. It jumps by 1/n at each distinct observation.
KS statistic
D = max over x of |F_n(x) − F0(x)|
Check at each ordered value x(i) both |i/n − F0(x(i))| and |(i−1)/n − F0(x(i))|. Take the largest of all.
KS decision rule
Reject H0 if D > critical value for n and the chosen significance level
Use the tables you are given. For large n the 5% critical value is roughly 1.36 ÷ √n. Check the tables book before relying on this.
Anderson-Darling statistic
A² = −n − (1/n) Σ (2i − 1)[ln z_i + ln(1 − z_(n+1−i))], where z_i = F0(x(i))
Ordered data x(1) ≤ … ≤ x(n). Larger A² means worse fit, with extra weight on the tails.
Signs test
Under H0, number of positive residuals ~ Binomial(n, ½)
Drop zero residuals and reduce n. Use a normal approximation with mean n/2 and variance n/4 for large n.
Runs test (mean)
E(R) = 2·n1·n2 ÷ (n1 + n2) + 1
n1 positive and n2 negative residuals. A run is an unbroken sequence of the same sign.
Runs test (variance)
Var(R) = 2·n1·n2·(2·n1·n2 − n1 − n2) ÷ [(n1 + n2)²·(n1 + n2 − 1)]
For large n1 and n2, R is approximately normal. Use z = (R ± 0.5 − E(R)) ÷ √Var(R) with a continuity correction.
Lag-1 serial correlation
r1 = Σ (e_t − ē)(e_(t+1) − ē) ÷ Σ (e_t − ē)²
Numerator sum runs over t = 1 to n−1. Under independence, r1 is approximately N(0, 1/n), so compare |r1| with 1.96 ÷ √n at 5%.
Q-Q plot
Plot ordered sample values x(i) against theoretical quantiles of the fitted distribution
Points near a straight line mean good fit. For normality, plot against standard normal quantiles.
Sign test statistic
S = number of positive differences; under H0, S ~ Binomial(n, ½)
Drop zero differences first and reduce n. For a test of a median m0, take differences x − m0. Use the binomial table or tail probabilities.
Wilcoxon signed-rank statistic
T+ = sum of ranks of positive differences; T− = sum of ranks of negative differences; T+ + T− = n(n + 1) ÷ 2
Rank the absolute differences, giving tied values the average rank. Drop zeros. Tables usually use the smaller of T+ and T−, and you reject if it is at or below the critical value.
Normal approximation for signed-rank
E(T) = n(n + 1) ÷ 4; Var(T) = n(n + 1)(2n + 1) ÷ 24; Z = (T − E(T)) ÷ √Var(T)
Use for larger n, with a continuity correction of 0.5. Ties need an adjusted variance.
Mann-Whitney U
U_X = R_X − m(m + 1) ÷ 2; U_X + U_Y = mn
X has size m, Y has size n, and R_X is the sum of X's ranks in the combined ranking. Check with U_X + U_Y = mn. Some books use W = R_X (rank-sum form) instead of U; the tests are equivalent.
Mann-Whitney normal approximation
E(U) = mn ÷ 2; Var(U) = mn(m + n + 1) ÷ 12
For the rank sum W = R_X: E(W) = m(m + n + 1) ÷ 2 and the variance is the same as for U.
Spearman rank correlation
r_s = 1 − 6Σd² ÷ (n(n² − 1))
d is the difference between the two ranks of each item. This shortcut holds only when there are no ties. With ties, compute Pearson's correlation on the ranks.
Kendall's tau
τ = (C − D) ÷ (n(n − 1) ÷ 2)
C is the number of concordant pairs and D the number of discordant pairs, with no ties. Values lie between −1 and +1. Under H0 of independence, Var(τ) = 2(2n + 5) ÷ (9n(n − 1)) for the normal approximation.
Permutation test p-value
p = (number of permutations with statistic at least as extreme as observed) ÷ (total number of permutations)
For a two-sided test, count extreme values in both tails. If you sample random permutations, the p-value is only an estimate.

Quick revision

  • H₀ is the default claim. You reject it or fail to reject it; you never prove it true.
  • The p-value is the probability, assuming H₀ is true, of a result at least as extreme as observed.
  • Reject H₀ when the p-value is less than or equal to the significance level.
  • Type I error is rejecting a true H₀. Its probability is the significance level.
  • Type II error is not rejecting a false H₀. Power = 1 − P(Type II error).
  • Use a t test for a mean when σ is unknown and the data are normal. Use the normal when σ is known.
  • The variance test uses (n − 1)S² ÷ σ₀², which is χ² with n − 1 degrees of freedom under normality.
  • Chi-square statistic = Σ (O − E)² ÷ E. Merge cells with small expected counts, commonly below 5.
  • Goodness of fit degrees of freedom = number of cells − 1 − number of estimated parameters.
  • For an r × c contingency table, E = row total × column total ÷ grand total, with (r − 1)(c − 1) degrees of freedom.
  • The Kolmogorov-Smirnov statistic is the largest gap between the empirical and hypothesised distribution functions.
  • Permutation tests build the null distribution by reshuffling labels, so they need few distributional assumptions.

Common mistakes

  • Saying 'accept H₀' when you fail to reject it. Fix: Write 'do not reject H₀' and add that there is insufficient evidence against it.
  • Treating the p-value as the probability that H₀ is true. Fix: Define it as the probability, assuming H₀ is true, of a result at least as extreme as observed.
  • Using a two-sample t-test on paired data. Fix: Ask whether each value in one sample belongs with one specific value in the other. If so, take differences and do a one-sample t-test with n − 1 degrees of freedom.
  • Using n instead of n − 1 in the sample variance or the degrees of freedom. Fix: Sample variance S² divides by n − 1. A one-sample t or chi-square has n − 1 degrees of freedom. A pooled t has n + m − 2.
  • Using the wrong direction in the Neyman-Pearson ratio, rejecting when L(θ0) ÷ L(θ1) is large. Fix: Reject H0 when the data are more likely under H1. So reject when L(θ0) ÷ L(θ1) is small, or L(θ1) ÷ L(θ0) is large. Sanity-check with the sign of θ1 − θ0.
  • Using p̂ in the standard error for a hypothesis test on a proportion. Fix: For a test, the standard error is √(p0(1 − p0) ÷ n), because the distribution is computed under H0. Use p̂ in the standard error only for confidence intervals.
  • Forgetting to subtract degrees of freedom for estimated parameters Fix: Always ask: did I estimate anything from this data? If yes, subtract one for each parameter. A Poisson with an estimated mean loses one more degree of freedom.
  • Counting cells before combining Fix: Combine first. Then count k from the final table. Combining reduces ν.
  • Using degrees of freedom r × c − 1 or r × c. Fix: For independence, estimating the margins removes r − 1 and c − 1 free values. Use (r − 1)(c − 1).
  • Calculating expected counts with the wrong totals, such as dividing by a row total. Fix: Always use row total × column total ÷ grand total. Check that each E row and column adds to the observed totals.

Exam tips

  • Always write H₀ and H₁ in terms of the parameter, with the test statistic and its null distribution. These are easy marks.
  • Finish with a conclusion in context. Many students lose the last mark by stopping at 'reject H₀'.
  • For power questions, standardise using the true mean from H₁, not the null mean.
  • In multiple-choice questions, check whether the test is one-sided or two-sided before reading the table.
  • In the computer-based paper, quote the p-value from R output and link it to α before concluding.
  • Write H0, H1, the statistic's distribution and its degrees of freedom every time. Method marks are awarded for each, even if your arithmetic slips.
  • State your assumptions explicitly: normality, independence, and equal variances for pooling. Examiners look for them.
  • In computer-based questions, show the R call or Excel formula you used, such as t.test with paired = TRUE or var.test, and read off the statistic, degrees of freedom and p-value. Then write the conclusion in words.