IAI Actuarial Core Principles · Actuarial Statistics
Hypothesis testing and goodness of fit: formula sheet
Key formulas
- Hypotheses
- H₀: θ = θ₀ against H₁: θ ≠ θ₀ (two-sided), θ > θ₀ or θ < θ₀ (one-sided)
- H₀ is a simple statement with equality. H₁ direction comes from the question wording, not from the data.
- Significance level
- α = P(reject H₀ | H₀ true)
- This is the Type I error probability. Set it before looking at the data.
- Type II error
- β(θ) = P(do not reject H₀ | true parameter is θ, θ in H₁)
- Depends on the true value θ.
- Power
- Power(θ) = 1 − β(θ) = P(reject H₀ | θ)
- At θ = θ₀ the power function equals α.
- p-value (one-sided upper)
- p = P(T ≥ t_obs | H₀)
- For lower-tailed use P(T ≤ t_obs). For two-sided with a symmetric distribution, p = 2 × P(T ≥ |t_obs|).
- Decision rule
- Reject H₀ if p ≤ α, or if t_obs lies in the critical region
- Both methods give the same decision for the same test.
- z statistic for a mean, known σ
- Z = (X̄ − μ₀) ÷ (σ ÷ √n) ~ N(0,1) under H₀
- Use the t statistic with n − 1 degrees of freedom when σ is estimated by S and the data are normal.
- One sample mean, σ known
- Z = (X̄ − μ0) ÷ (σ ÷ √n) ~ N(0, 1)
- Use for normal data with known σ. For large n with unknown σ, a z approximation is common but state it.
- One sample mean, σ unknown
- T = (X̄ − μ0) ÷ (S ÷ √n) ~ t(n − 1)
- S² = Σ(xi − x̄)² ÷ (n − 1). Degrees of freedom are n − 1.
- One sample variance
- χ² = (n − 1)S² ÷ σ0² ~ χ²(n − 1)
- Needs normal data. Upper-tail test for H1: σ² > σ0², lower-tail for H1: σ² < σ0².
- Two independent means, equal variances
- T = (X̄ − Ȳ − δ0) ÷ (Sp × √(1/n + 1/m)) ~ t(n + m − 2)
- δ0 is usually 0. Pooled variance Sp² = [(n − 1)SX² + (m − 1)SY²] ÷ (n + m − 2).
- Two independent means, variances known
- Z = (X̄ − Ȳ − δ0) ÷ √(σX²/n + σY²/m) ~ N(0, 1)
- Use when both population variances are given.
- Two independent means, variances unequal
- T = (X̄ − Ȳ − δ0) ÷ √(SX²/n + SY²/m)
- Welch approach: the distribution is only approximately t. Use it only if the question tells you to or the F-test rejects equal variances.
- F-test for variance ratio
- F = (SX² ÷ σX²) ÷ (SY² ÷ σY²) ~ F(n − 1, m − 1)
- Under H0: σX² = σY², F = SX² ÷ SY². The first degrees of freedom belong to the numerator.
- Lower F critical value
- F(1 − α; a, b) = 1 ÷ F(α; b, a)
- Tables often give only upper points. Use this to get the lower point. Here F(α; a, b) is the upper α point.
- Paired t-test
- Di = Xi − Yi; T = (D̄ − δ0) ÷ (SD ÷ √n) ~ t(n − 1)
- n is the number of pairs. Differences must be roughly normal.
- Size and power
- α = P(reject H0 | H0 true); power = P(reject H0 | H1 true) = 1 − β
- β is the probability of a Type II error.
- Test statistic for a proportion (large n)
- z = (p̂ − p0) ÷ √(p0(1 − p0) ÷ n), where p̂ = x ÷ n
- Approximately N(0,1) under H0. Use p0 in the standard error. Use a continuity correction if you need to approximate the exact binomial tail.
- Exact binomial test
- Under H0, X ~ Bin(n, p0); p-value = P(X ≥ x) for H1: p > p0
- Use the lower tail for H1: p < p0. Because X is discrete, the exact size may be below the nominal α.
- Neyman-Pearson lemma
- Reject H0 if L(θ0; x) ÷ L(θ1; x) ≤ k, with k chosen so that the size equals α
- Applies to simple H0 against simple H1. For continuous data this gives the most powerful test of size α.
- Likelihood ratio statistic
- λ = L(θ̂0) ÷ L(θ̂); test statistic = −2 ln λ = 2[ln L(θ̂) − ln L(θ̂0)]
- θ̂0 is the MLE under H0 and θ̂ is the unrestricted MLE. Reject H0 for large −2 ln λ.
- Asymptotic distribution of the LRT
- −2 ln λ ≈ χ²(r), r = (free parameters in full model) − (free parameters under H0)
- Needs nested models and a large sample. Reject at level α if the statistic exceeds the upper α point of χ²(r).
- Binomial LRT for H0: p = p0
- −2 ln λ = 2[x ln(x ÷ (n p0)) + (n − x) ln((n − x) ÷ (n(1 − p0)))]
- Degrees of freedom = 1. Close to z² when n is large.
- Wald statistic (one parameter)
- W = (θ̂ − θ0)² ÷ Var̂(θ̂) ≈ χ²(1)
- Var̂ is usually from the inverse of the information at the MLE. Equivalent to a squared z-test.
- Expected frequency
- E = n × p
- n is the total number of observations. p is the probability of the cell under H0. Check that ΣE = ΣO = n.
- Test statistic
- X² = Σ (O − E)² ÷ E
- Sum over all cells after combining. Large values indicate a poor fit.
- Alternative form
- X² = Σ (O² ÷ E) − n
- Useful for quick calculation. It gives the same value if you do not round too early.
- Degrees of freedom
- ν = k − 1 − m
- k is the number of cells after combining. m is the number of parameters estimated from the data. Use m = 0 if the model is fully specified.
- Decision rule
- Reject H0 if X² > χ²(ν) upper critical value at level α
- Equivalently, reject if the p-value P(χ²(ν) > X²) is less than α.
- Cell size rule of thumb
- Combine adjacent cells until E ≥ 5
- A guideline for the chi-square approximation. Combine cells that are adjacent or logically similar, such as tail cells.
- Expected count
- E(i,j) = (row i total × column j total) ÷ n
- n is the grand total. Row and column totals of the E table must equal those of the observed table.
- Test statistic
- X² = Σ (O − E)² ÷ E
- Sum over all r × c cells. Equivalent form: Σ O² ÷ E − n.
- Degrees of freedom
- ν = (r − 1)(c − 1)
- For an r by c table when expected counts are estimated from the margins. Nothing else is estimated.
- Decision rule
- Reject H0 if X² > upper critical value of χ²(ν) at the chosen significance level
- Upper-tail test. Equivalent: reject if the p-value is below the significance level.
- Hypotheses
- H0: the two factors are independent; H1: they are not independent
- Write them in words relating to the context.
- Empirical distribution function
- F_n(x) = (number of sample values ≤ x) ÷ n
- A step function. It jumps by 1/n at each distinct observation.
- KS statistic
- D = max over x of |F_n(x) − F0(x)|
- Check at each ordered value x(i) both |i/n − F0(x(i))| and |(i−1)/n − F0(x(i))|. Take the largest of all.
- KS decision rule
- Reject H0 if D > critical value for n and the chosen significance level
- Use the tables you are given. For large n the 5% critical value is roughly 1.36 ÷ √n. Check the tables book before relying on this.
- Anderson-Darling statistic
- A² = −n − (1/n) Σ (2i − 1)[ln z_i + ln(1 − z_(n+1−i))], where z_i = F0(x(i))
- Ordered data x(1) ≤ … ≤ x(n). Larger A² means worse fit, with extra weight on the tails.
- Signs test
- Under H0, number of positive residuals ~ Binomial(n, ½)
- Drop zero residuals and reduce n. Use a normal approximation with mean n/2 and variance n/4 for large n.
- Runs test (mean)
- E(R) = 2·n1·n2 ÷ (n1 + n2) + 1
- n1 positive and n2 negative residuals. A run is an unbroken sequence of the same sign.
- Runs test (variance)
- Var(R) = 2·n1·n2·(2·n1·n2 − n1 − n2) ÷ [(n1 + n2)²·(n1 + n2 − 1)]
- For large n1 and n2, R is approximately normal. Use z = (R ± 0.5 − E(R)) ÷ √Var(R) with a continuity correction.
- Lag-1 serial correlation
- r1 = Σ (e_t − ē)(e_(t+1) − ē) ÷ Σ (e_t − ē)²
- Numerator sum runs over t = 1 to n−1. Under independence, r1 is approximately N(0, 1/n), so compare |r1| with 1.96 ÷ √n at 5%.
- Q-Q plot
- Plot ordered sample values x(i) against theoretical quantiles of the fitted distribution
- Points near a straight line mean good fit. For normality, plot against standard normal quantiles.
- Sign test statistic
- S = number of positive differences; under H0, S ~ Binomial(n, ½)
- Drop zero differences first and reduce n. For a test of a median m0, take differences x − m0. Use the binomial table or tail probabilities.
- Wilcoxon signed-rank statistic
- T+ = sum of ranks of positive differences; T− = sum of ranks of negative differences; T+ + T− = n(n + 1) ÷ 2
- Rank the absolute differences, giving tied values the average rank. Drop zeros. Tables usually use the smaller of T+ and T−, and you reject if it is at or below the critical value.
- Normal approximation for signed-rank
- E(T) = n(n + 1) ÷ 4; Var(T) = n(n + 1)(2n + 1) ÷ 24; Z = (T − E(T)) ÷ √Var(T)
- Use for larger n, with a continuity correction of 0.5. Ties need an adjusted variance.
- Mann-Whitney U
- U_X = R_X − m(m + 1) ÷ 2; U_X + U_Y = mn
- X has size m, Y has size n, and R_X is the sum of X's ranks in the combined ranking. Check with U_X + U_Y = mn. Some books use W = R_X (rank-sum form) instead of U; the tests are equivalent.
- Mann-Whitney normal approximation
- E(U) = mn ÷ 2; Var(U) = mn(m + n + 1) ÷ 12
- For the rank sum W = R_X: E(W) = m(m + n + 1) ÷ 2 and the variance is the same as for U.
- Spearman rank correlation
- r_s = 1 − 6Σd² ÷ (n(n² − 1))
- d is the difference between the two ranks of each item. This shortcut holds only when there are no ties. With ties, compute Pearson's correlation on the ranks.
- Kendall's tau
- τ = (C − D) ÷ (n(n − 1) ÷ 2)
- C is the number of concordant pairs and D the number of discordant pairs, with no ties. Values lie between −1 and +1. Under H0 of independence, Var(τ) = 2(2n + 5) ÷ (9n(n − 1)) for the normal approximation.
- Permutation test p-value
- p = (number of permutations with statistic at least as extreme as observed) ÷ (total number of permutations)
- For a two-sided test, count extreme values in both tails. If you sample random permutations, the p-value is only an estimate.
Quick revision
- H₀ is the default claim. You reject it or fail to reject it; you never prove it true.
- The p-value is the probability, assuming H₀ is true, of a result at least as extreme as observed.
- Reject H₀ when the p-value is less than or equal to the significance level.
- Type I error is rejecting a true H₀. Its probability is the significance level.
- Type II error is not rejecting a false H₀. Power = 1 − P(Type II error).
- Use a t test for a mean when σ is unknown and the data are normal. Use the normal when σ is known.
- The variance test uses (n − 1)S² ÷ σ₀², which is χ² with n − 1 degrees of freedom under normality.
- Chi-square statistic = Σ (O − E)² ÷ E. Merge cells with small expected counts, commonly below 5.
- Goodness of fit degrees of freedom = number of cells − 1 − number of estimated parameters.
- For an r × c contingency table, E = row total × column total ÷ grand total, with (r − 1)(c − 1) degrees of freedom.
- The Kolmogorov-Smirnov statistic is the largest gap between the empirical and hypothesised distribution functions.
- Permutation tests build the null distribution by reshuffling labels, so they need few distributional assumptions.
Common mistakes
- Saying 'accept H₀' when you fail to reject it. Fix: Write 'do not reject H₀' and add that there is insufficient evidence against it.
- Treating the p-value as the probability that H₀ is true. Fix: Define it as the probability, assuming H₀ is true, of a result at least as extreme as observed.
- Using a two-sample t-test on paired data. Fix: Ask whether each value in one sample belongs with one specific value in the other. If so, take differences and do a one-sample t-test with n − 1 degrees of freedom.
- Using n instead of n − 1 in the sample variance or the degrees of freedom. Fix: Sample variance S² divides by n − 1. A one-sample t or chi-square has n − 1 degrees of freedom. A pooled t has n + m − 2.
- Using the wrong direction in the Neyman-Pearson ratio, rejecting when L(θ0) ÷ L(θ1) is large. Fix: Reject H0 when the data are more likely under H1. So reject when L(θ0) ÷ L(θ1) is small, or L(θ1) ÷ L(θ0) is large. Sanity-check with the sign of θ1 − θ0.
- Using p̂ in the standard error for a hypothesis test on a proportion. Fix: For a test, the standard error is √(p0(1 − p0) ÷ n), because the distribution is computed under H0. Use p̂ in the standard error only for confidence intervals.
- Forgetting to subtract degrees of freedom for estimated parameters Fix: Always ask: did I estimate anything from this data? If yes, subtract one for each parameter. A Poisson with an estimated mean loses one more degree of freedom.
- Counting cells before combining Fix: Combine first. Then count k from the final table. Combining reduces ν.
- Using degrees of freedom r × c − 1 or r × c. Fix: For independence, estimating the margins removes r − 1 and c − 1 free values. Use (r − 1)(c − 1).
- Calculating expected counts with the wrong totals, such as dividing by a row total. Fix: Always use row total × column total ÷ grand total. Check that each E row and column adds to the observed totals.
Exam tips
- Always write H₀ and H₁ in terms of the parameter, with the test statistic and its null distribution. These are easy marks.
- Finish with a conclusion in context. Many students lose the last mark by stopping at 'reject H₀'.
- For power questions, standardise using the true mean from H₁, not the null mean.
- In multiple-choice questions, check whether the test is one-sided or two-sided before reading the table.
- In the computer-based paper, quote the p-value from R output and link it to α before concluding.
- Write H0, H1, the statistic's distribution and its degrees of freedom every time. Method marks are awarded for each, even if your arithmetic slips.
- State your assumptions explicitly: normality, independence, and equal variances for pooling. Examiners look for them.
- In computer-based questions, show the R call or Excel formula you used, such as t.test with paired = TRUE or var.test, and read off the statistic, degrees of freedom and p-value. Then write the conclusion in words.