Skip to content

Actuarial Statistics · Generalised linear models

Model Selection and Hypothesis Testing in GLMs

Updated 11 October 2026 · Fact-checked

To compare nested GLMs, fit both, take the difference in their deviances and compare it with a chi-squared distribution whose degrees of freedom equal the number of extra parameters. If the difference is large, keep the bigger model. For non-nested models, compare AIC and choose the smaller. Use an F-test when the scale parameter is estimated.

Understand Model Selection and Hypothesis Testing in GLMs

A GLM is fitted by maximum likelihood. A bigger model with more parameters always fits the data at least as well as a smaller model nested inside it. So a better fit alone does not justify extra parameters. You need a way to decide if the improvement is real or just noise.

The likelihood ratio test does this for nested models. Model 1 is nested in Model 2 if you can get Model 1 by setting some parameters of Model 2 to fixed values (usually zero). The null hypothesis is that the smaller model is adequate. The test statistic is 2 × (log-likelihood of bigger model − log-likelihood of smaller model). For large samples it is approximately chi-squared with degrees of freedom equal to the number of extra parameters.

In GLM work you usually see this as a difference in deviances. The scaled deviance is 2 × (log-likelihood of the saturated model − log-likelihood of the fitted model). The saturated terms cancel when you subtract, so the deviance difference equals the likelihood ratio statistic. For the Poisson and binomial, the scale parameter φ is known (φ = 1), so you compare the difference directly with a chi-squared value. For the normal, gamma and similar families, φ is usually unknown and must be estimated. Then you use an F-test, which divides by the estimated scale.

AIC handles models that are not nested, and also gives a way to rank many models. AIC = −2 × log-likelihood + 2 × (number of parameters). It rewards fit and penalises complexity. You choose the model with the smallest AIC. There is no p-value, and only the differences between AIC values matter, not the size of any one value.

Variable selection methods such as forward selection, backward elimination and stepwise selection repeat these comparisons. Forward selection adds the variable that improves the model most. Backward elimination removes the least useful one. Stepwise allows both. They are practical but not perfect, so check the final model with residuals and common sense.

Key rules to remember

Likelihood ratio statistic
LR = 2 × (ℓ₂ − ℓ₁) ~ χ²(p₂ − p₁) approximately
Model 1 (p₁ parameters) must be nested in Model 2 (p₂ parameters). ℓ is the maximised log-likelihood. Large-sample result.
Scaled deviance
D* = 2 × (ℓ_saturated − ℓ_model)
Use the scaled deviance for the chi-squared test when φ is known.
Deviance difference test (φ known)
D*₁ − D*₂ ~ χ²(p₂ − p₁)
Reject the smaller model if the difference exceeds the chi-squared critical value. For Poisson and binomial, scaled deviance equals deviance.
F-test (φ unknown)
F = [(D₁ − D₂) ÷ (p₂ − p₁)] ÷ [D₂ ÷ (n − p₂)] ~ F(p₂ − p₁, n − p₂)
D is the unscaled deviance. The denominator estimates the scale parameter. Use for normal, gamma and similar families.
Akaike Information Criterion
AIC = −2ℓ + 2p
Choose the model with the lowest AIC. p counts all estimated parameters.
Degrees of freedom of a model
Residual df = n − p
n is the number of observations (or cells), p the number of fitted parameters.

How to solve Model Selection and Hypothesis Testing in GLMs questions

Use this method for any question that asks you to compare GLMs or test whether a variable should be kept.

  1. 1Write down the two models clearly and check that the smaller one is nested in the larger one. If not, plan to use AIC instead.
  2. 2State the hypotheses: H₀ is that the smaller model is adequate (the extra parameters are zero). H₁ is that the larger model is needed.
  3. 3Find the deviances (or log-likelihoods) and the number of parameters or residual degrees of freedom for each model.
  4. 4Compute the deviance difference. The degrees of freedom are the number of extra parameters, which equals the difference in residual degrees of freedom.
  5. 5Choose the reference distribution. Use χ² if the scale parameter is known (Poisson, binomial). Use F if it is estimated (normal, gamma).
  6. 6Compare with the critical value or p-value. A large statistic means reject H₀ and keep the larger model.
  7. 7State the conclusion in words, in terms of the variable or factor being tested.
  8. 8If asked to rank several models, compute AIC = −2ℓ + 2p for each and pick the lowest.

Quickest way: Deviance difference in three lines

When to use it: Use when a table of residual deviances and degrees of freedom is given and you must test one added term.

  1. Subtract the deviances (smaller model minus larger model). Subtract the residual degrees of freedom the same way.
  2. Compare the deviance difference with the χ² critical value on that many degrees of freedom. Rule of thumb: the 5% critical value of χ²(1) is 3.84, so a drop above about 3.84 for one extra parameter is significant.
  3. For AIC, just compute −2ℓ + 2p for each model and pick the lowest. No table needed.

Common mistakes in Model Selection and Hypothesis Testing in GLMs

  • Using a chi-squared test on models that are not nested.

    Students see two deviances and subtract without checking the structure.

    Fix: Always check that one model is a special case of the other. If not, use AIC.

  • Using the wrong degrees of freedom, such as the residual df of one model instead of the difference.

    The table lists several df values and it is easy to pick the wrong one.

    Fix: Degrees of freedom for the test equal the number of extra parameters, which is the difference in residual df.

  • Choosing the model with the highest AIC, or forgetting the 2p penalty.

    Students link a bigger number with better fit.

    Fix: AIC = −2ℓ + 2p and lower is better. Write the formula out each time.

  • Using a chi-squared test when the scale parameter is unknown.

    The chi-squared test is the one they practise most.

    Fix: For normal or gamma responses, estimate φ and use the F-test.

  • Testing the wrong direction: deviance of the larger model minus the smaller.

    Mixing up which model is which.

    Fix: The smaller model always has the larger deviance, so subtract larger model from smaller to get a positive value.

  • Treating a significant test as proof of the best model.

    Students stop once p < 0.05.

    Fix: Also check residuals, sensible interpretation and whether the variable makes practical sense.

Worked examples

Example 1

A Poisson GLM with log link is fitted to claim counts. Model A (age band only, 4 parameters) has deviance 62.4 on 36 residual df. Model B (age band plus vehicle type, 6 parameters) has deviance 55.1 on 34 residual df. Test at the 5% level whether vehicle type should be included. The 5% critical value of χ²(2) is 5.991.

Show the solution
  1. Model A is nested in Model B, since A is B with the vehicle parameters set to zero.
  2. H₀: the vehicle type parameters are zero. H₁: they are not.
  3. The Poisson scale parameter is known, so use χ².
  4. Deviance difference = 62.4 − 55.1 = 7.3.
  5. Degrees of freedom = 6 − 4 = 2, which matches 36 − 34.
  6. Compare: 7.3 > 5.991.

Answer: Reject H₀. Vehicle type significantly improves the model at the 5% level, so include it.

Example 2

Three GLMs are fitted to the same data. Model 1 has 3 parameters and maximised log-likelihood −210.5. Model 2 has 5 parameters and log-likelihood −205.0. Model 3 has 8 parameters and log-likelihood −203.8. Which model does AIC select?

Show the solution
  1. AIC = −2ℓ + 2p.
  2. Model 1: 421.0 + 6 = 427.0.
  3. Model 2: 410.0 + 10 = 420.0.
  4. Model 3: 407.6 + 16 = 423.6.
  5. The lowest AIC is Model 2.

Answer: AIC selects Model 2 (AIC = 420.0). Model 3 fits slightly better but the extra three parameters are not worth the penalty.

Exam tips

  • Always say whether the models are nested before choosing a test. Examiners award a mark for this.
  • Write the hypotheses and the degrees of freedom explicitly, even when the arithmetic is short.
  • Know which families have known scale (Poisson, binomial) and which do not (normal, gamma). This decides χ² versus F.
  • In computer-based questions, show the R output you use, such as the anova table or AIC values, then state your conclusion in words.
  • Give a conclusion in context, naming the rating factor, rather than only 'reject H₀'.

Practice questions from Generalised linear models

Model Selection and Hypothesis Testing in GLMs: frequently asked questions

What is the difference between deviance and the likelihood ratio statistic?

Deviance compares a fitted model with the saturated model. The likelihood ratio statistic compares two fitted models. The difference between two deviances equals the likelihood ratio statistic for those two models, because the saturated log-likelihood cancels.

When should I use AIC instead of a chi-squared test?

Use AIC when models are not nested, or when you are ranking several models at once. Use the chi-squared (or F) test when you have two nested models and want a formal hypothesis test.

Why does the F-test replace the chi-squared test for some GLMs?

When the scale parameter is unknown, it has to be estimated from the data. The F-test builds this estimate into the denominator, which gives a more accurate reference distribution in small samples.

Is stepwise selection reliable?

It is a useful tool but not a guarantee of the best model. It depends on the order of steps and the criterion used. Always check the chosen model with residuals and judgement.