IAI Actuarial Core Principles · Actuarial Statistics
Generalised Linear Models for IAI Actuarial Statistics
A **generalised linear model (GLM)** relates the mean of a response from an exponential family distribution to a linear predictor through a link function. You fit it by maximum likelihood, usually via IRLS, then judge it with deviance and residuals, and compare models using deviance differences or AIC.
What this chapter covers
This chapter extends ordinary linear regression. In a linear model the response is normal with constant variance and the mean equals the linear predictor. A GLM relaxes both conditions. The response can be Poisson, binomial, gamma or normal, and the mean is linked to the linear predictor by a function g, so that g(μ) = η = Xβ.
The chapter has a clear logic. First you learn the exponential family, because every GLM response belongs to it. Then you meet the three components: the distribution, the linear predictor and the link function. Next you see how parameters are estimated by maximum likelihood and iteratively reweighted least squares (IRLS). Finally you learn to check fit with deviance and residuals, choose between models, and apply GLMs to actuarial problems such as claim frequency and claim severity.
In the paper, this chapter builds directly on regression theory and on statistical inference, including likelihood, hypothesis tests and confidence intervals. In the 2026 syllabus, regression theory and applications carry the largest weighting in CS1, so GLMs sit in the heaviest part of the paper. It also links to CS2 through risk modelling distributions, and to pricing work in general insurance.
Regression theory and applications is the highest-weighted area in the 2026 CS1 syllabus, and GLMs are the most advanced part of it. Questions are often written, so they reward method: showing the exponential family form, identifying the canonical link, writing the linear predictor and interpreting coefficients. Much of this is procedural, so careful practice converts directly into marks. The chapter also supports the Paper B computer-based exam, where you fit and interpret GLMs in R.
Generalised linear models: topics in the order to study them
- 1Exponential Family of DistributionsEvery GLM response must fit this form, so the notation, mean and variance results come first.
- 2Components of a GLM: Link Functions and Linear PredictorOnce you know the family, you can see how the mean is tied to covariates through the link.
- 3Parameter Estimation: Maximum Likelihood and IRLSEstimation needs the family and link already in place, since the likelihood and weights depend on both.
- 4Deviance, Residuals and Goodness of FitYou can only measure fit after you know how a model is fitted and what the saturated model is.
- 5Model Selection and Hypothesis Testing in GLMsComparing models relies on deviance, so this follows the fit topic.
- 6Applications of GLMs in Actuarial WorkApplications pull everything together, so leave them until the theory is secure.
How to prepare Generalised linear models
Treat this chapter as a pipeline: family, link, fit, check, compare, apply. Study in that order and practise each stage on small examples.
- Learn the exponential family form and practise writing Poisson, binomial, normal and gamma in that form. Identify the natural parameter and the dispersion parameter each time.
- Derive the mean and variance from the cumulant function until you can do it without notes. Then state the variance function for each common distribution.
- List the canonical link for each common distribution. Practise writing the full model: distribution, link and linear predictor with named covariates.
- Work through one small likelihood and IRLS iteration by hand. Focus on what the weights and working response are, not on heavy arithmetic.
- Practise computing deviance and comparing nested models. Use the correct reference distribution for the test, and note how it differs when the dispersion is known or estimated.
- Fit GLMs in R on a practice data set. Read the output: coefficients, standard errors, null and residual deviance, and AIC. Then write a short interpretation in words.
- Finish with past-paper questions on claim frequency and severity, and time yourself on both the multiple-choice and written parts.
Common mistakes in Generalised linear models
Forgetting that the link function applies to the mean, not to the response itself.
Fix: Always write g(E(Y)) = Xβ. Say it aloud once: the link acts on the mean.
Mixing up the canonical link with the link you happen to choose.
Fix: Learn the canonical link for each distribution, but read the question. A log link for gamma is common and is not canonical.
Misreading coefficients under a log or logit link.
Fix: Under a log link, exponentiate the coefficient to get a multiplicative effect. Under a logit link, it gives an odds ratio.
Using the wrong reference distribution when comparing models by deviance.
Fix: Check whether the dispersion parameter is fixed, as for Poisson and binomial, or estimated. Choose chi-squared or F accordingly.
Comparing deviances or AIC across models fitted to different data or responses.
Fix: Only compare models fitted to the same data and the same response. Use deviance differences for nested models only.
Skipping assumptions and interpretation in written answers.
Fix: State the model, the distribution, the link and your assumptions. End by interpreting the result in words in terms of the problem.
Last-day revision: Generalised linear models
- A GLM has three parts: an exponential family response, a linear predictor η = Xβ, and a link function g with g(μ) = η.
- Exponential family form: f(y) = exp{[yθ − b(θ)] ÷ a(φ) + c(y, φ)}.
- Mean: E(Y) = b′(θ). Variance: Var(Y) = b″(θ) a(φ).
- The variance function V(μ) shows how variance depends on the mean, for example V(μ) = μ for Poisson.
- The canonical link sets g(μ) = θ. For Poisson it is log, for binomial it is logit, and for normal it is identity.
- The log link gives multiplicative effects: each coefficient exponentiated is a ratio change in the mean.
- Maximum likelihood estimates for GLMs usually have no closed form, so they are found by IRLS.
- Scaled deviance compares the fitted model with the saturated model, using twice the difference in log-likelihoods.
- Compare nested models by the difference in deviances. Use a chi-squared reference when the dispersion is known, and an F test when it is estimated.
- AIC = −2 × log-likelihood + 2 × (number of parameters). Lower is better, and it can compare non-nested models for the same data.
- Check residuals for patterns, since a clear pattern suggests a wrong link, a missing covariate or a wrong distribution.
- Frequency is commonly modelled with Poisson, and severity with gamma, often with a log link.
Generalised linear models practice questions
- A claim count N follows a Poisson distribution with mean μ. Written in exponential family form, what are the canonical parameter θ, the func…
- A GLM with 3 parameters is fitted to 20 observations and has deviance 30.0. Adding 2 further parameters reduces the deviance to 21.0. Assumi…
- For a Bernoulli response with success probability p, the pmf is written in exponential family form. Which choice gives the canonical paramet…
- A Poisson GLM with log link models claim counts with linear predictor η = β0 + β1·(age band B) + log(exposure). For a policy in band B with …
- Y is distributed as an exponential family member with b(θ) = −ln(−θ) for θ < 0 and φ = 1/ν for a constant ν > 0. Using E(Y) = b'(θ) and Var(…
- Two nested Poisson GLMs for claim counts are fitted to the same data. Model A (5 parameters) has deviance 312.4. Model B (8 parameters) adds…
- A Poisson GLM with log link is fitted to claim counts. The null model has deviance 180.4 on 99 degrees of freedom. After adding 3 rating fac…
- A Poisson GLM with log link has fitted coefficient 0.18 for a binary factor (1 = urban, 0 = rural), other covariates equal. By what percenta…
Generalised linear models: frequently asked questions
What is the difference between a linear model and a GLM?
A linear model assumes a normal response with constant variance and an identity link. A GLM allows any exponential family response and a chosen link, so variance can depend on the mean. The normal linear model is a special case of a GLM.
Do I need to learn IRLS in detail?
You should understand the idea well: each iteration solves a weighted least squares problem with updated weights and working response. Be ready to explain the steps and why they work. Heavy hand calculation is less likely than conceptual questions.
How do I choose the link function?
Start with the canonical link, which has convenient estimation properties. Then pick a link that suits the data: the log link keeps a mean positive and gives multiplicative effects. Check residuals to see whether the choice fits.
How is this chapter tested in Paper B?
Paper B is a 1 hour 45 minute computer-based exam, so expect to fit and interpret models in R. Practise reading the output, including coefficients, deviances and AIC, and writing short conclusions.
Which distributions should I know best for GLMs?
Know normal, Poisson, binomial and gamma thoroughly. For each one, learn the exponential family form, the mean and variance, the canonical link and a typical use in actuarial work.