Skip to content

Actuarial Statistics · Generalised linear models

Components of a GLM: Link Functions and the Linear Predictor

Updated 11 October 2026 · Fact-checked

A generalised linear model has three parts: a random component (the response follows an exponential family distribution), a systematic component (the linear predictor η = Xβ) and a link function g that connects them through g(μ) = η. You pick the distribution from the data type, then choose a suitable link.

Understand Components of a GLM: Link Functions and Linear Predictor

A linear model assumes the response is normal with constant variance and that its mean is a linear function of the covariates. This fails for claim counts (non-negative integers), claim amounts (positive and skewed) and yes/no outcomes. A generalised linear model (GLM) relaxes both assumptions.

A GLM has three parts.

  • Random component: the responses Y₁,…,Yₙ are independent and each has a distribution from the exponential family, such as normal, Poisson, binomial, gamma or inverse Gaussian. Its mean is μᵢ = E(Yᵢ).
  • Systematic component: the linear predictor ηᵢ = β₀ + β₁xᵢ₁ + … + βₚxᵢₚ. It is linear in the parameters β. The covariates themselves may be transformed, for example x² or log x, and factors enter through dummy variables.
  • Link function: a monotonic, differentiable function g with g(μᵢ) = ηᵢ. So μᵢ = g⁻¹(ηᵢ).

The link lets the mean sit in a sensible range. With a log link, μ = exp(η) is always positive, so a Poisson or gamma mean can never go negative. With a logit link, μ = e^η ÷ (1 + e^η) always lies between 0 and 1, which suits a probability.

The canonical link is the link for which g(μ) equals the canonical parameter θ of the exponential family form. It gives the neat result that η = θ. It also simplifies the likelihood equations and makes the sufficient statistics easy to use. The canonical links are: identity for normal, log for Poisson, logit for binomial, and reciprocal (1/μ) for gamma. For inverse Gaussian it is 1/μ².

The canonical link is a convenient default, not a rule. In practice the log link is very common for gamma and Poisson models because it makes effects multiplicative and keeps the mean positive. Choose the link using the range of the mean, how you want effects to behave (additive or multiplicative) and how well the model fits.

Key rules to remember

Link and linear predictor
g(μᵢ) = ηᵢ = β₀ + β₁xᵢ₁ + … + βₚxᵢₚ, so μᵢ = g⁻¹(ηᵢ)
Linear in the parameters β. The link is monotonic and differentiable.
Exponential family form
f(y; θ, φ) = exp{ [yθ − b(θ)] ÷ a(φ) + c(y, φ) }
θ is the canonical parameter. Usually a(φ) = φ ÷ w for a prior weight w.
Mean and variance
E(Y) = μ = b′(θ); Var(Y) = a(φ) b″(θ) = a(φ) V(μ)
V(μ) is the variance function. It depends on the distribution, not on the link.
Canonical link
g(μ) = θ, so η = θ
Equivalent to g = (b′)⁻¹.
Canonical links
Normal: g(μ) = μ; Poisson: g(μ) = ln μ; Binomial: g(μ) = ln[μ ÷ (1 − μ)]; Gamma: g(μ) = 1 ÷ μ
For the binomial, μ is the probability p of success.
Variance functions
Normal: V(μ) = 1; Poisson: V(μ) = μ; Binomial: V(μ) = μ(1 − μ); Gamma: V(μ) = μ²
For the gamma, Var(Y) = φμ², where φ = 1 ÷ α with shape α.
Log link interpretation
μ = exp(β₀) × exp(β₁x₁) × …
A one-unit rise in x₁ multiplies the mean by exp(β₁).

How to solve Components of a GLM: Link Functions and Linear Predictor questions

Use this method for any question that asks you to specify, interpret or justify a GLM.

  1. 1Identify the response type: count, positive continuous, proportion or binary, or unrestricted real. This decides the distribution.
  2. 2Pick the distribution from the exponential family: Poisson for counts, gamma or inverse Gaussian for positive skewed amounts, binomial for yes/no or proportions, normal for unrestricted real values.
  3. 3Write the random component: state independence, the distribution, μᵢ = E(Yᵢ) and the variance function.
  4. 4Write the linear predictor ηᵢ from the covariates, and define any factors with dummy variables and a baseline level.
  5. 5Choose the link g. State the canonical link, then say whether you keep it or use another. Check that g⁻¹(η) keeps μ in the allowed range.
  6. 6Write the model as g(μᵢ) = ηᵢ and invert it to give μᵢ.
  7. 7Interpret the parameters: additive effects on g(μ). For a log link, multiplicative effects exp(β) on the mean.
  8. 8If asked, justify the choice using the range of the mean, interpretation and fit.

Quickest way: Response type to model and link table

When to use it: Use this when the question asks you to choose or name a model and link quickly, especially in multiple-choice questions.

  1. Counts, or claim numbers per exposure: Poisson with log link. Use ln(exposure) as an offset.
  2. Positive skewed claim amounts: gamma with log link (canonical is reciprocal).
  3. Binary or proportion data: binomial with logit link.
  4. Unrestricted real data: normal with identity link.
  5. Write the canonical link as g(μ) = θ and quickly check that the link keeps μ in range.
  6. For interpretation, say: log link means multiplicative, identity link means additive.

Common mistakes in Components of a GLM: Link Functions and Linear Predictor

  • Saying the linear predictor is linear in the covariates only.

    The name sounds like it describes the x variables.

    Fix: Say it is linear in the parameters β. Covariates can be transformed, such as x² or ln x.

  • Modelling the mean of the transformed response, E[g(Y)], instead of g(E[Y]).

    It is confused with fitting a linear model to log Y.

    Fix: A GLM links the mean: g(E[Y]) = η. The response is not transformed, and the variance follows from the chosen distribution.

  • Treating the canonical link as compulsory.

    Notes stress that it has useful properties.

    Fix: Say it is a convenient default. Non-canonical links, such as log for gamma, are often better for interpretation and for keeping μ positive.

  • Giving the wrong canonical link, for example log for the binomial or logit for the Poisson.

    Links are memorised without linking them to θ.

    Fix: Derive θ from the density. Poisson: θ = ln μ. Binomial: θ = ln[p ÷ (1 − p)]. Gamma: θ is proportional to −1 ÷ μ, so the link is 1 ÷ μ.

  • Interpreting a log-link coefficient as an additive change in the mean.

    It is carried over from the linear model.

    Fix: A one-unit increase in x multiplies the mean by exp(β). For example, β = 0.2 gives a factor exp(0.2) ≈ 1.221.

  • Believing the link determines the variance.

    Link and variance function are mixed up.

    Fix: The variance function V(μ) comes from the distribution. The link only relates the mean to η.

Worked examples

Example 1

Claim counts Yᵢ for motor policies are modelled as Poisson with mean μᵢ. The covariate is the driver's age band (young = 1, otherwise 0), and a log link is used. The fitted model is ln μ = −2.1 + 0.4x. (a) Write the three components. (b) Find the fitted mean for young and other drivers. (c) By what factor is the mean larger for young drivers?

Show the solution
  1. (a) Random component: Yᵢ are independent Poisson(μᵢ), with Var(Yᵢ) = μᵢ. Systematic component: ηᵢ = β₀ + β₁xᵢ. Link: g(μ) = ln μ, which is the canonical link for the Poisson.
  2. (b) For x = 0: μ = exp(−2.1) = 0.1225.
  3. For x = 1: η = −2.1 + 0.4 = −1.7, so μ = exp(−1.7) = 0.1827.
  4. (c) The ratio is exp(0.4) = 1.4918, which agrees with 0.1827 ÷ 0.1225 ≈ 1.49.

Answer: Poisson with ηᵢ = β₀ + β₁xᵢ and a log link. Fitted means are 0.1225 (other drivers) and 0.1827 (young drivers). Young drivers' mean is about 1.49 times larger.

Example 2

Show that the canonical link for the Poisson distribution is the log link, and state the corresponding inverse link.

Show the solution
  1. The Poisson probability function is f(y) = e^(−μ) μ^y ÷ y!.
  2. Take logs: ln f = y ln μ − μ − ln y!.
  3. Compare with the exponential family form yθ − b(θ) + c(y), with a(φ) = 1. So θ = ln μ and b(θ) = μ = e^θ.
  4. Check: b′(θ) = e^θ = μ, and b″(θ) = e^θ = μ, so V(μ) = μ, as expected.
  5. The canonical link satisfies g(μ) = θ, so g(μ) = ln μ.
  6. The inverse link is μ = exp(η), which is positive for every η.

Answer: The canonical link is g(μ) = ln μ, with inverse μ = exp(η). The mean is therefore always positive.

Exam tips

  • Always name all three components: random, systematic and link. Marks are usually given for each.
  • If you are asked for the canonical link, derive θ from the density in the exponential family form. Showing this earns method marks.
  • For a link choice, give a reason: the allowed range of μ, how effects are interpreted, or fit.
  • Interpret coefficients on the right scale. A log link gives multiplicative effects, exp(β), and a logit link gives effects on the log odds.
  • In the computer-based paper, state the family and link explicitly in the model call, for example family = Gamma(link = "log"), and check the default link.

Practice questions from Generalised linear models

Components of a GLM: Link Functions and Linear Predictor: frequently asked questions

What is the difference between a linear model and a generalised linear model?

A linear model assumes a normal response with constant variance and a mean that is linear in the covariates. A GLM allows any exponential family distribution and connects the mean to the linear predictor through a link function. The variance can then depend on the mean.

What is a canonical link function in a GLM?

It is the link for which g(μ) equals the canonical parameter θ of the exponential family form of the distribution. It makes the linear predictor equal to θ and simplifies estimation. Examples are identity for normal, log for Poisson and logit for binomial.

How do I choose a link function for Poisson and gamma GLMs?

For both, the log link is the usual choice because it keeps the mean positive and gives multiplicative effects. It is canonical for the Poisson. For the gamma the canonical link is the reciprocal, but the log link is often preferred in practice.

What is the linear predictor?

It is η = β₀ + β₁x₁ + … + βₚxₚ, a linear combination of the parameters and covariates. It is the quantity that the link function sets equal to g(μ). The fitted mean is found by applying the inverse link to η.