Actuarial Statistics · Generalised linear models
GLM Parameter Estimation: Maximum Likelihood and IRLS Explained
Updated 11 October 2026 · Fact-checked
GLM parameters are estimated by maximum likelihood. You write the log-likelihood for the exponential family, differentiate to get score equations, and solve them numerically with iteratively reweighted least squares (IRLS), which repeats a weighted regression on a working response. The scale parameter is then estimated from Pearson residuals or the deviance.
Understand Parameter Estimation: Maximum Likelihood and IRLS
A GLM assumes each response Yᵢ comes from an exponential family with mean μᵢ. The mean links to the linear predictor ηᵢ = xᵢᵀβ through a link function g, so g(μᵢ) = ηᵢ. Your job is to find the β that makes the observed data most likely.
For the exponential family in canonical form, the log-likelihood of one observation is ℓᵢ = (yᵢθᵢ − b(θᵢ)) ÷ a(φ) + c(yᵢ, φ). Here E(Yᵢ) = b′(θᵢ) and Var(Yᵢ) = a(φ) b″(θᵢ). Write a(φ) = φ ÷ wᵢ, where wᵢ is a known prior weight. The variance function is V(μ) = b″(θ), so Var(Yᵢ) = φ V(μᵢ) ÷ wᵢ.
Differentiating the log-likelihood with respect to each βⱼ gives the score equations. By the chain rule, they simplify to Σᵢ wᵢ (yᵢ − μᵢ) xᵢⱼ ÷ (V(μᵢ) g′(μᵢ)) = 0 for each j. The scale parameter φ cancels, so you can find β̂ without knowing φ. With the canonical link, the equations reduce to Σᵢ wᵢ (yᵢ − μᵢ) xᵢⱼ = 0. For Poisson with log link and no weights, this means the fitted totals match the observed totals.
These equations are non-linear in β for most links, so there is no closed form. IRLS solves them using Fisher scoring. At each step you build a working response zᵢ = ηᵢ + (yᵢ − μᵢ) g′(μᵢ) and working weights Wᵢ = wᵢ ÷ (V(μᵢ) g′(μᵢ)²). You then do a weighted least squares regression of z on X with weights W. Update, recompute μ, and repeat until β converges.
Once β̂ is found, you estimate φ if it is unknown. Poisson and binomial have φ = 1 by assumption. Normal, gamma and inverse Gaussian need φ estimated. The usual estimate is Pearson's: φ̂ = (1 ÷ (n − p)) Σ wᵢ (yᵢ − μ̂ᵢ)² ÷ V(μ̂ᵢ).
Key rules to remember
- Exponential family log-likelihood
- ℓ = Σ [ (yᵢθᵢ − b(θᵢ)) ÷ a(φ) + c(yᵢ, φ) ]
- With a(φ) = φ ÷ wᵢ. Mean = b′(θ), variance = a(φ) b″(θ).
- Score equations
- Σᵢ wᵢ (yᵢ − μᵢ) xᵢⱼ ÷ (V(μᵢ) g′(μᵢ)) = 0, for j = 1,…,p
- φ cancels. Canonical link: Σᵢ wᵢ (yᵢ − μᵢ) xᵢⱼ = 0.
- Working response
- zᵢ = ηᵢ + (yᵢ − μᵢ) g′(μᵢ)
- Computed using current estimates of ηᵢ and μᵢ.
- Working weights
- Wᵢ = wᵢ ÷ ( V(μᵢ) [g′(μᵢ)]² )
- Recomputed at each iteration.
- IRLS update
- β_new = (XᵀWX)⁻¹ XᵀWz
- W is the diagonal matrix of working weights.
- Approximate variance of β̂
- Var(β̂) ≈ φ (XᵀWX)⁻¹
- W evaluated at β̂. Asymptotic result.
- Pearson estimate of scale
- φ̂ = (1 ÷ (n − p)) Σ wᵢ (yᵢ − μ̂ᵢ)² ÷ V(μ̂ᵢ)
- p is the number of parameters in the linear predictor.
- Deviance-based estimate of scale
- φ̂ = D ÷ (n − p)
- D is the scaled-free deviance, defined as φ times the scaled deviance. Often less reliable than Pearson's.
How to solve Parameter Estimation: Maximum Likelihood and IRLS questions
Use this order for any question on likelihood, score equations, IRLS or the scale parameter.
- 1Identify the distribution, its V(μ), the link g and any prior weights wᵢ. Write down whether φ is known.
- 2Write the log-likelihood in terms of μᵢ, then replace μᵢ by g⁻¹(xᵢᵀβ).
- 3Differentiate with respect to each βⱼ. Use the chain rule through θ, μ and η. State the score equations, noting that φ cancels.
- 4If the link is canonical, simplify to Σ wᵢ (yᵢ − μᵢ) xᵢⱼ = 0. Check whether the equations can be solved directly, as with a single categorical factor.
- 5If not, set up IRLS. Compute zᵢ and Wᵢ from the starting values, then solve (XᵀWX)β = XᵀWz.
- 6Update ηᵢ and μᵢ, and repeat until the change in β or the deviance is small.
- 7Estimate φ if needed using the Pearson formula with n − p in the divisor. State the assumption for Poisson or binomial that φ = 1.
- 8Give the final answer with units and state any approximation, such as the asymptotic variance.
Quickest way: Shortcut: canonical link with a single factor
When to use it: Use when the model has one categorical factor, or only an intercept, and a canonical link. The score equations then give the fitted means directly.
- Write the score equation Σ wᵢ (yᵢ − μᵢ) xᵢⱼ = 0 for each parameter.
- For each level, the equation says the sum of fitted values equals the sum of observed values for that level.
- So μ̂ for a level is the weighted average of observations in that level. Convert to the parameter through the link.
- For an intercept-only model, μ̂ = Σ wᵢyᵢ ÷ Σ wᵢ.
- Skip IRLS and go straight to the scale estimate if needed.
Common mistakes in Parameter Estimation: Maximum Likelihood and IRLS
Including φ in the score equations or thinking β̂ depends on it.
φ appears in the log-likelihood, so students keep it after differentiating.
Fix: Note that a(φ) is a common factor in the derivative. It cancels, so β̂ is found without φ. φ matters only for standard errors and testing.
Using n instead of n − p in the Pearson estimate of φ.
Students copy the sample variance idea and forget parameters were fitted.
Fix: Always divide by the degrees of freedom n − p, where p counts every fitted parameter including the intercept.
Forgetting g′(μ) squared in the working weight, or using g′ instead of its square.
The two working quantities look similar, and the powers are easy to mix up.
Fix: Remember z uses g′ once, W uses g′ squared. Check with the log link, where g′ = 1/μ, so W = wμ² ÷ V(μ).
Applying the simplified canonical-link score equation to a non-canonical link.
The short form is used so often that students forget its condition.
Fix: Check that g is the canonical link for the distribution. If not, keep the full form with V(μ) g′(μ) in the denominator.
Estimating φ for a Poisson or binomial model without being told to.
Students apply the Pearson formula to every GLM.
Fix: For Poisson and binomial, φ = 1 by definition. Only estimate it if the question allows for over-dispersion.
Using fitted μ̂ᵢ from the wrong iteration when computing z and W.
Students mix old and new values during hand calculations.
Fix: Finish each iteration by updating η, then μ = g⁻¹(η), and only then compute the next z and W.
Worked examples
Example 1
Claim counts Yᵢ are Poisson with mean μᵢ and a log link, with a single intercept: log μᵢ = β₀. Observed counts for five policies, each with exposure 1, are 2, 0, 3, 1, 4. Find β̂₀ and then Var(β̂₀) approximately.
Show the solution
- Poisson with log link is canonical. The score equation is Σ (yᵢ − μ) = 0 with μ = e^β₀.
- Σ yᵢ = 2 + 0 + 3 + 1 + 4 = 10, so 10 − 5μ = 0 and μ̂ = 2.
- β̂₀ = ln 2 ≈ 0.6931.
- For the variance, V(μ) = μ and g′(μ) = 1/μ, so W = μ²÷μ = μ = 2 for each policy. XᵀWX = 5 × 2 = 10.
- With φ = 1, Var(β̂₀) ≈ 1 ÷ 10 = 0.1.
Answer: β̂₀ = ln 2 ≈ 0.693 and Var(β̂₀) ≈ 0.1.
Example 2
A gamma GLM with a log link and an intercept-only model is fitted to n = 4 claim sizes (all prior weights 1): 100, 200, 300, 400. The fitted mean is μ̂ = 250. Variance function is V(μ) = μ². Estimate the scale parameter φ using the Pearson statistic.
Show the solution
- For an intercept-only model p = 1, so n − p = 3.
- Residuals yᵢ − μ̂: −150, −50, 50, 150.
- Squared residuals: 22,500, 2,500, 2,500, 22,500. Sum = 50,000.
- Divide each by V(μ̂) = 250² = 62,500. The sum becomes 50,000 ÷ 62,500 = 0.8.
- φ̂ = 0.8 ÷ 3 ≈ 0.2667.
Answer: φ̂ ≈ 0.267.
Exam tips
- Start every written answer by naming V(μ), the link and whether it is canonical. Marks are often given for this alone.
- Show the chain rule clearly when deriving the score equations. Examiners award marks for each link in the chain.
- For IRLS questions, set out z and W in a small table by observation so errors are easy to follow.
- State that φ cancels in the score equations, and that Poisson and binomial have φ = 1.
- In computer-based papers, you can fit with glm() in R. Be ready to say what IRLS is doing, and read the dispersion from the summary output.
Practice questions from Generalised linear models
- For a Bernoulli response with success probability p, the pmf is written in exponential family form. Which choice gives the canonical paramet…
- A Poisson log-link GLM for claim frequency gives a coefficient of 0.30 for 'urban' relative to the base level 'rural'. Holding other factors…
- Y is distributed as an exponential family member with b(θ) = −ln(−θ) for θ < 0 and φ = 1/ν for a constant ν > 0. Using E(Y) = b'(θ) and Var(…
- Two nested Poisson GLMs for claim counts are fitted to the same data. Model A (5 parameters) has deviance 312.4. Model B (8 parameters) adds…
- A Poisson GLM with log link is fitted to claim counts. The null model has deviance 180.4 on 99 degrees of freedom. After adding 3 rating fac…
Parameter Estimation: Maximum Likelihood and IRLS: frequently asked questions
Why do GLMs use IRLS instead of solving the score equations directly?
The score equations are non-linear in β for most links, so no closed-form solution exists. IRLS turns each step into a weighted least squares problem that can be solved with matrix algebra. It repeats until the estimates converge.
How do you estimate the dispersion parameter in a GLM?
The usual method is the Pearson estimate: the sum of weighted squared residuals divided by V(μ̂), then divided by n − p. A deviance-based estimate, D ÷ (n − p), is also used. For Poisson and binomial models, φ is taken as 1 unless over-dispersion is being modelled.
Does the scale parameter affect the estimates of β?
No. φ cancels from the score equations, so β̂ is the same whatever value φ takes. It does affect the standard errors of β̂ and any test statistics based on them.
Is IRLS the same as Newton-Raphson?
Not in general. IRLS is Fisher scoring, which uses the expected information. With a canonical link, the expected and observed information are equal, so it matches Newton-Raphson.