Skip to content

Actuarial Statistics · Generalised linear models

Deviance, Residuals and Goodness of Fit in GLMs

Updated 11 October 2026 · Fact-checked

Scaled deviance measures how far a fitted GLM is from the saturated model: twice the difference in log-likelihoods. Pearson residuals are standardised raw errors; deviance residuals are signed square roots of each point's deviance contribution. To assess fit, compare the deviance with a chi-square distribution on n − p degrees of freedom and inspect the residuals.

Understand Deviance, Residuals and Goodness of Fit

A GLM is fitted by maximum likelihood. So the natural question is: how good is the fitted likelihood compared with the best possible one? The best possible model is the saturated model. It has one parameter per observation, so it fits every data point exactly (μ̂ᵢ = yᵢ). Its log-likelihood is the upper limit for any model on the same data.

The scaled deviance is D* = 2[l(saturated) − l(fitted)]. It is always ≥ 0. A small value means the fitted model is close to the saturated one. The deviance is D = φ × D*, where φ is the dispersion parameter. For the Poisson and binomial distributions φ = 1, so deviance and scaled deviance are equal. For the normal distribution, the deviance is the residual sum of squares and the scaled deviance is RSS ÷ σ².

The deviance is a sum of one term per observation. Taking the signed square root of each term gives the deviance residual. The Pearson residual is a different idea: the raw error (y − μ̂) divided by the standard deviation implied by the variance function. Both types should look roughly like random noise with no pattern if the model is adequate. Plot them against fitted values and against each covariate.

There are two uses of deviance. First, absolute fit: if the model is correct and the data are not sparse, the scaled deviance is approximately χ² with n − p degrees of freedom (p = number of fitted parameters). Second, comparing nested models: the drop in scaled deviance between a smaller and a larger model is approximately χ² with degrees of freedom equal to the number of extra parameters. The second use is more reliable. The first use is poor when counts are small or the data are individual 0/1 values.

Key rules to remember

Scaled deviance
D* = 2[l(saturated) − l(fitted)]
Always ≥ 0. The saturated model has μ̂ᵢ = yᵢ for every observation.
Deviance and dispersion
D = φ × D*
For Poisson and binomial, φ = 1, so D = D*. For the normal, D = Σ(yᵢ − μ̂ᵢ)².
Poisson deviance
D = 2 Σ [ yᵢ ln(yᵢ ÷ μ̂ᵢ) − (yᵢ − μ̂ᵢ) ]
If yᵢ = 0 the term is 2μ̂ᵢ. If the model has an intercept with the log link, Σ(yᵢ − μ̂ᵢ) = 0, so D = 2 Σ yᵢ ln(yᵢ ÷ μ̂ᵢ).
Binomial deviance (yᵢ successes out of nᵢ)
D = 2 Σ [ yᵢ ln(yᵢ ÷ μ̂ᵢ) + (nᵢ − yᵢ) ln((nᵢ − yᵢ) ÷ (nᵢ − μ̂ᵢ)) ]
Here μ̂ᵢ = nᵢ p̂ᵢ. A term with a zero count is taken as 0 × ln 0 = 0.
Gamma deviance
D = 2 Σ [ −ln(yᵢ ÷ μ̂ᵢ) + (yᵢ − μ̂ᵢ) ÷ μ̂ᵢ ]
Variance function V(μ) = μ².
Pearson residual
rᴾᵢ = (yᵢ − μ̂ᵢ) ÷ √V(μ̂ᵢ)
Poisson: (y − μ̂) ÷ √μ̂. Divide also by √φ if you want it on the scale of the dispersion.
Pearson statistic
X² = Σ (yᵢ − μ̂ᵢ)² ÷ V(μ̂ᵢ)
Sum of squared Pearson residuals. Usually close to D when the model fits well.
Deviance residual
rᴰᵢ = sign(yᵢ − μ̂ᵢ) × √dᵢ, where D = Σ dᵢ
Sum of squared deviance residuals equals the deviance.
Goodness of fit test
D* ≈ χ² with n − p degrees of freedom
Approximate. Unreliable for small counts or ungrouped binary data.
Nested model comparison
D*(smaller) − D*(larger) ≈ χ² with q degrees of freedom
q = number of extra parameters in the larger model. If φ is unknown, use an F test with the estimated φ.
Estimated dispersion
φ̂ = X² ÷ (n − p)
Pearson-based estimate. A value well above 1 for Poisson or binomial suggests overdispersion.

How to solve Deviance, Residuals and Goodness of Fit questions

Use this method for any question on deviance, residuals or goodness of fit in a GLM.

  1. 1Identify the distribution, the link function and the number of fitted parameters p. Note n, the number of observations (or groups).
  2. 2Find the fitted values μ̂ᵢ. If only the linear predictor is given, invert the link: for the log link, μ̂ = e^η.
  3. 3Write the unit deviance dᵢ for the distribution. Compute each term and add them to get D. Divide by φ if the scaled deviance is wanted and φ ≠ 1.
  4. 4For residuals, compute the Pearson residual (y − μ̂) ÷ √V(μ̂), or the deviance residual sign(y − μ̂) × √dᵢ. Keep the sign.
  5. 5To test absolute fit, compare D* with χ² on n − p degrees of freedom. A value far above the mean n − p points to poor fit or overdispersion.
  6. 6To compare nested models, subtract the scaled deviances, use q = difference in parameters as the degrees of freedom, and compare with the χ² critical value.
  7. 7State the conclusion in words, with the assumptions: approximate χ² result, sufficiently large counts, and a known or well-estimated φ. Mention residual plots as a further check.

Quickest way: Compare deviances first, then check one residual plot

When to use it: Use this under time pressure when you are given deviances for two nested models, or fitted and observed values for a handful of points.

  1. For nested models, subtract the two deviances at once. Do not recompute anything.
  2. Count extra parameters q. Use χ²(q) at 5%: 3.841 for q = 1, 5.991 for q = 2, 7.815 for q = 3.
  3. For q = 2 you can get the exact p-value as e^(−x ÷ 2), where x is the drop in scaled deviance.
  4. For absolute fit, compare D* with n − p. If D* is about equal to n − p, the fit is fine. If it is far above, suspect misfit or overdispersion.
  5. For a Poisson model with an intercept and log link, drop the −(y − μ̂) part and use D = 2 Σ y ln(y ÷ μ̂). Check Σ μ̂ = Σ y first.

Common mistakes in Deviance, Residuals and Goodness of Fit

  • Forgetting the −(y − μ̂) term in the Poisson deviance.

    Students remember the shortened form 2 Σ y ln(y ÷ μ̂) and apply it to every model.

    Fix: The short form needs Σ(y − μ̂) = 0, which holds when the model has an intercept with the log link. If the model has no intercept, use the full formula.

  • Using the wrong degrees of freedom for the deviance test.

    Students use n − 1 or n, or use the number of covariates instead of parameters.

    Fix: Absolute fit uses n − p, where p counts every fitted parameter including the intercept. Model comparison uses the difference in p between the two models.

  • Dropping the sign when computing deviance residuals.

    The square root of dᵢ is always positive, so the sign is easy to forget.

    Fix: Always multiply by sign(y − μ̂). The sign tells you whether the model over- or under-predicts that point.

  • Treating the χ² result for deviance as exact.

    The test is written as D* ~ χ² in notes and the word approximately is lost.

    Fix: Say it is approximate. For sparse counts or ungrouped binary data the absolute test is unreliable, but the comparison of nested models is still reasonable.

  • Ignoring the dispersion parameter when it is not 1.

    Poisson and binomial have φ = 1, and students carry this over to the normal or gamma.

    Fix: For normal, gamma and quasi-models, scaled deviance is D ÷ φ. Estimate φ (for example φ̂ = X² ÷ (n − p)) before using a χ² or F test.

  • Saying a model fits because the deviance is small, without looking at residuals.

    A single number feels conclusive.

    Fix: Also check residual plots for patterns, outliers and non-constant spread. A model can have an acceptable total deviance and still have a clear structural misfit.

Worked examples

Example 1

A Poisson GLM with log link and an intercept is fitted to four observations, using p = 2 parameters. The observed claim counts are 2, 5, 8, 5 and the fitted values are 3, 4, 6, 7. (a) Calculate the deviance. (b) Calculate the Pearson and deviance residuals for the fourth observation. (c) Comment on the fit.

Show the solution
  1. Check: Σ y = 20 and Σ μ̂ = 20, so Σ(y − μ̂) = 0. The short formula D = 2 Σ y ln(y ÷ μ̂) can be used, but we keep the full unit deviance dᵢ = 2[y ln(y ÷ μ̂) − (y − μ̂)] so each term is a valid residual input.
  2. Obs 1: y = 2, μ̂ = 3. 2 ln(2/3) = −0.81093. Minus (2 − 3) gives +1. Bracket = 0.18907. d₁ = 0.3781.
  3. Obs 2: y = 5, μ̂ = 4. 5 ln(1.25) = 1.11572. Minus (1) gives 0.11572. d₂ = 0.2314.
  4. Obs 3: y = 8, μ̂ = 6. 8 ln(4/3) = 2.30146. Minus (2) gives 0.30146. d₃ = 0.6029.
  5. Obs 4: y = 5, μ̂ = 7. 5 ln(5/7) = −1.68236. Minus (−2) gives 0.31764. d₄ = 0.6353.
  6. D = 0.3781 + 0.2314 + 0.6029 + 0.6353 = 1.848 (to 3 decimal places).
  7. (b) Pearson residual for obs 4 = (5 − 7) ÷ √7 = −2 ÷ 2.6458 = −0.756.
  8. Deviance residual for obs 4 = −√0.6353 = −0.797. Both are negative because the model over-predicts this point.
  9. (c) Degrees of freedom = n − p = 4 − 2 = 2. For Poisson φ = 1, so D* = 1.848. The 5% critical value of χ² on 2 degrees of freedom is 5.991. Since 1.848 < 5.991, there is no evidence of lack of fit. With only four small counts this χ² test is rough, so treat the conclusion with caution.

Answer: D ≈ 1.848. For observation 4 the Pearson residual is ≈ −0.756 and the deviance residual is ≈ −0.797. Since 1.848 is below 5.991 (χ² on 2 df at 5%), there is no evidence of lack of fit, subject to the approximation being rough for small counts.

Example 2

Two Poisson GLMs are fitted to n = 40 observations. Model A has 3 parameters and deviance 45.2. Model B contains all of Model A's terms plus 2 more, so it has 5 parameters, and has deviance 38.9. Test at the 5% level whether the extra terms are needed, and comment on the fit of Model B.

Show the solution
  1. Poisson has φ = 1, so scaled deviance equals deviance.
  2. Drop in scaled deviance = 45.2 − 38.9 = 6.3.
  3. Extra parameters q = 5 − 3 = 2. Under H₀ (the extra terms are not needed), the drop is approximately χ² with 2 degrees of freedom.
  4. The 5% critical value for χ² on 2 degrees of freedom is 5.991. Since 6.3 > 5.991, reject H₀.
  5. Exact p-value for 2 degrees of freedom is e^(−6.3 ÷ 2) = e^(−3.15) ≈ 0.043, which is below 0.05, confirming the result.
  6. Fit of Model B: degrees of freedom n − p = 40 − 5 = 35. A χ² variable on 35 degrees of freedom has mean 35 and standard deviation √70 ≈ 8.4. A deviance of 38.9 is close to the mean, so there is no sign of overdispersion or lack of fit.
  7. State the caveat: both tests are approximate and rely on counts being reasonably large. Residual plots should also be checked.

Answer: The drop in deviance is 6.3 on 2 degrees of freedom (p ≈ 0.043), which exceeds the 5% critical value 5.991. So the two extra terms are needed and Model B is preferred. Its deviance of 38.9 on 35 degrees of freedom shows no evidence of lack of fit, subject to the approximation.

Exam tips

  • Write the unit deviance formula before substituting numbers. Marks are given for the correct formula and for a clear table of terms.
  • Always state the degrees of freedom and where they come from (n − p, or the difference in parameters). This is a frequent lost mark.
  • If a question asks you to compare Pearson and deviance residuals, say that deviance residuals are usually closer to normal for non-normal data, and Pearson residuals are simpler to compute. Both sum-of-squares measures should be similar when the fit is good.
  • In computer-based (Paper B) work, read the deviance, null deviance and residual degrees of freedom from the model summary. Then say what they imply in a sentence.
  • Mention the approximation, the dispersion parameter and the residual plots in your conclusion. Examiners reward interpretation, not only the arithmetic.

Practice questions from Generalised linear models

Deviance, Residuals and Goodness of Fit: frequently asked questions

What is the difference between deviance and scaled deviance?

Scaled deviance is 2[l(saturated) − l(fitted)]. Deviance is the scaled deviance multiplied by the dispersion parameter φ. For Poisson and binomial models φ = 1, so the two are the same. For the normal model, the deviance is the residual sum of squares and the scaled deviance is RSS ÷ σ².

What is the difference between Pearson and deviance residuals?

A Pearson residual is the raw error divided by the standard deviation implied by the variance function. A deviance residual is the signed square root of an observation's contribution to the deviance. The squares of Pearson residuals sum to X², and the squares of deviance residuals sum to D. Deviance residuals are often closer to normal for skewed distributions.

How do I calculate deviance for a Poisson GLM?

Compute the fitted values μ̂ᵢ, then take D = 2 Σ [yᵢ ln(yᵢ ÷ μ̂ᵢ) − (yᵢ − μ̂ᵢ)]. If yᵢ = 0 the term is 2μ̂ᵢ. For a model with an intercept and the log link, the second part sums to zero, so the formula simplifies.

How do I know if a GLM fits well?

Compare the scaled deviance with χ² on n − p degrees of freedom. A value near n − p suggests a good fit, and a much larger value suggests poor fit or overdispersion. Also check residual plots, because the χ² approximation is poor for small counts.

Can I use the deviance to compare any two models?

No. The deviance difference test works for nested models fitted to the same data with the same distribution and link. For non-nested models you need other tools, such as AIC.