Skip to content

Actuarial Statistics · Linear regression models

Residuals and Model Diagnostics in Linear Regression

Updated 11 October 2026 · Fact-checked

Residual analysis checks whether a fitted linear regression meets its assumptions. A residual is the observed value minus the fitted value. Plot residuals against fitted values to check linearity and constant variance, use a normal QQ plot to check normality, and use leverage and Cook's distance to find influential points. Fix problems by transforming, adding terms or reviewing points.

Understand Residuals and Model Diagnostics

A linear regression model assumes that the errors are independent, have mean zero, have constant variance σ², and (for exact inference) are normally distributed. You never see the errors. You see the residuals, eᵢ = yᵢ − ŷᵢ, which estimate them. If the model is right, the residuals should look like random noise with no pattern.

So diagnostics means looking for patterns. A residuals vs fitted values plot is the main tool. A random horizontal band around zero is good. A curve means the mean is not linear in the predictors, so add a term or transform. A funnel shape means the variance changes with the mean (heteroscedasticity). A plot of residuals against time or observation order that shows runs of positive then negative values suggests autocorrelation, meaning the errors are not independent.

To check normality, use a normal QQ plot of standardised residuals. Points close to the straight line support normality. An S-shape suggests heavy or light tails. A bend at one end suggests skewness. You can also use a formal test, but with small samples the plot is more informative.

Not all points have equal pull on the fitted line. Leverage hᵢᵢ measures how unusual a point's x-values are. An outlier is a point with an unusually large residual. A point with high leverage and a large residual can be influential: removing it changes the fit a lot. Cook's distance combines both ideas into one number.

A useful point for exams: raw residuals do not have equal variance, even when the errors do. Var(eᵢ) = σ²(1 − hᵢᵢ). That is why you standardise them before comparing them.

Key rules to remember

Residual
eᵢ = yᵢ − ŷᵢ
With an intercept in the model, the residuals sum to zero. This is always true and is not evidence of a good fit.
Variance estimate
σ̂² = Σ eᵢ² ÷ (n − p)
p is the number of fitted parameters including the intercept. For simple linear regression p = 2.
Variance of a residual
Var(eᵢ) = σ²(1 − hᵢᵢ)
Residuals have unequal variances when leverages differ.
Standardised residual
rᵢ = eᵢ ÷ (σ̂ √(1 − hᵢᵢ))
Values beyond about ±2 deserve a look, and beyond ±3 are rare under a good model. This is a rule of thumb, not a test.
Leverage
hᵢᵢ = xᵢᵀ(XᵀX)⁻¹xᵢ; for simple regression hᵢᵢ = 1/n + (xᵢ − x̄)² ÷ Sxx
The leverages sum to p, so the average leverage is p/n. A common rule of thumb flags hᵢᵢ > 2p/n.
Cook's distance
Dᵢ = (rᵢ² ÷ p) × hᵢᵢ ÷ (1 − hᵢᵢ)
Here rᵢ is the standardised residual above. Large values mean the point is influential. Compare with the other points in the data, and treat fixed cut-offs as rough guides.

How to solve Residuals and Model Diagnostics questions

Use this order for any question that asks you to assess a regression model from residuals or diagnostics.

  1. 1State the assumptions being tested: linearity, independence, constant variance, normality, and no undue influence.
  2. 2Compute or read the residuals eᵢ = yᵢ − ŷᵢ. If asked, standardise them using rᵢ = eᵢ ÷ (σ̂ √(1 − hᵢᵢ)).
  3. 3Look at residuals vs fitted values (and vs each predictor or vs order). Name the pattern you see: curve, funnel, runs, or random scatter.
  4. 4Link the pattern to the assumption it breaks. Curve means non-linearity, funnel means non-constant variance, runs means dependence.
  5. 5Check normality with a QQ plot of standardised residuals. Say what the shape shows about tails or skewness.
  6. 6Check leverage and influence. Compare hᵢᵢ with 2p/n, look at large |rᵢ|, and compute Cook's distance if the data allow.
  7. 7Say what the consequence is for the model (for example, unreliable standard errors and intervals) and propose a remedy: transform y or x, add a term, use weights or a GLM, or investigate the data point.
  8. 8Do not delete a point only because it is unusual. Say you would check it for data error first.

Quickest way: Pattern-to-problem shortcut

When to use it: Use this for multiple-choice questions and short written parts that show a plot or describe one.

  1. Curved band in residuals vs fitted: the mean structure is wrong. Add a squared term or transform.
  2. Funnel (spread grows with fitted value): non-constant variance. Try a log transform of y, weights, or a GLM with a suitable variance function.
  3. Runs of same-sign residuals in time order: positive autocorrelation. Standard errors are too small, so t-tests look too good.
  4. QQ plot S-shaped or bending at the ends: non-normal errors. Tails or skewness are the issue.
  5. One point far out in x: high leverage. One point far out in y: outlier. Both together: likely influential, so check Cook's distance.
  6. For numbers, compute h = 1/n + (x − x̄)²/Sxx and compare with 2p/n, then rᵢ, then D.

Common mistakes in Residuals and Model Diagnostics

  • Treating heteroscedasticity and autocorrelation as the same thing.

    Both are described as 'residuals not behaving', and both affect standard errors.

    Fix: Heteroscedasticity is non-constant variance and shows as a funnel against fitted values. Autocorrelation is dependence between errors, shown by runs when residuals are plotted in order. Name the assumption broken in each case.

  • Using raw residuals to find outliers when leverages differ.

    Raw residuals are easy to compute, so students skip the standardising step.

    Fix: Divide by σ̂√(1 − hᵢᵢ). High-leverage points pull the line towards themselves, so their raw residuals are often small.

  • Saying a high-leverage point is automatically bad or must be removed.

    Students confuse leverage with influence.

    Fix: High leverage only means unusual x-values. A high-leverage point that lies close to the fitted line is harmless. Check the residual and Cook's distance, and investigate the data before any removal.

  • Reading a QQ plot with too much precision, or using it on raw residuals with unequal variances.

    Students look for perfect straightness, and small samples always wobble.

    Fix: Use standardised residuals. Look for clear systematic departures such as strong curvature, not small deviations.

  • Treating residuals summing to zero as a check that the model is good.

    It looks like a successful test.

    Fix: With an intercept, least squares forces Σeᵢ = 0. It tells you nothing about fit. Only patterns in the residuals do.

  • Using n instead of n − p in σ̂², or using p = 1 for simple regression.

    Students copy the sample variance denominator.

    Fix: Count every fitted parameter. Simple linear regression has p = 2, so σ̂² = RSS ÷ (n − 2).

Worked examples

Example 1

A simple linear regression is fitted to n = 20 observations. The mean of x is 10 and Sxx = 400. The estimate of σ is σ̂ = 2. For the observation with x = 18, the residual is 3.5. (a) Find its leverage and say whether it is high by the 2p/n rule. (b) Find the standardised residual. (c) Find Cook's distance.

Show the solution
  1. Here p = 2 and n = 20, so the threshold is 2p/n = 4/20 = 0.2.
  2. (a) h = 1/20 + (18 − 10)² ÷ 400 = 0.05 + 64/400 = 0.05 + 0.16 = 0.21.
  3. Since 0.21 > 0.2, the point is flagged as high leverage.
  4. (b) 1 − h = 0.79 and √0.79 = 0.8888. So r = 3.5 ÷ (2 × 0.8888) = 3.5 ÷ 1.7776 = 1.969.
  5. (c) D = (r² ÷ p) × h ÷ (1 − h). r² = 3.877, and r²/p = 1.938. Also h/(1 − h) = 0.21/0.79 = 0.2658.
  6. So D = 1.938 × 0.2658 = 0.515.

Answer: Leverage = 0.21 (just above the 0.2 threshold), standardised residual ≈ 1.97, Cook's distance ≈ 0.52. The point has high leverage and a fairly large residual, so it is worth checking, but the evidence is not extreme.

Example 2

An actuary models claim cost per policy (y) against policyholder age (x) using ŷ = 20 + 3x. For a policy with x = 5 and y = 41, σ̂ = 4 and the leverage is 0.1. (a) Find the residual and standardised residual. (b) The plot of standardised residuals against fitted values shows a funnel that widens to the right. Which assumption fails, what is the effect, and what would you do?

Show the solution
  1. (a) The fitted value is ŷ = 20 + 3 × 5 = 35. The residual is e = 41 − 35 = 6.
  2. 1 − h = 0.9 and √0.9 = 0.9487. So r = 6 ÷ (4 × 0.9487) = 6 ÷ 3.795 = 1.58.
  3. (b) A widening funnel shows that the spread of the residuals grows as the fitted value grows. The constant variance assumption (homoscedasticity) fails.
  4. The least squares estimates of the coefficients remain unbiased, but the usual standard errors, confidence intervals and tests are unreliable, and the fit is not the most efficient.
  5. Remedies: model log(y) if the spread grows in proportion to the mean, use weighted least squares, or use a GLM (for example gamma) whose variance rises with the mean. Then re-check the residual plot.

Answer: Residual = 6, standardised residual ≈ 1.58 (not unusual). The funnel shows non-constant variance. Inference is unreliable, so transform y, use weights, or use a GLM, then re-check the plots.

Exam tips

  • When a question says 'comment on the plot', do three things: describe the pattern, name the assumption it breaks, and give a remedy. Marks are split across all three.
  • Always say what the plot is against: fitted values, a predictor, or time order. Autocorrelation only shows up in order plots.
  • Show the formula with hᵢᵢ and p stated before you substitute. Count p as the number of parameters including the intercept.
  • Treat 2p/n and ±2 as rules of thumb in your wording. Do not call them tests.
  • In the computer-based paper, produce the plots from the fitted model, label the axes, and write one sentence of interpretation under each plot.

Practice questions from Linear regression models

Residuals and Model Diagnostics in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Residuals and Model Diagnostics: frequently asked questions

How do I check normality of residuals in regression using a QQ plot?

Plot the standardised residuals against theoretical normal quantiles. If the points lie close to a straight line, normality is reasonable. Systematic curves or S-shapes point to skewness or heavy or light tails. With few points, expect some wobble.

What is the difference between heteroscedasticity and autocorrelation in residuals?

Heteroscedasticity means the error variance is not constant, often seen as a funnel in residuals vs fitted values. Autocorrelation means errors are correlated with each other, often seen as runs of same-sign residuals when plotted in time order. They break different assumptions.

What is the difference between an outlier and a high-leverage point?

An outlier has an unusually large residual, so its y-value is far from the fitted line. A high-leverage point has unusual x-values. A point that is both can be influential, and Cook's distance measures that combined effect.

Why do we standardise residuals?

Even when the errors have constant variance, the residuals do not, because Var(eᵢ) = σ²(1 − hᵢᵢ). Dividing by σ̂√(1 − hᵢᵢ) puts the residuals on a comparable scale, so you can judge which are unusually large.