Skip to content

IAI Actuarial Core Principles · Actuarial Statistics

Linear regression models: formula sheet

Full chapter guide

Key formulas

Simple linear regression model
Yᵢ = α + βxᵢ + εᵢ, i = 1, …, n
Y is the response, x the explanatory variable (fixed), ε the random error.
Mean of the response
E[Yᵢ] = α + βxᵢ
Follows because E[εᵢ] = 0 and xᵢ is fixed.
Variance of the response
Var(Yᵢ) = σ²
Same at every xᵢ. This is constant variance.
Error assumptions
E[εᵢ] = 0, Var(εᵢ) = σ², Cov(εᵢ, εⱼ) = 0 for i ≠ j
These are the basic assumptions for least squares to behave well.
Normal error model
εᵢ ~ N(0, σ²) independent, so Yᵢ ~ N(α + βxᵢ, σ²)
Needed for exact t-tests and confidence intervals.
Sxx
Sxx = Σ(xi − x̄)² = Σxi² − n x̄²
Use the second form for calculations. It needs only Σx and Σx².
Sxy
Sxy = Σ(xi − x̄)(yi − ȳ) = Σxiyi − n x̄ ȳ
Sign of Sxy is the sign of the slope.
Syy
Syy = Σ(yi − ȳ)² = Σyi² − n ȳ²
Needed for σ̂², SSE and R².
Normal equations
Σy = nα + βΣx; Σxy = αΣx + βΣx²
Obtained by setting ∂S/∂α = 0 and ∂S/∂β = 0.
Least squares slope
β̂ = Sxy ÷ Sxx
Also equal to Σ(xi − x̄)yi ÷ Sxx.
Least squares intercept
α̂ = ȳ − β̂ x̄
Fitted line passes through (x̄, ȳ).
Mean of estimators
E(β̂) = β and E(α̂) = α
Both are unbiased under the model assumptions.
Variance of slope
Var(β̂) = σ² ÷ Sxx
Larger spread in x gives smaller variance.
Variance of intercept
Var(α̂) = σ² (1/n + x̄²/Sxx)
Smallest when x̄ = 0.
Covariance
Cov(α̂, β̂) = −σ² x̄ ÷ Sxx
Zero only if x̄ = 0.
Residual sum of squares
SSE = Syy − Sxy² ÷ Sxx
Equals Σ(yi − ŷi)².
Unbiased estimator of σ²
σ̂² = SSE ÷ (n − 2)
Divide by n − 2 because two parameters are estimated.
Model
Yi = α + β xi + εi, with εi independent N(0, σ²)
State these assumptions before using any t result.
Sums of squares
Sxx = Σx² − n x̄²; Sxy = Σxy − n x̄ ȳ; Syy = Σy² − n ȳ²
Most exam questions give these or the raw sums.
Least squares estimates
β̂ = Sxy ÷ Sxx; α̂ = ȳ − β̂ x̄
The fitted line always passes through (x̄, ȳ).
Variances of estimators
Var(β̂) = σ² ÷ Sxx; Var(α̂) = σ² (1/n + x̄²/Sxx); Cov(α̂, β̂) = −σ² x̄ ÷ Sxx
Both estimators are unbiased and normally distributed.
Error variance estimate
σ̂² = SSres ÷ (n − 2), where SSres = Syy − Sxy² ÷ Sxx
Unbiased for σ². Also (n − 2)σ̂²/σ² ~ χ² with n − 2 degrees of freedom.
t pivot for the slope
(β̂ − β) ÷ (σ̂ ÷ √Sxx) ~ t(n − 2)
For the test of H0: β = 0, the statistic is β̂ ÷ (σ̂ ÷ √Sxx).
Confidence interval for the slope
β̂ ± t(n − 2, 1 − γ/2) × σ̂ ÷ √Sxx
Here the confidence level is 1 − γ, for example 95% when γ = 0.05.
Confidence interval for the intercept
α̂ ± t(n − 2, 1 − γ/2) × σ̂ √(1/n + x̄²/Sxx)
Often the intercept is outside the data range, so interpret with care.
Confidence interval for the mean response at x0
(α̂ + β̂ x0) ± t × σ̂ √(1/n + (x0 − x̄)²/Sxx)
Estimates E[Y | x0], the average response.
Prediction interval for a new observation at x0
(α̂ + β̂ x0) ± t × σ̂ √(1 + 1/n + (x0 − x̄)²/Sxx)
The extra 1 is the variance of the new error. It is always wider than the mean-response interval.
Link with F
t² = F, with 1 and n − 2 degrees of freedom
For the slope test in simple regression, the t-test and ANOVA F-test agree.
Sums of squares
SST = Σ(yᵢ − ȳ)² = Syy; SSR = Σ(ŷᵢ − ȳ)²; SSE = Σ(yᵢ − ŷᵢ)²
Syy is the usual notation for the corrected sum of squares of y.
Decomposition
SST = SSR + SSE
Holds when the model includes an intercept.
Coefficient of determination
R² = SSR ÷ SST = 1 − SSE ÷ SST
Proportion of variation in y explained by the model.
Simple regression shortcuts
SSR = Sxy² ÷ Sxx = β̂ × Sxy; SSE = Syy − Sxy² ÷ Sxx
Here Sxx = Σ(xᵢ − x̄)², Sxy = Σ(xᵢ − x̄)(yᵢ − ȳ) and β̂ = Sxy ÷ Sxx.
Link to correlation
r = Sxy ÷ √(Sxx × Syy); R² = r²
Valid for simple linear regression only.
Degrees of freedom
SSR: 1; SSE: n − 2; SST: n − 1
For multiple regression with k explanatory variables: k, n − k − 1, n − 1.
Mean squares and F statistic
MSR = SSR ÷ 1; MSE = SSE ÷ (n − 2) = σ̂²; F = MSR ÷ MSE ~ F(1, n − 2) under H₀: β = 0
Reject H₀ for large F. The test is one-sided in F but equals a two-sided t test.
F and t link
F = t², where t = β̂ ÷ se(β̂)
t has n − 2 degrees of freedom.
F in terms of R²
F = R² × (n − 2) ÷ (1 − R²)
Useful when only R² and n are given.
Model in matrix form
y = Xβ + ε, with E(ε) = 0 and Var(ε) = σ²I
X is n × p with p = k + 1. The first column of X is all 1s.
Normal equations and estimator
X'X β̂ = X'y, so β̂ = (X'X)⁻¹X'y
Needs X'X invertible, i.e. no column of X is a linear combination of the others.
Variance of the estimator
Var(β̂) = σ²(X'X)⁻¹
The standard error of β̂ⱼ is σ̂ times the square root of the j-th diagonal entry of (X'X)⁻¹.
Variance estimate
σ̂² = RSS ÷ (n − p), where RSS = (y − Xβ̂)'(y − Xβ̂)
This is unbiased for σ². Degrees of freedom are n − p, not n − k.
t test for one coefficient
t = β̂ⱼ ÷ se(β̂ⱼ), compared with t on n − p degrees of freedom
Tests H₀: βⱼ = 0, given the other variables are in the model.
R² and adjusted R²
R² = 1 − RSS ÷ TSS; adjusted R² = 1 − [RSS ÷ (n − k − 1)] ÷ [TSS ÷ (n − 1)]
TSS = Σ(yᵢ − ȳ)². Use adjusted R² to compare models with different numbers of variables.
Overall F test
F = [(TSS − RSS) ÷ k] ÷ [RSS ÷ (n − k − 1)], on k and n − k − 1 degrees of freedom
Tests H₀: β₁ = … = βₖ = 0.
Partial F test for nested models
F = [(RSS_small − RSS_large) ÷ q] ÷ [RSS_large ÷ (n − p_large)]
q is the number of extra parameters in the larger model. Compare with F on q and n − p_large degrees of freedom.
Hat matrix and fitted values
ŷ = Hy, where H = X(X'X)⁻¹X'
The diagonal entries hᵢᵢ are leverages. Their sum equals p.
Dummy variable count
A factor with m levels needs m − 1 dummy variables
Keeping all m with an intercept makes X'X singular.
Residual
eᵢ = yᵢ − ŷᵢ
With an intercept in the model, the residuals sum to zero. This is always true and is not evidence of a good fit.
Variance estimate
σ̂² = Σ eᵢ² ÷ (n − p)
p is the number of fitted parameters including the intercept. For simple linear regression p = 2.
Variance of a residual
Var(eᵢ) = σ²(1 − hᵢᵢ)
Residuals have unequal variances when leverages differ.
Standardised residual
rᵢ = eᵢ ÷ (σ̂ √(1 − hᵢᵢ))
Values beyond about ±2 deserve a look, and beyond ±3 are rare under a good model. This is a rule of thumb, not a test.
Leverage
hᵢᵢ = xᵢᵀ(XᵀX)⁻¹xᵢ; for simple regression hᵢᵢ = 1/n + (xᵢ − x̄)² ÷ Sxx
The leverages sum to p, so the average leverage is p/n. A common rule of thumb flags hᵢᵢ > 2p/n.
Cook's distance
Dᵢ = (rᵢ² ÷ p) × hᵢᵢ ÷ (1 − hᵢᵢ)
Here rᵢ is the standardised residual above. Large values mean the point is influential. Compare with the other points in the data, and treat fixed cut-offs as rough guides.

Quick revision

  • Model: Yi = β0 + β1xi + εi, with errors independent, mean 0 and constant variance σ².
  • For inference, assume errors are normal: εi ~ N(0, σ²).
  • β̂1 = Sxy ÷ Sxx and β̂0 = ȳ − β̂1x̄, where Sxy = Σ(xi − x̄)(yi − ȳ) and Sxx = Σ(xi − x̄)².
  • Var(β̂1) = σ² ÷ Sxx. Replace σ² with its estimate to get the standard error.
  • σ̂² = SSE ÷ (n − 2) in simple regression, and SSE ÷ (n − p) with p parameters in general.
  • The t statistic for a slope is (β̂1 − β1) ÷ se(β̂1), with n − 2 degrees of freedom in simple regression.
  • A prediction interval for a new observation is wider than the confidence interval for the mean response, because it adds the error variance.
  • SST = SSR + SSE, and R² = SSR ÷ SST.
  • In simple regression, the F statistic equals the square of the slope t statistic.
  • Multiple regression: β̂ = (XᵀX)⁻¹Xᵀy, and Var(β̂) = σ²(XᵀX)⁻¹.
  • Adjusted R² penalises extra explanatory variables, while plain R² never falls when you add one.
  • Residuals against fitted values should show no pattern. A curve suggests a missing term, and a funnel suggests non-constant variance.

Common mistakes

  • Saying the explanatory variable x is random with a variance. Fix: In this model x is fixed and known. Only ε, and therefore Y, is random.
  • Writing the assumptions about Y instead of ε, or forgetting that the mean of Y changes with x. Fix: The mean of Y is α + βxᵢ, which varies. Only the variance σ² is constant.
  • Dividing by n − 1 or n instead of n − 2 when estimating σ². Fix: Regression estimates two parameters, α and β. So σ̂² = SSE ÷ (n − 2) is the unbiased estimator.
  • Using Σx² in place of Sxx, or Σxy in place of Sxy. Fix: Always write Sxx = Σx² − n x̄² and Sxy = Σxy − n x̄ ȳ before computing the slope.
  • Using n − 1 or n as the degrees of freedom for the t distribution. Fix: In simple linear regression two parameters are estimated, so use n − 2 for both σ̂² and the t critical value.
  • Using the mean-response interval when the question asks for a prediction interval, or the reverse. Fix: Check the wording. Average or expected response means no extra 1. A single new observation means include the 1 inside the square root.
  • Using n − 1 degrees of freedom for SSE. Fix: SSE has n − 2 degrees of freedom because two parameters, α and β, are estimated. n − 1 belongs to SST.
  • Saying R² = 0.9 means the model is correct or that x causes y. Fix: Say only that 90% of the variation in y is explained by the fitted line. Check residual plots for the model assumptions and remember that correlation is not causation.
  • Using n − k instead of n − k − 1 (or n − p) as the residual degrees of freedom. Fix: Count every estimated β, including the intercept. Residual degrees of freedom are always n − p.
  • Including m dummy variables for a factor with m levels alongside an intercept. Fix: Use m − 1 dummies and name the baseline level. The all-dummy columns would add up to the intercept column, so X'X would be singular.

Exam tips

  • Always write the model with subscripts and say that x is fixed and ε is random. Examiners give marks for this.
  • State each assumption separately. Do not bundle them into one phrase, as marks are awarded per assumption.
  • When asked to interpret, use the units and context in the question, not just 'a unit increase'.
  • In written answers on diagnostics, name the plot, describe what you see and say which assumption it challenges.
  • In the computer-based paper, the same assumptions guide your checks of the residual plots in R, so link each plot to an assumption.
  • Show the three sums Sxx, Sxy and Syy clearly. Marks are often given for each, even if the final arithmetic slips.
  • For derivation questions, set up β̂ as Σ ci Yi with ci = (xi − x̄)/Sxx. This makes unbiasedness and variance short.
  • State which assumptions you use at each step. Examiners separate results that need only zero mean and constant variance from those that need normality.