IAI Actuarial Core Principles · Actuarial Statistics
Linear regression models: formula sheet
Key formulas
- Simple linear regression model
- Yᵢ = α + βxᵢ + εᵢ, i = 1, …, n
- Y is the response, x the explanatory variable (fixed), ε the random error.
- Mean of the response
- E[Yᵢ] = α + βxᵢ
- Follows because E[εᵢ] = 0 and xᵢ is fixed.
- Variance of the response
- Var(Yᵢ) = σ²
- Same at every xᵢ. This is constant variance.
- Error assumptions
- E[εᵢ] = 0, Var(εᵢ) = σ², Cov(εᵢ, εⱼ) = 0 for i ≠ j
- These are the basic assumptions for least squares to behave well.
- Normal error model
- εᵢ ~ N(0, σ²) independent, so Yᵢ ~ N(α + βxᵢ, σ²)
- Needed for exact t-tests and confidence intervals.
- Sxx
- Sxx = Σ(xi − x̄)² = Σxi² − n x̄²
- Use the second form for calculations. It needs only Σx and Σx².
- Sxy
- Sxy = Σ(xi − x̄)(yi − ȳ) = Σxiyi − n x̄ ȳ
- Sign of Sxy is the sign of the slope.
- Syy
- Syy = Σ(yi − ȳ)² = Σyi² − n ȳ²
- Needed for σ̂², SSE and R².
- Normal equations
- Σy = nα + βΣx; Σxy = αΣx + βΣx²
- Obtained by setting ∂S/∂α = 0 and ∂S/∂β = 0.
- Least squares slope
- β̂ = Sxy ÷ Sxx
- Also equal to Σ(xi − x̄)yi ÷ Sxx.
- Least squares intercept
- α̂ = ȳ − β̂ x̄
- Fitted line passes through (x̄, ȳ).
- Mean of estimators
- E(β̂) = β and E(α̂) = α
- Both are unbiased under the model assumptions.
- Variance of slope
- Var(β̂) = σ² ÷ Sxx
- Larger spread in x gives smaller variance.
- Variance of intercept
- Var(α̂) = σ² (1/n + x̄²/Sxx)
- Smallest when x̄ = 0.
- Covariance
- Cov(α̂, β̂) = −σ² x̄ ÷ Sxx
- Zero only if x̄ = 0.
- Residual sum of squares
- SSE = Syy − Sxy² ÷ Sxx
- Equals Σ(yi − ŷi)².
- Unbiased estimator of σ²
- σ̂² = SSE ÷ (n − 2)
- Divide by n − 2 because two parameters are estimated.
- Model
- Yi = α + β xi + εi, with εi independent N(0, σ²)
- State these assumptions before using any t result.
- Sums of squares
- Sxx = Σx² − n x̄²; Sxy = Σxy − n x̄ ȳ; Syy = Σy² − n ȳ²
- Most exam questions give these or the raw sums.
- Least squares estimates
- β̂ = Sxy ÷ Sxx; α̂ = ȳ − β̂ x̄
- The fitted line always passes through (x̄, ȳ).
- Variances of estimators
- Var(β̂) = σ² ÷ Sxx; Var(α̂) = σ² (1/n + x̄²/Sxx); Cov(α̂, β̂) = −σ² x̄ ÷ Sxx
- Both estimators are unbiased and normally distributed.
- Error variance estimate
- σ̂² = SSres ÷ (n − 2), where SSres = Syy − Sxy² ÷ Sxx
- Unbiased for σ². Also (n − 2)σ̂²/σ² ~ χ² with n − 2 degrees of freedom.
- t pivot for the slope
- (β̂ − β) ÷ (σ̂ ÷ √Sxx) ~ t(n − 2)
- For the test of H0: β = 0, the statistic is β̂ ÷ (σ̂ ÷ √Sxx).
- Confidence interval for the slope
- β̂ ± t(n − 2, 1 − γ/2) × σ̂ ÷ √Sxx
- Here the confidence level is 1 − γ, for example 95% when γ = 0.05.
- Confidence interval for the intercept
- α̂ ± t(n − 2, 1 − γ/2) × σ̂ √(1/n + x̄²/Sxx)
- Often the intercept is outside the data range, so interpret with care.
- Confidence interval for the mean response at x0
- (α̂ + β̂ x0) ± t × σ̂ √(1/n + (x0 − x̄)²/Sxx)
- Estimates E[Y | x0], the average response.
- Prediction interval for a new observation at x0
- (α̂ + β̂ x0) ± t × σ̂ √(1 + 1/n + (x0 − x̄)²/Sxx)
- The extra 1 is the variance of the new error. It is always wider than the mean-response interval.
- Link with F
- t² = F, with 1 and n − 2 degrees of freedom
- For the slope test in simple regression, the t-test and ANOVA F-test agree.
- Sums of squares
- SST = Σ(yᵢ − ȳ)² = Syy; SSR = Σ(ŷᵢ − ȳ)²; SSE = Σ(yᵢ − ŷᵢ)²
- Syy is the usual notation for the corrected sum of squares of y.
- Decomposition
- SST = SSR + SSE
- Holds when the model includes an intercept.
- Coefficient of determination
- R² = SSR ÷ SST = 1 − SSE ÷ SST
- Proportion of variation in y explained by the model.
- Simple regression shortcuts
- SSR = Sxy² ÷ Sxx = β̂ × Sxy; SSE = Syy − Sxy² ÷ Sxx
- Here Sxx = Σ(xᵢ − x̄)², Sxy = Σ(xᵢ − x̄)(yᵢ − ȳ) and β̂ = Sxy ÷ Sxx.
- Link to correlation
- r = Sxy ÷ √(Sxx × Syy); R² = r²
- Valid for simple linear regression only.
- Degrees of freedom
- SSR: 1; SSE: n − 2; SST: n − 1
- For multiple regression with k explanatory variables: k, n − k − 1, n − 1.
- Mean squares and F statistic
- MSR = SSR ÷ 1; MSE = SSE ÷ (n − 2) = σ̂²; F = MSR ÷ MSE ~ F(1, n − 2) under H₀: β = 0
- Reject H₀ for large F. The test is one-sided in F but equals a two-sided t test.
- F and t link
- F = t², where t = β̂ ÷ se(β̂)
- t has n − 2 degrees of freedom.
- F in terms of R²
- F = R² × (n − 2) ÷ (1 − R²)
- Useful when only R² and n are given.
- Model in matrix form
- y = Xβ + ε, with E(ε) = 0 and Var(ε) = σ²I
- X is n × p with p = k + 1. The first column of X is all 1s.
- Normal equations and estimator
- X'X β̂ = X'y, so β̂ = (X'X)⁻¹X'y
- Needs X'X invertible, i.e. no column of X is a linear combination of the others.
- Variance of the estimator
- Var(β̂) = σ²(X'X)⁻¹
- The standard error of β̂ⱼ is σ̂ times the square root of the j-th diagonal entry of (X'X)⁻¹.
- Variance estimate
- σ̂² = RSS ÷ (n − p), where RSS = (y − Xβ̂)'(y − Xβ̂)
- This is unbiased for σ². Degrees of freedom are n − p, not n − k.
- t test for one coefficient
- t = β̂ⱼ ÷ se(β̂ⱼ), compared with t on n − p degrees of freedom
- Tests H₀: βⱼ = 0, given the other variables are in the model.
- R² and adjusted R²
- R² = 1 − RSS ÷ TSS; adjusted R² = 1 − [RSS ÷ (n − k − 1)] ÷ [TSS ÷ (n − 1)]
- TSS = Σ(yᵢ − ȳ)². Use adjusted R² to compare models with different numbers of variables.
- Overall F test
- F = [(TSS − RSS) ÷ k] ÷ [RSS ÷ (n − k − 1)], on k and n − k − 1 degrees of freedom
- Tests H₀: β₁ = … = βₖ = 0.
- Partial F test for nested models
- F = [(RSS_small − RSS_large) ÷ q] ÷ [RSS_large ÷ (n − p_large)]
- q is the number of extra parameters in the larger model. Compare with F on q and n − p_large degrees of freedom.
- Hat matrix and fitted values
- ŷ = Hy, where H = X(X'X)⁻¹X'
- The diagonal entries hᵢᵢ are leverages. Their sum equals p.
- Dummy variable count
- A factor with m levels needs m − 1 dummy variables
- Keeping all m with an intercept makes X'X singular.
- Residual
- eᵢ = yᵢ − ŷᵢ
- With an intercept in the model, the residuals sum to zero. This is always true and is not evidence of a good fit.
- Variance estimate
- σ̂² = Σ eᵢ² ÷ (n − p)
- p is the number of fitted parameters including the intercept. For simple linear regression p = 2.
- Variance of a residual
- Var(eᵢ) = σ²(1 − hᵢᵢ)
- Residuals have unequal variances when leverages differ.
- Standardised residual
- rᵢ = eᵢ ÷ (σ̂ √(1 − hᵢᵢ))
- Values beyond about ±2 deserve a look, and beyond ±3 are rare under a good model. This is a rule of thumb, not a test.
- Leverage
- hᵢᵢ = xᵢᵀ(XᵀX)⁻¹xᵢ; for simple regression hᵢᵢ = 1/n + (xᵢ − x̄)² ÷ Sxx
- The leverages sum to p, so the average leverage is p/n. A common rule of thumb flags hᵢᵢ > 2p/n.
- Cook's distance
- Dᵢ = (rᵢ² ÷ p) × hᵢᵢ ÷ (1 − hᵢᵢ)
- Here rᵢ is the standardised residual above. Large values mean the point is influential. Compare with the other points in the data, and treat fixed cut-offs as rough guides.
Quick revision
- Model: Yi = β0 + β1xi + εi, with errors independent, mean 0 and constant variance σ².
- For inference, assume errors are normal: εi ~ N(0, σ²).
- β̂1 = Sxy ÷ Sxx and β̂0 = ȳ − β̂1x̄, where Sxy = Σ(xi − x̄)(yi − ȳ) and Sxx = Σ(xi − x̄)².
- Var(β̂1) = σ² ÷ Sxx. Replace σ² with its estimate to get the standard error.
- σ̂² = SSE ÷ (n − 2) in simple regression, and SSE ÷ (n − p) with p parameters in general.
- The t statistic for a slope is (β̂1 − β1) ÷ se(β̂1), with n − 2 degrees of freedom in simple regression.
- A prediction interval for a new observation is wider than the confidence interval for the mean response, because it adds the error variance.
- SST = SSR + SSE, and R² = SSR ÷ SST.
- In simple regression, the F statistic equals the square of the slope t statistic.
- Multiple regression: β̂ = (XᵀX)⁻¹Xᵀy, and Var(β̂) = σ²(XᵀX)⁻¹.
- Adjusted R² penalises extra explanatory variables, while plain R² never falls when you add one.
- Residuals against fitted values should show no pattern. A curve suggests a missing term, and a funnel suggests non-constant variance.
Common mistakes
- Saying the explanatory variable x is random with a variance. Fix: In this model x is fixed and known. Only ε, and therefore Y, is random.
- Writing the assumptions about Y instead of ε, or forgetting that the mean of Y changes with x. Fix: The mean of Y is α + βxᵢ, which varies. Only the variance σ² is constant.
- Dividing by n − 1 or n instead of n − 2 when estimating σ². Fix: Regression estimates two parameters, α and β. So σ̂² = SSE ÷ (n − 2) is the unbiased estimator.
- Using Σx² in place of Sxx, or Σxy in place of Sxy. Fix: Always write Sxx = Σx² − n x̄² and Sxy = Σxy − n x̄ ȳ before computing the slope.
- Using n − 1 or n as the degrees of freedom for the t distribution. Fix: In simple linear regression two parameters are estimated, so use n − 2 for both σ̂² and the t critical value.
- Using the mean-response interval when the question asks for a prediction interval, or the reverse. Fix: Check the wording. Average or expected response means no extra 1. A single new observation means include the 1 inside the square root.
- Using n − 1 degrees of freedom for SSE. Fix: SSE has n − 2 degrees of freedom because two parameters, α and β, are estimated. n − 1 belongs to SST.
- Saying R² = 0.9 means the model is correct or that x causes y. Fix: Say only that 90% of the variation in y is explained by the fitted line. Check residual plots for the model assumptions and remember that correlation is not causation.
- Using n − k instead of n − k − 1 (or n − p) as the residual degrees of freedom. Fix: Count every estimated β, including the intercept. Residual degrees of freedom are always n − p.
- Including m dummy variables for a factor with m levels alongside an intercept. Fix: Use m − 1 dummies and name the baseline level. The all-dummy columns would add up to the intercept column, so X'X would be singular.
Exam tips
- Always write the model with subscripts and say that x is fixed and ε is random. Examiners give marks for this.
- State each assumption separately. Do not bundle them into one phrase, as marks are awarded per assumption.
- When asked to interpret, use the units and context in the question, not just 'a unit increase'.
- In written answers on diagnostics, name the plot, describe what you see and say which assumption it challenges.
- In the computer-based paper, the same assumptions guide your checks of the residual plots in R, so link each plot to an assumption.
- Show the three sums Sxx, Sxy and Syy clearly. Marks are often given for each, even if the final arithmetic slips.
- For derivation questions, set up β̂ as Σ ci Yi with ci = (xi − x̄)/Sxx. This makes unbiasedness and variance short.
- State which assumptions you use at each step. Examiners separate results that need only zero mean and constant variance from those that need normality.