Skip to content

IAI Actuarial Core Principles · Actuarial Statistics

Linear Regression Models for IAI CS1 Actuarial Statistics

Linear regression models describe how a response variable depends on one or more explanatory variables, using Y = β0 + β1x + ε. You estimate the parameters by least squares, test them with t and F tests, check fit with R² and ANOVA, then check residuals to see whether the assumptions hold.

What this chapter covers

This chapter builds the standard linear model. You start with one explanatory variable and the assumptions on the errors. You then derive the least squares estimators, find their distributions, and use them for confidence intervals, prediction intervals and hypothesis tests. After that you split the total variation into explained and unexplained parts, and move to several explanatory variables using matrix notation. You finish by checking whether the model is sensible.

Regression theory and applications is the largest block in the CS1 syllabus weightings, so this chapter carries real weight. It also draws on earlier work. You need the normal, t, chi-square and F distributions from random variables and distributions. You need estimation and hypothesis testing from statistical inference. If those are weak, regression will feel harder than it is.

The chapter also connects forward. Generalised linear models, time series and survival models in CS2 reuse the ideas of fitting a model, testing parameters and examining residuals. Paper B, the computer-based exam, often asks you to fit a regression in R, read the output and comment. So you need both the algebra and the interpretation.

Regression theory and applications has the highest topic weighting in the CS1 syllabus, so this chapter is worth serious effort. It suits both exam styles. Multiple-choice questions test quick facts such as degrees of freedom, the meaning of R² and which assumption a plot checks. Written questions ask you to derive estimators, build intervals, complete an ANOVA table and comment on diagnostics. Paper B asks you to run the model in R and interpret the output. Students who learn the method once can pick up marks across all three formats, and the same skills carry into CS2.

Linear regression models: topics in the order to study them

  1. 1Simple Linear Regression Model and AssumptionsEverything later depends on the model form and the assumptions on the errors, so learn these first.
  2. 2Least Squares Estimation of ParametersYou need the estimators, their means and their variances before you can test or interval anything.
  3. 3Inference: Confidence Intervals and Hypothesis TestsOnce you know the distribution of the estimators, you can build t-based intervals and tests, including prediction.
  4. 4Goodness of Fit and ANOVA for RegressionThis splits the total sum of squares and links the F test to the t test, which you can only follow after inference.
  5. 5Multiple Linear Regression and Matrix FormIt generalises the simple case, so the simple results act as a check on the matrix results.
  6. 6Residuals and Model DiagnosticsStudy it last so you can judge the assumptions in all the earlier models and tie the chapter together.

How to prepare Linear regression models

Work from the simple model outward, and practise by hand and in R. Aim to be able to reproduce each result and explain it in words.

  1. Write the model and the error assumptions from memory: errors independent, mean zero, constant variance σ², and normal for inference. Say which results need normality and which do not.
  2. Derive the least squares estimators for the simple model. Learn Sxx, Sxy and Syy, and the forms β̂1 = Sxy ÷ Sxx and β̂0 = ȳ − β̂1x̄. Do this until it takes a few minutes.
  3. Practise inference with real numbers. Use the estimate of σ² with n − 2 degrees of freedom, build the t statistic, and give confidence intervals for the slope, the mean response and a new observation. Be clear why the last one is wider.
  4. Build the ANOVA table by hand: SST, SSR, SSE, degrees of freedom, mean squares and F. Check that R² = SSR ÷ SST and see how F relates to the square of the slope t statistic in the simple model.
  5. Learn the matrix form: β̂ = (XᵀX)⁻¹Xᵀy. Practise setting up the design matrix, then interpret coefficients, adjusted R² and the F test for the whole model. Use n − p degrees of freedom, where p counts all parameters including the intercept.
  6. Fit models in R with lm() and read summary() and anova() output. Practise plotting residuals against fitted values and a normal Q-Q plot, then write two or three sentences of comment on what you see and what you would change.
  7. Finish with timed past-paper questions. Check each answer for correct degrees of freedom, units and a stated conclusion.

Common mistakes in Linear regression models

  • Using the wrong degrees of freedom for t tests and σ̂².

    Fix: Count the estimated parameters p, including the intercept, and use n − p. Write this down before you look up any table.

  • Mixing up the confidence interval for the mean response and the prediction interval for a new observation.

    Fix: Ask what is being estimated. For a single new value, add σ² to the variance, which gives the extra 1 inside the square root. Check that your prediction interval is the wider one.

  • Treating a large R² as proof that the model is correct.

    Fix: Say that R² measures the share of variation explained, nothing more. Always check the residual plots for curvature, outliers and changing spread.

  • Stating a conclusion from a test without linking it to the hypotheses and the context.

    Fix: State H0 and H1, the statistic, the critical value or p-value, the decision and a one-line conclusion in context.

  • Confusing the t test on one coefficient with the F test for the whole multiple regression model.

    Fix: Remember that the F test asks whether any explanatory variable helps, while each t test asks about one coefficient given the others in the model.

  • Reading R output without stating assumptions or commenting on diagnostics in Paper B.

    Fix: For each output, write what the number means, whether the assumptions look reasonable from the plots, and what you would do next.

Last-day revision: Linear regression models

  • Model: Yi = β0 + β1xi + εi, with errors independent, mean 0 and constant variance σ².
  • For inference, assume errors are normal: εi ~ N(0, σ²).
  • β̂1 = Sxy ÷ Sxx and β̂0 = ȳ − β̂1x̄, where Sxy = Σ(xi − x̄)(yi − ȳ) and Sxx = Σ(xi − x̄)².
  • Var(β̂1) = σ² ÷ Sxx. Replace σ² with its estimate to get the standard error.
  • σ̂² = SSE ÷ (n − 2) in simple regression, and SSE ÷ (n − p) with p parameters in general.
  • The t statistic for a slope is (β̂1 − β1) ÷ se(β̂1), with n − 2 degrees of freedom in simple regression.
  • A prediction interval for a new observation is wider than the confidence interval for the mean response, because it adds the error variance.
  • SST = SSR + SSE, and R² = SSR ÷ SST.
  • In simple regression, the F statistic equals the square of the slope t statistic.
  • Multiple regression: β̂ = (XᵀX)⁻¹Xᵀy, and Var(β̂) = σ²(XᵀX)⁻¹.
  • Adjusted R² penalises extra explanatory variables, while plain R² never falls when you add one.
  • Residuals against fitted values should show no pattern. A curve suggests a missing term, and a funnel suggests non-constant variance.

Linear regression models practice questions

Linear regression models in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Linear regression models: frequently asked questions

How much of this chapter is derivation and how much is interpretation?

Both matter. Written questions often ask you to derive or show a result for the simple model, and then to interpret numbers. Paper B focuses on fitting the model in R and commenting on the output.

Do I need to memorise the matrix results?

Yes, learn β̂ = (XᵀX)⁻¹Xᵀy and Var(β̂) = σ²(XᵀX)⁻¹, and know how to set up the design matrix. Check them against the simple regression case to build confidence.

Which assumptions need normal errors?

Least squares estimates and their means and variances need only the zero-mean, constant-variance and uncorrelated error assumptions. The exact t and F distributions used for intervals and tests rely on normal errors.

What should I do if the residual plot shows a pattern?

Say what the pattern suggests, such as a missing curved term or non-constant variance. Then propose a fix, such as adding a term or transforming the response, and refit the model.