Actuarial Statistics · Linear regression models
Multiple Linear Regression and Matrix Form Explained
Updated 11 October 2026 · Fact-checked
Multiple linear regression models a response as a linear combination of several explanatory variables plus error. In matrix form, y = Xβ + ε. Solve the normal equations X'Xβ̂ = X'y to get β̂ = (X'X)⁻¹X'y. Then test coefficients with t, test the model with F, and compare models using adjusted R².
Understand Multiple Linear Regression and Matrix Form
Simple linear regression uses one explanatory variable. Multiple linear regression uses k of them. The model is Yᵢ = β₀ + β₁xᵢ₁ + … + βₖxᵢₖ + εᵢ. The errors are assumed independent, with mean 0 and constant variance σ². For inference you also assume they are normal. Each coefficient βⱼ is the change in the mean response for a one-unit rise in xⱼ, holding the other variables fixed.
Writing every equation out is slow, so we stack the data. Let y be the n × 1 vector of responses. Let X be the n × p design matrix, where p = k + 1. Its first column is all 1s for the intercept. Its other columns hold the explanatory variables. Then y = Xβ + ε. The least squares estimate minimises the sum of squared errors (y − Xβ)'(y − Xβ). Setting the derivative to zero gives the normal equations: X'X β̂ = X'y. If X has full column rank, X'X is invertible and β̂ = (X'X)⁻¹X'y.
The estimator β̂ is unbiased. Its variance matrix is σ²(X'X)⁻¹. The diagonal entries give the variances of the individual coefficients. Because we do not know σ², we estimate it by σ̂² = RSS ÷ (n − p). Replacing σ² with σ̂² gives standard errors and a t distribution with n − p degrees of freedom.
Categorical variables enter through dummy (indicator) variables. A factor with m levels needs m − 1 dummies, because one level is the baseline. If you include all m dummies and an intercept, the columns of X are linearly dependent and X'X cannot be inverted. Each dummy coefficient is the difference in mean response between that level and the baseline, other variables held fixed.
Model selection balances fit against simplicity. R² never falls when you add a variable, so it cannot choose between models of different size. Adjusted R² penalises extra parameters. You can also use the partial F test for nested models, AIC or BIC, and stepwise methods (forward selection, backward elimination). Always check that the chosen model makes sense and passes residual diagnostics.
Key rules to remember
- Model in matrix form
- y = Xβ + ε, with E(ε) = 0 and Var(ε) = σ²I
- X is n × p with p = k + 1. The first column of X is all 1s.
- Normal equations and estimator
- X'X β̂ = X'y, so β̂ = (X'X)⁻¹X'y
- Needs X'X invertible, i.e. no column of X is a linear combination of the others.
- Variance of the estimator
- Var(β̂) = σ²(X'X)⁻¹
- The standard error of β̂ⱼ is σ̂ times the square root of the j-th diagonal entry of (X'X)⁻¹.
- Variance estimate
- σ̂² = RSS ÷ (n − p), where RSS = (y − Xβ̂)'(y − Xβ̂)
- This is unbiased for σ². Degrees of freedom are n − p, not n − k.
- t test for one coefficient
- t = β̂ⱼ ÷ se(β̂ⱼ), compared with t on n − p degrees of freedom
- Tests H₀: βⱼ = 0, given the other variables are in the model.
- R² and adjusted R²
- R² = 1 − RSS ÷ TSS; adjusted R² = 1 − [RSS ÷ (n − k − 1)] ÷ [TSS ÷ (n − 1)]
- TSS = Σ(yᵢ − ȳ)². Use adjusted R² to compare models with different numbers of variables.
- Overall F test
- F = [(TSS − RSS) ÷ k] ÷ [RSS ÷ (n − k − 1)], on k and n − k − 1 degrees of freedom
- Tests H₀: β₁ = … = βₖ = 0.
- Partial F test for nested models
- F = [(RSS_small − RSS_large) ÷ q] ÷ [RSS_large ÷ (n − p_large)]
- q is the number of extra parameters in the larger model. Compare with F on q and n − p_large degrees of freedom.
- Hat matrix and fitted values
- ŷ = Hy, where H = X(X'X)⁻¹X'
- The diagonal entries hᵢᵢ are leverages. Their sum equals p.
- Dummy variable count
- A factor with m levels needs m − 1 dummy variables
- Keeping all m with an intercept makes X'X singular.
How to solve Multiple Linear Regression and Matrix Form questions
Use this order for most exam questions on multiple regression, whether the data are given or only summary matrices.
- 1Write down n, k and p = k + 1. Identify which columns of X are the intercept, numeric variables and dummy variables.
- 2Set up the model y = Xβ + ε and state the assumptions the question needs (independent errors, constant variance, normality for tests).
- 3If X'X and X'y are given, solve X'X β̂ = X'y. Use the inverse for a 2 × 2 or diagonal matrix. Otherwise solve the equations directly.
- 4Get fitted values or predictions by substituting into β̂₀ + β̂₁x₁ + …. Then find RSS and σ̂² = RSS ÷ (n − p).
- 5For inference, compute standard errors from σ̂²(X'X)⁻¹. Use t on n − p degrees of freedom for a coefficient or a confidence interval. Use F for the whole model or for nested models.
- 6For model choice, compare adjusted R², the partial F test, or AIC/BIC. Prefer the simpler model if the extra variables are not clearly useful.
- 7State the conclusion in words, in the context of the question. Mention any assumption you rely on, such as normal errors.
Quickest way: Fast route for summary-statistics questions
When to use it: Use when the question gives you TSS, RSS, n and k, or gives X'X and X'y, and asks for estimates or a model comparison.
- Compute p = k + 1 and the degrees of freedom n − p first. Most lost marks come from using the wrong one.
- If X'X is diagonal, divide each entry of X'y by the matching diagonal entry. This gives β̂ at once.
- For model comparison, compute RSS ÷ (n − p) for each model. The model with the smaller value has the larger adjusted R², since TSS ÷ (n − 1) is the same for both.
- For a nested comparison, find the partial F from the drop in RSS. Compare it with the F critical value from the tables.
- Check the answer: R² must lie between 0 and 1, and adjusted R² must not exceed R².
Common mistakes in Multiple Linear Regression and Matrix Form
Using n − k instead of n − k − 1 (or n − p) as the residual degrees of freedom.
In simple regression you remember n − 2. In multiple regression the intercept is easy to forget.
Fix: Count every estimated β, including the intercept. Residual degrees of freedom are always n − p.
Including m dummy variables for a factor with m levels alongside an intercept.
It feels natural to give every level its own column.
Fix: Use m − 1 dummies and name the baseline level. The all-dummy columns would add up to the intercept column, so X'X would be singular.
Choosing the model with the highest R².
R² always rises when a variable is added, even a useless one.
Fix: Compare models with adjusted R², a partial F test or AIC/BIC. Remember that these penalise extra parameters.
Reading a coefficient as the effect of that variable alone, ignoring the other variables.
Students carry over the simple regression interpretation.
Fix: Say 'holding the other explanatory variables fixed'. A coefficient can change in size or sign when other variables are added or removed.
Writing X'X as XX', or inverting in the wrong order, when solving for β̂.
Matrix products are not commutative and dimensions are not checked.
Fix: Check dimensions: X'X is p × p and X'y is p × 1. So β̂ = (X'X)⁻¹X'y is p × 1.
Treating a significant overall F test as proof that every variable is useful.
F tests that at least one slope is non-zero, which is a weak claim.
Fix: Use individual t tests or partial F tests to judge each variable or group of variables. Note that t tests depend on which other variables are in the model.
Worked examples
Example 1
A model y = β₀ + β₁x₁ + β₂x₂ + ε is fitted to n = 6 observations. The columns of X are orthogonal, with X'X = diag(6, 4, 2) and X'y = (30, 8, −3)'. The residual sum of squares is 12. (a) Find β̂. (b) Find σ̂². (c) Predict y when x₁ = 1 and x₂ = 2. (d) Find the standard error of β̂₂.
Show the solution
- Here k = 2, so p = 3. The residual degrees of freedom are n − p = 6 − 3 = 3.
- (a) Since X'X is diagonal, (X'X)⁻¹ = diag(1/6, 1/4, 1/2). So β̂ = (30/6, 8/4, −3/2)' = (5, 2, −1.5)'.
- (b) σ̂² = RSS ÷ (n − p) = 12 ÷ 3 = 4.
- (c) Prediction = 5 + 2(1) + (−1.5)(2) = 5 + 2 − 3 = 4.
- (d) Var(β̂₂) is estimated by σ̂² × (1/2) = 4 × 0.5 = 2. So se(β̂₂) = √2 ≈ 1.414.
Answer: β̂ = (5, 2, −1.5)'; σ̂² = 4; predicted y = 4; se(β̂₂) ≈ 1.414.
Example 2
Two models are fitted to n = 20 observations. TSS = 500. Model A has 2 explanatory variables and RSS = 200. Model B adds one more variable (3 in total) and has RSS = 190. (a) Find R² and adjusted R² for each model. (b) Test at the 5% level whether the extra variable in Model B is needed. The 5% critical value of F(1, 16) is 4.49.
Show the solution
- (a) Model A: R² = 1 − 200/500 = 0.60. Residual degrees of freedom = 20 − 2 − 1 = 17. RSS ÷ 17 = 11.765. TSS ÷ 19 = 26.316. Adjusted R² = 1 − 11.765/26.316 = 1 − 0.4471 = 0.553.
- Model B: R² = 1 − 190/500 = 0.62. Residual degrees of freedom = 16. RSS ÷ 16 = 11.875. Adjusted R² = 1 − 11.875/26.316 = 1 − 0.4513 = 0.549.
- So R² rises from 0.60 to 0.62, but adjusted R² falls from 0.553 to 0.549.
- (b) H₀: the extra coefficient is 0. F = [(200 − 190) ÷ 1] ÷ [190 ÷ 16] = 10 ÷ 11.875 = 0.842.
- Since 0.842 < 4.49, we do not reject H₀ at the 5% level.
Answer: Model A: R² = 0.60, adjusted R² ≈ 0.553. Model B: R² = 0.62, adjusted R² ≈ 0.549. The partial F statistic is about 0.84, below 4.49, so the extra variable is not needed. Choose Model A.
Exam tips
- Show the dimensions of X, y and β in your answer. Examiners give marks for a correct set-up, and it helps you avoid algebra slips.
- Always say what the baseline level is when you use dummy variables, and interpret each dummy coefficient as a difference from that baseline.
- For model selection questions, name the criterion you use and explain why. A bare 'model B is better' earns few marks.
- In computer-based (R) questions, quote the lm() output you rely on: the coefficient estimate, its standard error, the t value, and the residual degrees of freedom. Then comment on it in words.
- In written questions, state assumptions (independent errors, constant variance, normality) before you do any hypothesis test.
Practice questions from Linear regression models
- In a simple linear regression fitted by least squares with an intercept, which statement about the raw residuals e_i = y_i - fitted y_i is a…
- In a simple linear regression with n = 27 observations, the total sum of squares is 500 and the regression sum of squares is 300. What is th…
- In the simple linear regression model Y_i = a + b x_i + e_i, which one of the following sets of assumptions on the errors is the standard on…
- For a simple linear regression with n = 10 observations, x̄ = 4, ȳ = 30, Sxx = 40 and Sxy = 100. The fitted line is used to predict y at x =…
- A simple linear regression with n = 14 has residual standard error s = 4, mean x of 10, and Sxx = 200. At x0 = 14 the fitted value is 50. Wh…
Multiple Linear Regression and Matrix Form in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Multiple Linear Regression and Matrix Form: frequently asked questions
Why does the design matrix have a column of 1s?
The column of 1s carries the intercept β₀. Each row of Xβ then equals β₀ + β₁xᵢ₁ + … + βₖxᵢₖ. If you leave it out, the fitted plane is forced through the origin.
Why can adjusted R² fall when I add a variable?
Adjusted R² divides RSS by n − k − 1, so each new variable removes a degree of freedom. If the drop in RSS is too small to make up for that, the ratio rises and adjusted R² falls. This is a sign the variable adds little.
How many dummy variables do I need for a factor with several levels?
You need one fewer than the number of levels, if the model has an intercept. The omitted level is the baseline. Each dummy coefficient measures the difference from the baseline level.
What is the difference between the t test and the F test in multiple regression?
A t test checks one coefficient, given the other variables in the model. The overall F test checks whether all slope coefficients are zero together. A partial F test checks a chosen group of coefficients, and for a single coefficient it gives the square of the t statistic.
When is X'X not invertible?
It is not invertible when the columns of X are linearly dependent, for example when you include all dummies plus an intercept, or when one variable is an exact combination of others. It can also fail when there are fewer observations than parameters. Remove the redundant column to fix it.