Skip to content

FRM Exam Part I · Regression with Multiple Explanatory Variables

Multiple Regression Model and OLS Assumptions Explained

Updated 11 October 2026 · Fact-checked

A multiple regression models one dependent variable as a linear function of several explanatory variables plus an error. Each slope shows the expected change in Y for a one-unit change in that variable, holding the others constant. OLS estimates are unbiased when the classical assumptions hold, including zero conditional mean of errors and no perfect multicollinearity.

Understand Multiple Regression Model and OLS Assumptions

A multiple regression extends simple regression by using more than one explanatory variable. The model is Yi = β0 + β1X1i + β2X2i + ... + βkXki + εi. Y is the dependent variable. The X variables are explanatory variables (regressors). ε is the error term, which captures everything the model does not explain.

OLS (ordinary least squares) picks the coefficient values that minimise the sum of squared residuals. A residual is the actual Y minus the fitted Y. With several regressors you rarely compute this by hand. The exam gives you the output and asks you to interpret it.

The key idea in interpretation is holding other variables constant. A slope β1 is the expected change in Y for a one-unit rise in X1 when X2 to Xk do not change. This is why adding a variable can change the other slopes. If X1 and X2 are correlated, a simple regression of Y on X1 alone mixes the effect of X1 with part of the effect of X2. The multiple regression separates them. The intercept β0 is the expected Y when all regressors equal zero. It may have no practical meaning.

The classical OLS assumptions make the estimates well behaved. The main ones: (1) the model is linear in the parameters; (2) the error has zero conditional mean given the regressors, E(ε | X1,...,Xk) = 0; (3) observations are independent and identically distributed (or, in a looser version, errors are uncorrelated); (4) large outliers are unlikely, so X and Y have finite fourth moments; (5) there is no perfect multicollinearity, meaning no regressor is an exact linear combination of the others. Under these, OLS estimators are unbiased and consistent. Adding homoskedasticity (constant error variance) and no serial correlation gives the Gauss-Markov result: OLS is BLUE, the best linear unbiased estimator. Adding normal errors allows exact t and F tests in small samples.

The assumption that matters most for bias is zero conditional mean. If you leave out a variable that affects Y and is correlated with an included regressor, that regressor picks up the omitted effect. This is omitted variable bias. Heteroskedasticity and serial correlation do not bias the slopes. They make the usual standard errors wrong. Imperfect multicollinearity does not bias them either. It inflates standard errors.

Key formulas to remember

Multiple regression model
Yi = β0 + β1X1i + β2X2i + ... + βkXki + εi
k regressors, k + 1 coefficients including the intercept.
Fitted value and residual
Ŷi = b0 + b1X1i + ... + bkXki; ei = Yi − Ŷi
OLS minimises Σei².
Interpretation of slope
ΔŶ = bj × ΔXj, other X held constant
Use this to predict the change in Y from a given change in one regressor.
Zero conditional mean
E(εi | X1i, ..., Xki) = 0
Key assumption for unbiasedness. Violated by omitted variables correlated with regressors.
Degrees of freedom
df = n − k − 1
Used for t tests of coefficients. n is observations, k is number of slope coefficients.
Standard error of regression
SER = √[SSR ÷ (n − k − 1)]
Estimate of the standard deviation of the error.
Perfect multicollinearity
Xj = c0 + c1X1 + ... (exact linear relation among regressors)
OLS cannot be computed. Classic case is the dummy variable trap.

How to solve Multiple Regression Model and OLS Assumptions questions

Use this method for most questions on setup, interpretation and assumptions.

  1. 1Write the model and identify Y, each regressor and k. Note the units of each variable.
  2. 2For interpretation, read each slope as the change in Y per one-unit change in that X, with the other regressors held constant. Convert the units asked about.
  3. 3For prediction, substitute the given values into the fitted equation and compute step by step. Include the intercept.
  4. 4For a test of one coefficient, compute t = (b − hypothesised value) ÷ SE(b) and compare with the critical value using n − k − 1 degrees of freedom.
  5. 5For an assumption question, name the violated assumption and state its effect: bias in slopes, wrong standard errors, or inability to estimate.
  6. 6Choose the answer that matches the effect. Omitted correlated variable means biased slopes. Heteroskedasticity means unbiased slopes but invalid standard errors. Perfect multicollinearity means OLS fails.
  7. 7Check the answer for sign and size. Does it make sense in the units given?

Quickest way: Plug in and match the violation

When to use it: Use it when the question gives a fitted equation or describes a data problem and offers four options.

  1. For numeric questions, multiply each slope by its change in X and add. Ignore the intercept if only a change is asked.
  2. For assumption questions, ask one thing: does the problem bias the slopes, or only the standard errors?
  3. Eliminate options that say heteroskedasticity or multicollinearity biases the coefficients.
  4. If an exact linear relation exists among regressors, pick the option that says OLS cannot be estimated.

Common mistakes in Multiple Regression Model and OLS Assumptions

  • Reading a slope without saying other variables are held constant, or reading it as the simple regression effect.

    Students carry over simple regression thinking.

    Fix: Always interpret each slope as a partial effect. Adding a correlated regressor can change it.

  • Saying heteroskedasticity or multicollinearity makes OLS coefficients biased.

    Students mix up bias with unreliable standard errors.

    Fix: Remember: both leave slopes unbiased. Heteroskedasticity distorts standard errors. Imperfect multicollinearity inflates them.

  • Using n − k as degrees of freedom.

    Forgetting the intercept.

    Fix: Use n − k − 1 when k counts slope coefficients only.

  • Confusing perfect and imperfect multicollinearity.

    Both use the same word.

    Fix: Perfect means an exact linear relation and OLS cannot run. Imperfect means high correlation, so estimates exist but are imprecise.

  • Treating the intercept as always meaningful.

    Students assume zero values for all regressors are realistic.

    Fix: Interpret it only if all regressors can sensibly equal zero.

  • Thinking normality of errors is needed for unbiasedness.

    Mixing assumptions used for different results.

    Fix: Unbiasedness needs zero conditional mean and no perfect multicollinearity. Normality is for exact small-sample tests.

Worked examples

Example 1

A regression of a fund's monthly excess return (%) on market excess return (%) and a size factor (%) gives: Ŷ = 0.20 + 1.10 X1 + 0.40 X2. If the market excess return is 3% and the size factor is −2%, what is the predicted excess return? If X1 rises by 1 point with X2 unchanged, what is the change in Ŷ?

Show the solution
  1. Substitute: Ŷ = 0.20 + 1.10 × 3 + 0.40 × (−2).
  2. Compute 1.10 × 3 = 3.30 and 0.40 × (−2) = −0.80.
  3. Sum: 0.20 + 3.30 − 0.80 = 2.70.
  4. For the change, ΔŶ = 1.10 × 1 = 1.10 with X2 held constant.

Answer: Predicted excess return is 2.70%. A 1-point rise in X1 with X2 fixed raises Ŷ by 1.10 points.

Example 2

In a regression with 62 observations and 3 explanatory variables, the estimated coefficient on X2 is 0.84 with a standard error of 0.30. Test H0: β2 = 0 at the 5% level, two-tailed. The critical t value for 58 degrees of freedom is about 2.00. What do you conclude?

Show the solution
  1. Degrees of freedom = n − k − 1 = 62 − 3 − 1 = 58.
  2. t = (0.84 − 0) ÷ 0.30 = 2.80.
  3. Compare |2.80| with the critical value 2.00. 2.80 > 2.00.
  4. So reject H0 at the 5% level.

Answer: t = 2.80 exceeds 2.00, so reject H0. The coefficient on X2 is statistically significant, holding the other regressors constant.

Exam tips

  • Questions often ask which assumption violation biases the coefficients. Only omitted correlated variables (zero conditional mean failure) do so among the usual options.
  • Always compute degrees of freedom as n − k − 1 before looking up a critical value.
  • In interpretation questions, watch units. A coefficient on a variable in percent versus decimals changes the answer by 100.
  • Know the dummy variable trap as the standard example of perfect multicollinearity.

Practice questions from Regression with Multiple Explanatory Variables

Multiple Regression Model and OLS Assumptions in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Multiple Regression Model and OLS Assumptions: frequently asked questions

What are the OLS assumptions in multiple regression?

Linear model, zero conditional mean of errors, independent and identically distributed observations, no large outliers, and no perfect multicollinearity. For BLUE you also need homoskedasticity and no serial correlation. Normal errors are added for exact small-sample tests.

How do I interpret a coefficient in multiple regression?

It is the expected change in Y for a one-unit increase in that variable, holding all other regressors constant. It is a partial effect, so it can differ from the slope in a simple regression.

Does multicollinearity bias OLS estimates?

No. Imperfect multicollinearity leaves estimates unbiased but raises standard errors, so t statistics shrink. Perfect multicollinearity makes estimation impossible.

What is omitted variable bias?

It arises when a left-out variable affects Y and is correlated with an included regressor. The included regressor's coefficient then absorbs part of the omitted effect and is biased.