FRM Exam Part I · Regression Diagnostics
Multiple Regression and OLS Assumptions for FRM Part I
Updated 11 October 2026 · Fact-checked
Multiple regression models a dependent variable as a linear function of two or more regressors plus an error. OLS picks the coefficients that minimise squared residuals. If the classical assumptions hold, the Gauss-Markov theorem says OLS is BLUE: the best linear unbiased estimator. Each slope is the effect of one regressor, holding the others fixed.
Understand Multiple Regression and OLS Assumptions
A multiple linear regression explains a dependent variable Y using several explanatory variables X1, X2, ..., Xk. The model is Y = β0 + β1X1 + ... + βkXk + ε. The error ε holds everything the model does not capture.
OLS (ordinary least squares) chooses the coefficient estimates that minimise the sum of squared residuals (SSR). A residual is the gap between the actual Y and the fitted Y. Squaring stops positive and negative gaps from cancelling and penalises large misses.
The key interpretation point: a slope coefficient βj is the expected change in Y for a one-unit change in Xj, holding all other regressors constant. This is why the slope on X1 usually changes when you add X2. In a simple regression, X1 also picks up the effect of any related omitted variable. The intercept is the expected Y when every regressor equals zero.
The classical assumptions are: (1) the model is linear in the parameters and correctly specified; (2) the errors have zero conditional mean given the regressors, E(ε | X) = 0; (3) the regressors show no perfect multicollinearity (no regressor is an exact linear combination of the others); (4) the errors have constant variance (homoskedasticity); (5) the errors are uncorrelated with each other (no serial correlation). Normality of errors is not needed for Gauss-Markov. It is used for exact t and F tests in small samples.
The Gauss-Markov theorem says that under assumptions (1) to (5), OLS is BLUE: among all linear unbiased estimators, it has the smallest variance. "Best" means minimum variance. "Linear" means the estimator is a linear function of Y. "Unbiased" means its expected value equals the true β. If errors are heteroskedastic or correlated, OLS stays unbiased (given zero conditional mean) but is no longer best, and the usual standard errors are wrong. If the zero-mean assumption fails, for example from an omitted variable correlated with a regressor, OLS is biased.
Key formulas to remember
- Multiple regression model
- Yi = β0 + β1X1i + β2X2i + ... + βkXki + εi
- k regressors plus an intercept, so k + 1 coefficients.
- OLS objective
- minimise SSR = Σ (Yi − Ŷi)² = Σ ei²
- OLS chooses the estimates that give the smallest sum of squared residuals.
- Fitted value and residual
- Ŷi = b0 + b1X1i + ... + bkXki; ei = Yi − Ŷi
- Residuals sum to zero when the model has an intercept.
- Degrees of freedom
- df = n − k − 1
- n observations, k slope coefficients. Used for t-tests and the standard error of regression.
- Standard error of regression
- SER = √(SSR ÷ (n − k − 1))
- Estimate of the standard deviation of the error term.
- Zero conditional mean
- E(ε | X1, ..., Xk) = 0
- Failure causes biased and inconsistent OLS estimates.
- Homoskedasticity
- Var(ε | X) = σ², a constant
- Violation is heteroskedasticity.
- No serial correlation
- Cov(εi, εj) = 0 for i ≠ j
- Violation is autocorrelation, common in time series.
- Gauss-Markov / BLUE
- Under the classical assumptions, OLS has the minimum variance among linear unbiased estimators
- Does not require normal errors.
How to solve Multiple Regression and OLS Assumptions questions
Use this routine for questions on interpreting coefficients, assumptions, or the properties of OLS.
- 1Identify what is asked: coefficient meaning, a forecast, an assumption, or an estimator property.
- 2Write the fitted equation with the given estimates and define each variable and its units.
- 3For interpretation, state each slope as the change in Y per one-unit change in that X, holding the others constant.
- 4For prediction, substitute the given X values and compute Ŷ. Check that units match, such as percent versus decimal.
- 5For assumption questions, match the symptom to the assumption: omitted variable points to zero conditional mean; unequal residual spread points to heteroskedasticity; correlated residuals point to serial correlation; near-duplicate regressors point to multicollinearity.
- 6Decide the consequence: bias, loss of efficiency, or wrong standard errors.
- 7For BLUE, check which of the assumptions hold. If all five hold, OLS is BLUE.
- 8Choose the option that fits, then re-read the question for words such as unbiased, efficient, or consistent.
Quickest way: Assumption-to-consequence shortcut
When to use it: Use when a question describes a data problem and asks what it does to OLS.
- Ask first: does the problem make the estimates biased? Only a violation of zero conditional mean (omitted variable, measurement error in X, endogeneity) does so.
- If the problem is heteroskedasticity or serial correlation, think: still unbiased, no longer minimum variance, standard errors unreliable.
- If the problem is multicollinearity, think: still unbiased and still BLUE, but standard errors are large and t-stats weak.
- If it is perfect multicollinearity, think: OLS cannot be computed.
- For a slope interpretation, always add the phrase holding other variables constant.
Common mistakes in Multiple Regression and OLS Assumptions
Saying heteroskedasticity makes OLS coefficients biased.
Students link any assumption violation with bias.
Fix: Remember that heteroskedasticity and serial correlation leave estimates unbiased but make them inefficient and the standard errors incorrect.
Treating normality of errors as a Gauss-Markov requirement.
Normality appears in the same list of assumptions in many texts.
Fix: Gauss-Markov needs only linearity, zero conditional mean, no perfect multicollinearity, homoskedasticity and no serial correlation. Normality is for exact small-sample inference.
Interpreting a slope without holding other variables constant.
Students carry over the simple regression reading.
Fix: State the ceteris paribus condition every time. The slope is a partial effect.
Using n − k as degrees of freedom.
Confusion over whether k counts the intercept.
Fix: With k slope coefficients plus an intercept, df = n − k − 1.
Claiming that multicollinearity biases coefficients.
The word suggests a serious defect.
Fix: Imperfect multicollinearity inflates standard errors but OLS remains unbiased and BLUE. Only perfect multicollinearity breaks estimation.
Mixing up unbiased with efficient.
Both are described as good properties.
Fix: Unbiased means the average estimate equals the true value. Efficient means smallest variance among the unbiased candidates.
Worked examples
Example 1
A regression of monthly fund excess return (%) on market excess return (%) and a size factor (%) gives: Ŷ = 0.20 + 0.95 × MKT + 0.30 × SIZE. If MKT = 4% and SIZE = −2%, what is the predicted excess return? How do you interpret the 0.30?
Show the solution
- Write the equation: Ŷ = 0.20 + 0.95 × MKT + 0.30 × SIZE.
- Substitute: Ŷ = 0.20 + 0.95 × 4 + 0.30 × (−2).
- Compute: 0.95 × 4 = 3.80; 0.30 × (−2) = −0.60.
- Add: 0.20 + 3.80 − 0.60 = 3.40.
- Interpret 0.30: a one-percentage-point rise in SIZE raises the expected fund excess return by 0.30 percentage points, holding the market factor constant.
Answer: Predicted excess return = 3.40%. The coefficient 0.30 is the partial effect of SIZE, holding MKT constant.
Example 2
An analyst fits a model with n = 60 observations and 4 regressors, and finds SSR = 220.5. (a) What is the standard error of the regression? (b) A colleague notes that the residual variance grows with the size of one regressor but the model has no omitted variables. Which OLS property is affected?
Show the solution
- Degrees of freedom: n − k − 1 = 60 − 4 − 1 = 55.
- SER² = SSR ÷ df = 220.5 ÷ 55 = 4.00909.
- SER = √4.00909 ≈ 2.002.
- For (b): residual variance rising with a regressor is heteroskedasticity.
- With no omitted variables, zero conditional mean still holds, so OLS remains unbiased.
- But the constant variance assumption fails, so Gauss-Markov no longer applies: OLS is not minimum variance and the usual standard errors are unreliable.
Answer: (a) SER ≈ 2.00. (b) OLS remains unbiased but is no longer BLUE, and its standard errors are incorrect.
Exam tips
- Expect conceptual questions that ask which assumption violation causes bias and which only hurts efficiency.
- Memorise that df = n − k − 1 and compute it first in any test-statistic question.
- When a coefficient is asked to be interpreted, include units and the holding-constant phrase to rule out wrong options.
- Read BLUE word by word: best (minimum variance), linear, unbiased, estimator.
- Watch for the distractor that adds normality to Gauss-Markov. It is not required.
Practice questions from Regression Diagnostics
- A regression of bond spread changes on equity returns shows a clear U-shaped pattern in the residuals when plotted against the fitted values…
- The true model is Y = 1.0 + 2.0*X1 + 3.0*X2 + e. An analyst omits X2 and regresses Y on X1 only. In the sample, Cov(X1, X2) = 0.8 and Var(X1…
- In a multiple regression, the auxiliary regression of explanatory variable X1 on the other explanatory variables yields an R-squared of 0.90…
- A regression of daily portfolio returns on a market factor shows that the variance of residuals is much larger in high-volatility periods th…
- A risk analyst's regression of portfolio returns on two highly correlated factors (sample correlation 0.97) is used only to forecast returns…
Multiple Regression and OLS Assumptions in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Multiple Regression and OLS Assumptions: frequently asked questions
What are the OLS assumptions for FRM Part I?
The model is linear in parameters, errors have zero conditional mean, there is no perfect multicollinearity, errors have constant variance, and errors are uncorrelated across observations. Normal errors are added for exact small-sample tests. Know what each one protects.
What does BLUE mean in the Gauss-Markov theorem?
BLUE stands for best linear unbiased estimator. Under the classical assumptions, OLS has the smallest variance among all linear unbiased estimators. It does not claim OLS beats nonlinear or biased estimators.
How do I interpret a multiple regression coefficient?
It is the expected change in the dependent variable for a one-unit increase in that regressor, with all other regressors held constant. Always give the units. The intercept is the expected value of Y when all regressors are zero.
Does multicollinearity violate the Gauss-Markov assumptions?
Only perfect multicollinearity does, because OLS cannot then be computed. High but imperfect multicollinearity leaves OLS unbiased and BLUE. It raises standard errors, which makes individual coefficients look insignificant.