FRM Exam Part I · Regression with Multiple Explanatory Variables
Heteroskedasticity and Omitted Variable Bias in Regression
Updated 11 October 2026 · Fact-checked
Heteroskedasticity means the regression error variance changes with the explanatory variables. OLS slopes stay unbiased, but standard errors are wrong, so you use robust standard errors. Omitted variable bias arises when you drop a variable that affects Y and is correlated with an included regressor. The bias is β2 × Cov(X1, X2) ÷ Var(X1).
Understand Heteroskedasticity and Omitted Variable Bias
OLS assumes the error term has the same variance for every observation. This is homoskedasticity: Var(ε | X) = σ². When the error variance changes with the regressors, you have heteroskedasticity. A common case in finance is a model where the scatter of residuals widens as firm size or market volatility rises.
Heteroskedasticity does not bias the OLS coefficients. They stay unbiased and consistent, provided the other assumptions hold. What breaks is the standard error. The usual formula is wrong, so t-statistics, p-values and confidence intervals cannot be trusted. OLS is also no longer BLUE, because it is not the most efficient estimator. The usual fix is robust (White) standard errors. Another fix is weighted least squares, which is a form of GLS.
GARP distinguishes two types. Conditional heteroskedasticity means the error variance is related to the level of the regressors. This is the problem case for inference. Unconditional heteroskedasticity means the variance changes over time or across observations but is not related to the regressors. It does not cause major problems for inference.
You detect it by plotting residuals against the regressors, or with a formal test. The Breusch-Pagan test regresses the squared residuals on the regressors. The White test also adds squares and cross-products of the regressors. In both, the null hypothesis is homoskedasticity.
Omitted variable bias (OVB) is a different problem and a more serious one. It appears when you leave out a variable that (1) truly affects Y and (2) is correlated with an included regressor. The included regressor then picks up part of the omitted variable's effect. The coefficient is biased, and the bias does not shrink as the sample grows, so it is inconsistent. If the omitted variable is uncorrelated with your regressors, the slopes are not biased. Only the fit and the error variance suffer.
The reverse mistake is including an irrelevant variable. Coefficients stay unbiased, but their variances increase, especially if the extra variable is correlated with the others. This reduces precision. It is a milder error than omitting a relevant variable.
Key formulas to remember
- Homoskedasticity assumption
- Var(ε | X1, ..., Xk) = σ² (constant)
- Heteroskedasticity means this variance depends on X.
- Breusch-Pagan LM statistic
- LM = n × R² ~ χ²(k)
- R² comes from regressing squared residuals on the k regressors. Null: homoskedasticity. A large LM rejects the null.
- White test
- LM = n × R² ~ χ²(q)
- The auxiliary regression uses regressors, their squares and cross-products. q is the number of auxiliary regressors, excluding the intercept.
- Omitted variable bias (two-variable case)
- Bias in β1-hat = β2 × Cov(X1, X2) ÷ Var(X1)
- True model Y = β0 + β1X1 + β2X2 + ε, with X2 omitted. The estimate converges to β1 plus this term. The sign is the sign of β2 times the sign of the correlation.
- Consequences of heteroskedasticity
- Coefficients: unbiased and consistent. Standard errors: invalid. OLS: not BLUE.
- The standard error is not valid, so t-tests and confidence intervals are unreliable until you correct them.
- Remedies
- Robust (White) standard errors, or WLS/GLS
- Robust standard errors keep the OLS coefficients and change only the standard errors.
- Irrelevant variable
- Coefficients unbiased, variances larger
- Adjusted R² penalises useless regressors.
How to solve Heteroskedasticity and Omitted Variable Bias questions
Use this sequence for any question on non-constant variance or omitted variables.
- 1Identify the problem. Is the issue about the error variance (heteroskedasticity) or about a missing or extra regressor (specification)?
- 2For heteroskedasticity, state the effect. Coefficients stay unbiased and consistent, standard errors are invalid, and OLS is not BLUE.
- 3If a test is asked for, set H0 as homoskedasticity. Compute LM = n × R² from the auxiliary regression of squared residuals.
- 4Compare LM with the χ² critical value, using degrees of freedom equal to the number of auxiliary regressors. Reject H0 if LM exceeds it.
- 5If heteroskedasticity is found, choose the remedy. Use robust standard errors or WLS. Do not change the coefficients just because the variance is non-constant.
- 6For omitted variables, check both conditions. Does the omitted variable affect Y, and is it correlated with an included regressor? If either fails, there is no bias in the slope.
- 7Compute the bias as β2 × Cov(X1, X2) ÷ Var(X1). Add it to β1 for the large-sample value of the estimate. Check the sign by multiplying the signs of β2 and the correlation.
- 8For extra variables, conclude: no bias, but less precision.
Quickest way: Two-question shortcut
When to use it: Use it when an MCQ asks what is biased, what is wrong, or what a test concludes.
- Ask: is the problem in the variance of errors or in the regressors chosen? Variance problem means only standard errors are affected. Regressor problem means coefficients can be biased.
- For a test, compute n × R² and compare with the 5% χ² value. Useful values: df 1 is 3.841, df 2 is 5.991, df 3 is 7.815, df 4 is 9.488.
- For OVB direction, multiply the two signs. Same sign means upward bias, opposite signs mean downward bias.
- For the size of bias, compute Cov ÷ Var first, then multiply by β2. A calculator is enough.
Common mistakes in Heteroskedasticity and Omitted Variable Bias
Saying heteroskedasticity makes the OLS coefficients biased.
Students link any violated assumption to bias.
Fix: Remember that only the standard errors and efficiency are damaged. Coefficients stay unbiased and consistent. Bias comes from omitted variables or correlation between the error and the regressors.
Using the wrong null in Breusch-Pagan or White.
Students read a large LM as supporting the model.
Fix: The null is homoskedasticity. A large LM, above the critical value, rejects it and signals heteroskedasticity.
Assuming any omitted variable causes bias.
The word omitted sounds like automatic harm.
Fix: Bias needs two conditions: the omitted variable affects Y, and it is correlated with an included regressor. If either one is missing, the slope is not biased.
Getting the sign of the OVB wrong.
Students forget to use the sign of both β2 and the correlation.
Fix: Bias = β2 × Cov(X1, X2) ÷ Var(X1). Positive times positive, or negative times negative, gives upward bias. Mixed signs give downward bias.
Treating conditional and unconditional heteroskedasticity as equally serious.
Both are described as non-constant variance.
Fix: Only conditional heteroskedasticity, where variance is related to the regressors, causes the main inference problems. Unconditional does not.
Thinking that adding more variables is always safe.
Students know omission is serious and overcorrect.
Fix: Irrelevant variables do not bias the coefficients, but they raise variances and reduce precision. Adjusted R² penalises them.
Worked examples
Example 1
A regression of fund returns on 3 explanatory variables uses 200 observations. You regress the squared residuals on the same 3 variables and get R² = 0.06. Run a Breusch-Pagan test at the 5% level. The χ² critical value with 3 degrees of freedom is 7.815. What do you conclude, and what should you do?
Show the solution
- H0: homoskedasticity. H1: the error variance depends on the regressors.
- LM = n × R² = 200 × 0.06 = 12.0.
- Degrees of freedom = number of regressors in the auxiliary regression = 3. Critical value = 7.815.
- 12.0 > 7.815, so reject H0.
- The error variance is related to the regressors. The coefficients are still unbiased but the usual standard errors are invalid. Re-estimate with robust standard errors or use WLS.
Answer: LM = 12.0 exceeds 7.815. Reject homoskedasticity. Use heteroskedasticity-robust standard errors for the t-tests.
Example 2
The true model is Y = β0 + 0.80 X1 + 0.50 X2 + ε. An analyst omits X2 and regresses Y on X1 only. Var(X1) = 4 and Cov(X1, X2) = 1.2. In a very large sample, the estimated slope on X1 will be approximately: A) 0.65 B) 0.80 C) 0.95 D) 1.30
Show the solution
- Both conditions for bias hold. X2 affects Y, since β2 = 0.50, and X2 is correlated with X1.
- Compute the slope of X2 on X1: Cov(X1, X2) ÷ Var(X1) = 1.2 ÷ 4 = 0.30.
- Bias = β2 × 0.30 = 0.50 × 0.30 = 0.15.
- The sign is positive, because β2 and the covariance are both positive.
- Large-sample estimate = 0.80 + 0.15 = 0.95.
Answer: C) 0.95. The slope is biased upward by 0.15 and the bias does not disappear as the sample grows.
Exam tips
- Know the exact consequences list for heteroskedasticity: coefficients unbiased and consistent, standard errors invalid, OLS not BLUE. Many MCQs offer a wrong mix of these.
- When a question gives R² from an auxiliary regression and n, compute n × R² and compare with a χ² value. Do not use the original model R².
- For OVB questions, check the two conditions before calculating. A trap option often says no bias when the omitted variable is uncorrelated with the included ones.
- Do not confuse OVB with multicollinearity. OVB biases coefficients. Multicollinearity inflates standard errors and leaves coefficients unbiased.
- Remember that robust standard errors change only the standard errors, not the coefficient estimates.
Practice questions from Regression with Multiple Explanatory Variables
- Which statement about the OLS estimators in a multiple regression is correct when the classical assumptions hold except that the error varia…
- An analyst omits a relevant explanatory variable that is positively correlated with an included variable, and whose true coefficient is posi…
- Model A has k = 2 regressors and R-squared of 0.500. Model B adds 3 more regressors (k = 5) and has R-squared of 0.530. Both use n = 31 obse…
- An analyst regresses a firm's monthly excess return (Y) on the market excess return (X1) and omits a relevant variable X2 that is positively…
- An analyst regresses monthly fund returns on the market excess return and a dummy variable D equal to 1 for months in a recession and 0 othe…
Heteroskedasticity and Omitted Variable Bias: frequently asked questions
What is the difference between heteroskedasticity and homoskedasticity?
Under homoskedasticity the error variance is the same for all observations. Under heteroskedasticity it changes, often with the level of a regressor. The second case leaves the OLS coefficients unbiased but makes the usual standard errors invalid.
How do you detect heteroskedasticity with the Breusch-Pagan test?
Estimate the model, then regress the squared residuals on the regressors. Compute LM = n × R² from that auxiliary regression and compare it with a χ² critical value, with degrees of freedom equal to the number of regressors. A large LM rejects the null of homoskedasticity.
What is conditional versus unconditional heteroskedasticity?
Conditional heteroskedasticity means the error variance is related to the regressors, and it is the case that causes problems for inference. Unconditional heteroskedasticity means the variance changes but is not related to the regressors, and it does not cause major problems.
When does omitting a variable cause bias?
Bias occurs only if the omitted variable affects the dependent variable and is correlated with an included regressor. In the two-variable case the bias is β2 × Cov(X1, X2) ÷ Var(X1). If the omitted variable is uncorrelated with the included regressors, the slopes are not biased.
Is it a problem to include an irrelevant variable?
It does not bias the coefficients. It does increase the variance of the estimates, so they are less precise. This is usually a smaller problem than omitting a relevant variable.