Skip to content

FRM Exam Part I · Regression with Multiple Explanatory Variables

Heteroskedasticity and Omitted Variable Bias in Regression

Updated 11 October 2026 · Fact-checked

Heteroskedasticity means the regression error variance changes with the explanatory variables. OLS slopes stay unbiased, but standard errors are wrong, so you use robust standard errors. Omitted variable bias arises when you drop a variable that affects Y and is correlated with an included regressor. The bias is β2 × Cov(X1, X2) ÷ Var(X1).

Understand Heteroskedasticity and Omitted Variable Bias

OLS assumes the error term has the same variance for every observation. This is homoskedasticity: Var(ε | X) = σ². When the error variance changes with the regressors, you have heteroskedasticity. A common case in finance is a model where the scatter of residuals widens as firm size or market volatility rises.

Heteroskedasticity does not bias the OLS coefficients. They stay unbiased and consistent, provided the other assumptions hold. What breaks is the standard error. The usual formula is wrong, so t-statistics, p-values and confidence intervals cannot be trusted. OLS is also no longer BLUE, because it is not the most efficient estimator. The usual fix is robust (White) standard errors. Another fix is weighted least squares, which is a form of GLS.

GARP distinguishes two types. Conditional heteroskedasticity means the error variance is related to the level of the regressors. This is the problem case for inference. Unconditional heteroskedasticity means the variance changes over time or across observations but is not related to the regressors. It does not cause major problems for inference.

You detect it by plotting residuals against the regressors, or with a formal test. The Breusch-Pagan test regresses the squared residuals on the regressors. The White test also adds squares and cross-products of the regressors. In both, the null hypothesis is homoskedasticity.

Omitted variable bias (OVB) is a different problem and a more serious one. It appears when you leave out a variable that (1) truly affects Y and (2) is correlated with an included regressor. The included regressor then picks up part of the omitted variable's effect. The coefficient is biased, and the bias does not shrink as the sample grows, so it is inconsistent. If the omitted variable is uncorrelated with your regressors, the slopes are not biased. Only the fit and the error variance suffer.

The reverse mistake is including an irrelevant variable. Coefficients stay unbiased, but their variances increase, especially if the extra variable is correlated with the others. This reduces precision. It is a milder error than omitting a relevant variable.

Key formulas to remember

Homoskedasticity assumption
Var(ε | X1, ..., Xk) = σ² (constant)
Heteroskedasticity means this variance depends on X.
Breusch-Pagan LM statistic
LM = n × R² ~ χ²(k)
R² comes from regressing squared residuals on the k regressors. Null: homoskedasticity. A large LM rejects the null.
White test
LM = n × R² ~ χ²(q)
The auxiliary regression uses regressors, their squares and cross-products. q is the number of auxiliary regressors, excluding the intercept.
Omitted variable bias (two-variable case)
Bias in β1-hat = β2 × Cov(X1, X2) ÷ Var(X1)
True model Y = β0 + β1X1 + β2X2 + ε, with X2 omitted. The estimate converges to β1 plus this term. The sign is the sign of β2 times the sign of the correlation.
Consequences of heteroskedasticity
Coefficients: unbiased and consistent. Standard errors: invalid. OLS: not BLUE.
The standard error is not valid, so t-tests and confidence intervals are unreliable until you correct them.
Remedies
Robust (White) standard errors, or WLS/GLS
Robust standard errors keep the OLS coefficients and change only the standard errors.
Irrelevant variable
Coefficients unbiased, variances larger
Adjusted R² penalises useless regressors.

How to solve Heteroskedasticity and Omitted Variable Bias questions

Use this sequence for any question on non-constant variance or omitted variables.

  1. 1Identify the problem. Is the issue about the error variance (heteroskedasticity) or about a missing or extra regressor (specification)?
  2. 2For heteroskedasticity, state the effect. Coefficients stay unbiased and consistent, standard errors are invalid, and OLS is not BLUE.
  3. 3If a test is asked for, set H0 as homoskedasticity. Compute LM = n × R² from the auxiliary regression of squared residuals.
  4. 4Compare LM with the χ² critical value, using degrees of freedom equal to the number of auxiliary regressors. Reject H0 if LM exceeds it.
  5. 5If heteroskedasticity is found, choose the remedy. Use robust standard errors or WLS. Do not change the coefficients just because the variance is non-constant.
  6. 6For omitted variables, check both conditions. Does the omitted variable affect Y, and is it correlated with an included regressor? If either fails, there is no bias in the slope.
  7. 7Compute the bias as β2 × Cov(X1, X2) ÷ Var(X1). Add it to β1 for the large-sample value of the estimate. Check the sign by multiplying the signs of β2 and the correlation.
  8. 8For extra variables, conclude: no bias, but less precision.

Quickest way: Two-question shortcut

When to use it: Use it when an MCQ asks what is biased, what is wrong, or what a test concludes.

  1. Ask: is the problem in the variance of errors or in the regressors chosen? Variance problem means only standard errors are affected. Regressor problem means coefficients can be biased.
  2. For a test, compute n × R² and compare with the 5% χ² value. Useful values: df 1 is 3.841, df 2 is 5.991, df 3 is 7.815, df 4 is 9.488.
  3. For OVB direction, multiply the two signs. Same sign means upward bias, opposite signs mean downward bias.
  4. For the size of bias, compute Cov ÷ Var first, then multiply by β2. A calculator is enough.

Common mistakes in Heteroskedasticity and Omitted Variable Bias

  • Saying heteroskedasticity makes the OLS coefficients biased.

    Students link any violated assumption to bias.

    Fix: Remember that only the standard errors and efficiency are damaged. Coefficients stay unbiased and consistent. Bias comes from omitted variables or correlation between the error and the regressors.

  • Using the wrong null in Breusch-Pagan or White.

    Students read a large LM as supporting the model.

    Fix: The null is homoskedasticity. A large LM, above the critical value, rejects it and signals heteroskedasticity.

  • Assuming any omitted variable causes bias.

    The word omitted sounds like automatic harm.

    Fix: Bias needs two conditions: the omitted variable affects Y, and it is correlated with an included regressor. If either one is missing, the slope is not biased.

  • Getting the sign of the OVB wrong.

    Students forget to use the sign of both β2 and the correlation.

    Fix: Bias = β2 × Cov(X1, X2) ÷ Var(X1). Positive times positive, or negative times negative, gives upward bias. Mixed signs give downward bias.

  • Treating conditional and unconditional heteroskedasticity as equally serious.

    Both are described as non-constant variance.

    Fix: Only conditional heteroskedasticity, where variance is related to the regressors, causes the main inference problems. Unconditional does not.

  • Thinking that adding more variables is always safe.

    Students know omission is serious and overcorrect.

    Fix: Irrelevant variables do not bias the coefficients, but they raise variances and reduce precision. Adjusted R² penalises them.

Worked examples

Example 1

A regression of fund returns on 3 explanatory variables uses 200 observations. You regress the squared residuals on the same 3 variables and get R² = 0.06. Run a Breusch-Pagan test at the 5% level. The χ² critical value with 3 degrees of freedom is 7.815. What do you conclude, and what should you do?

Show the solution
  1. H0: homoskedasticity. H1: the error variance depends on the regressors.
  2. LM = n × R² = 200 × 0.06 = 12.0.
  3. Degrees of freedom = number of regressors in the auxiliary regression = 3. Critical value = 7.815.
  4. 12.0 > 7.815, so reject H0.
  5. The error variance is related to the regressors. The coefficients are still unbiased but the usual standard errors are invalid. Re-estimate with robust standard errors or use WLS.

Answer: LM = 12.0 exceeds 7.815. Reject homoskedasticity. Use heteroskedasticity-robust standard errors for the t-tests.

Example 2

The true model is Y = β0 + 0.80 X1 + 0.50 X2 + ε. An analyst omits X2 and regresses Y on X1 only. Var(X1) = 4 and Cov(X1, X2) = 1.2. In a very large sample, the estimated slope on X1 will be approximately: A) 0.65 B) 0.80 C) 0.95 D) 1.30

Show the solution
  1. Both conditions for bias hold. X2 affects Y, since β2 = 0.50, and X2 is correlated with X1.
  2. Compute the slope of X2 on X1: Cov(X1, X2) ÷ Var(X1) = 1.2 ÷ 4 = 0.30.
  3. Bias = β2 × 0.30 = 0.50 × 0.30 = 0.15.
  4. The sign is positive, because β2 and the covariance are both positive.
  5. Large-sample estimate = 0.80 + 0.15 = 0.95.

Answer: C) 0.95. The slope is biased upward by 0.15 and the bias does not disappear as the sample grows.

Exam tips

  • Know the exact consequences list for heteroskedasticity: coefficients unbiased and consistent, standard errors invalid, OLS not BLUE. Many MCQs offer a wrong mix of these.
  • When a question gives R² from an auxiliary regression and n, compute n × R² and compare with a χ² value. Do not use the original model R².
  • For OVB questions, check the two conditions before calculating. A trap option often says no bias when the omitted variable is uncorrelated with the included ones.
  • Do not confuse OVB with multicollinearity. OVB biases coefficients. Multicollinearity inflates standard errors and leaves coefficients unbiased.
  • Remember that robust standard errors change only the standard errors, not the coefficient estimates.

Practice questions from Regression with Multiple Explanatory Variables

Heteroskedasticity and Omitted Variable Bias: frequently asked questions

What is the difference between heteroskedasticity and homoskedasticity?

Under homoskedasticity the error variance is the same for all observations. Under heteroskedasticity it changes, often with the level of a regressor. The second case leaves the OLS coefficients unbiased but makes the usual standard errors invalid.

How do you detect heteroskedasticity with the Breusch-Pagan test?

Estimate the model, then regress the squared residuals on the regressors. Compute LM = n × R² from that auxiliary regression and compare it with a χ² critical value, with degrees of freedom equal to the number of regressors. A large LM rejects the null of homoskedasticity.

What is conditional versus unconditional heteroskedasticity?

Conditional heteroskedasticity means the error variance is related to the regressors, and it is the case that causes problems for inference. Unconditional heteroskedasticity means the variance changes but is not related to the regressors, and it does not cause major problems.

When does omitting a variable cause bias?

Bias occurs only if the omitted variable affects the dependent variable and is correlated with an included regressor. In the two-variable case the bias is β2 × Cov(X1, X2) ÷ Var(X1). If the omitted variable is uncorrelated with the included regressors, the slopes are not biased.

Is it a problem to include an irrelevant variable?

It does not bias the coefficients. It does increase the variance of the estimates, so they are less precise. This is usually a smaller problem than omitting a relevant variable.