Skip to content

FRM Exam Part I · Linear Regression

Heteroskedasticity, Multicollinearity and Model Misspecification Explained

Updated 11 October 2026 · Fact-checked

These are three ways a regression can break OLS assumptions. Heteroskedasticity means error variance is not constant and standard errors are wrong. Multicollinearity means regressors are highly correlated and standard errors inflate. Omitted variable bias means a missing relevant variable biases coefficients. Detect with plots, tests, VIF and theory; fix with robust errors, dropping variables or adding variables.

Understand Heteroskedasticity, Multicollinearity and Model Misspecification

OLS gives reliable results only when certain assumptions hold. This topic covers three common failures. Each damages a different part of your output, so the exam asks you to match the problem to its effect and its fix.

Heteroskedasticity means the variance of the error term is not constant across observations. A typical case is that residuals get larger as the size of a regressor grows. The coefficient estimates stay unbiased and consistent, but the usual standard errors are wrong. So t-statistics and confidence intervals cannot be trusted. Often the standard errors are understated, which makes variables look more significant than they are. You detect it with a residual plot (a fan or funnel shape) or a formal test such as the Breusch-Pagan test or the White test. You fix it with robust (White) standard errors, or by using weighted least squares.

Multicollinearity means two or more regressors are highly correlated with each other. Perfect multicollinearity (one regressor is an exact linear combination of others) makes OLS impossible to compute. Imperfect multicollinearity still allows estimation, and the estimates stay unbiased. But standard errors become large, so t-statistics are small and individual coefficients look insignificant. The classic sign is a high R² and a significant F-statistic while individual t-statistics are insignificant. You detect it with pairwise correlations and the variance inflation factor (VIF). Fixes include dropping one of the redundant variables, combining them, or collecting more data.

Model misspecification means the model has the wrong form. The key case is omitted variable bias: you leave out a variable that affects Y and is correlated with an included regressor. The included coefficient then absorbs part of the missing variable's effect and becomes biased and inconsistent. If the omitted variable is uncorrelated with the included regressors, there is no bias in the slopes. The opposite error, including irrelevant variables, does not cause bias but raises standard errors. Other misspecification includes wrong functional form, such as fitting a straight line to a curved relationship.

A simple memory aid: heteroskedasticity hurts the standard errors only. Multicollinearity hurts the standard errors of the affected coefficients only. Omitted variables hurt the coefficients themselves.

Key formulas to remember

Variance inflation factor
VIF_j = 1 ÷ (1 − R²_j)
R²_j is the R² from regressing regressor j on all the other regressors. A common rule of thumb flags VIF above 10, but it is a guideline, not a strict rule.
Omitted variable bias (one included, one omitted regressor)
Bias in b1 ≈ β2 × [Cov(X1, X2) ÷ Var(X1)]
β2 is the true effect of the omitted variable. Bias direction is the sign of β2 times the sign of the correlation between X1 and X2. No bias if the correlation is zero.
Breusch-Pagan test statistic
n × R² from regressing squared residuals on the regressors ~ χ² with k degrees of freedom
Null hypothesis: homoskedasticity. A large statistic rejects the null.
Summary of effects
Heteroskedasticity: coefficients unbiased, standard errors wrong. Multicollinearity: coefficients unbiased, standard errors large. Omitted variable: coefficients biased
Use this as a quick matching table in the exam.

How to solve Heteroskedasticity, Multicollinearity and Model Misspecification questions

Use this sequence for any question on violations of OLS assumptions.

  1. 1Identify the symptom in the question: fan-shaped residuals, high R² with insignificant t-statistics, a surprising coefficient sign, or a missing variable.
  2. 2Name the problem: non-constant error variance is heteroskedasticity, correlated regressors is multicollinearity, a missing or wrongly shaped variable is misspecification.
  3. 3State the effect: does it bias the coefficients or only distort the standard errors? Remember that only omitted variables (correlated with included ones) bias the slopes.
  4. 4If a calculation is needed, apply the formula: VIF = 1 ÷ (1 − R²_j), or the bias direction from the signs of β2 and the correlation.
  5. 5Pick the detection method: residual plot, Breusch-Pagan or White test for heteroskedasticity; correlation matrix or VIF for multicollinearity; theory and residual patterns for specification.
  6. 6Pick the fix: robust standard errors, dropping or combining correlated variables, or adding the missing variable or correcting the functional form.
  7. 7Check that your answer matches the exact wording, such as unbiased versus efficient, or inconsistent versus inefficient.

Quickest way: Symptom-to-problem matching

When to use it: Use for conceptual multiple-choice questions where you must name the problem, its effect or its fix.

  1. Funnel-shaped residuals or a test on squared residuals: heteroskedasticity. Fix: robust standard errors.
  2. High R², significant F, insignificant individual t-statistics: multicollinearity. Check VIF.
  3. Coefficient seems biased, or a relevant variable is missing and correlated with a regressor: omitted variable bias. Fix: add the variable.
  4. For bias sign, multiply the sign of the omitted variable's true effect by the sign of its correlation with the included regressor.
  5. For VIF, compute 1 ÷ (1 − R²) directly; R² of 0.90 gives 10.

Common mistakes in Heteroskedasticity, Multicollinearity and Model Misspecification

  • Saying heteroskedasticity makes the OLS coefficients biased.

    Students link any violated assumption with bias.

    Fix: Remember that heteroskedasticity leaves coefficients unbiased and consistent. It makes the usual standard errors unreliable and OLS no longer the most efficient.

  • Saying multicollinearity biases the coefficients.

    Large, unstable estimates look like bias.

    Fix: Imperfect multicollinearity leaves estimates unbiased but raises standard errors. Look for high R² with low t-statistics.

  • Computing VIF as 1 ÷ R² or using the overall regression R².

    The formula is misremembered, and the R² used is the wrong one.

    Fix: Use 1 ÷ (1 − R²_j), where R²_j comes from regressing that regressor on the other regressors, not on Y.

  • Assuming any omitted variable causes bias.

    Students forget the correlation condition.

    Fix: Bias requires the omitted variable to affect Y and be correlated with an included regressor. If uncorrelated, slopes are not biased.

  • Getting the direction of omitted variable bias wrong.

    Students guess instead of combining signs.

    Fix: Multiply the sign of the omitted variable's true coefficient by the sign of its correlation with the included variable. Positive times positive is upward bias.

  • Fixing heteroskedasticity by dropping variables.

    Students confuse it with the multicollinearity remedy.

    Fix: Use robust standard errors or weighted least squares. Dropping variables risks creating omitted variable bias.

Worked examples

Example 1

In a regression of fund returns on four factors, the regression of factor 2 on the other three factors gives R² = 0.84. Compute the VIF for factor 2 and state whether it signals a multicollinearity concern under the common rule of thumb of 10.

Show the solution
  1. Use VIF = 1 ÷ (1 − R²_j).
  2. R²_j = 0.84, so 1 − 0.84 = 0.16.
  3. VIF = 1 ÷ 0.16 = 6.25.
  4. Compare with the rule of thumb of 10: 6.25 is below 10.

Answer: VIF = 6.25. It is below the common threshold of 10, so it does not signal severe multicollinearity by that rule, though it shows moderate inflation of the variance.

Example 2

You regress a bank's loan losses (Y) on unemployment (X1). The true model also includes house price decline (X2), which has a positive effect on losses (β2 > 0) and is positively correlated with unemployment. You omit X2. What is the likely effect on the estimated coefficient on X1?

Show the solution
  1. Omitted variable bias is present because X2 affects Y and is correlated with X1.
  2. Sign of β2 is positive.
  3. Sign of Cov(X1, X2) is positive.
  4. Bias direction = (+) × (+) = positive.
  5. So the estimated coefficient on X1 absorbs part of the effect of X2 and is too high.

Answer: The coefficient on unemployment is biased upward (overstated), and it is inconsistent, so a larger sample does not fix it.

Exam tips

  • Always separate what is affected: coefficients (bias) versus standard errors (inference). Many wrong options swap them.
  • Learn the multicollinearity signature: high R², significant F-test, insignificant t-statistics.
  • For VIF questions, do the arithmetic 1 ÷ (1 − R²) and check the R² came from the regressor-on-regressors regression.
  • For omitted variable bias, write the two signs down before choosing an option.
  • Know that robust standard errors fix heteroskedasticity inference without changing the coefficient estimates.

Practice questions from Linear Regression

Heteroskedasticity, Multicollinearity and Model Misspecification in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Heteroskedasticity, Multicollinearity and Model Misspecification: frequently asked questions

What is the difference between heteroskedasticity and multicollinearity?

Heteroskedasticity is about the errors: their variance is not constant. Multicollinearity is about the regressors: they are highly correlated with each other. Neither biases the coefficients, but both make standard errors unreliable or large.

How do you detect multicollinearity using VIF?

Regress each explanatory variable on the others, take that R², and compute VIF = 1 ÷ (1 − R²). A larger VIF means more variance inflation. A VIF above about 10 is a common warning level.

What are robust standard errors?

They are standard errors recalculated so they stay valid when the error variance is not constant. The coefficient estimates do not change. Only the standard errors, t-statistics and confidence intervals change.

When does omitted variable bias occur?

It occurs when you leave out a variable that affects the dependent variable and is correlated with an included regressor. The included coefficient then picks up part of the missing effect. If the omitted variable is uncorrelated with the included ones, the slopes are not biased.