Skip to content

CFA Level II Exam · Extensions of Multiple Regression

Multicollinearity in Multiple Regression for CFA Level II

Updated 7 October 2026 · Fact-checked

Multicollinearity means two or more independent variables in a regression are highly correlated with each other. Slope estimates stay unbiased, but standard errors rise, t-statistics fall and coefficients become unstable. Detect it with the variance inflation factor, VIF = 1 ÷ (1 − Rj²). Fix it by dropping or combining variables.

Understand Multicollinearity

A multiple regression estimates the effect of each independent variable while holding the others constant. That works only if each variable carries some information the others do not. Multicollinearity occurs when two or more independent variables are highly correlated, or when one variable is close to a linear combination of the others.

When this happens, the model cannot tell which variable is doing the work. The regression as a whole may still fit well and predict well. But the individual coefficients become imprecise. Their standard errors are inflated, so t-statistics shrink and variables that truly matter can look insignificant.

The classic symptom is a high R² and a significant F-statistic, together with insignificant t-statistics on individual slopes. Coefficient signs may also look wrong, and estimates can swing sharply when you add or remove a variable or a few observations.

Multicollinearity is not a violation that biases the coefficients. The estimates remain unbiased and consistent. The damage is to precision and to hypothesis tests on individual coefficients. This is the key contrast with heteroskedasticity and serial correlation, which concern the regression errors, not the relationships between independent variables.

There are two levels of problem. Perfect multicollinearity (an exact linear relationship, such as including all dummy categories plus an intercept) means the regression cannot be estimated. High but imperfect multicollinearity is the usual exam case.

Key formulas to remember

Variance inflation factor
VIFj = 1 ÷ (1 − Rj²)
Rj² comes from regressing independent variable j on all the other independent variables. It is not the R² of the main model.
VIF rules of thumb
VIF > 5: investigate. VIF > 10: serious multicollinearity
These are conventions, not exact tests. A VIF of 1 means no collinearity with the other variables.
Classic symptom pattern
High R², significant F-test, insignificant individual t-statistics
Suggests multicollinearity but does not prove it. Confirm with VIF or pairwise correlations.
Effect on estimates
Slopes: unbiased and consistent. Standard errors: inflated. Type II errors: more likely
You fail to reject a false null more often because t-statistics are too small.

How to solve Multicollinearity questions

Use this sequence for any multicollinearity question in an item set.

  1. 1Find the regression output in the exhibit: R², F-statistic, coefficients, standard errors, t-statistics, and any VIFs or correlation matrix.
  2. 2Compare the overall fit with the individual tests. High R² and a significant F, but insignificant t-statistics, point to multicollinearity.
  3. 3If a VIF is given, apply the rule of thumb (above 5 investigate, above 10 serious). If only Rj² is given, compute VIF = 1 ÷ (1 − Rj²).
  4. 4Check the pairwise correlations if shown, but remember that a low pairwise correlation does not rule out multicollinearity among three or more variables.
  5. 5State the consequence: unbiased coefficients, inflated standard errors, lower t-statistics, higher chance of a Type II error.
  6. 6Choose the remedy: drop one of the correlated variables, combine them into one variable or index, use a different proxy, or use principal components or penalized methods.
  7. 7Answer exactly what is asked. If the question mentions errors being non-constant or correlated over time, it is about heteroskedasticity or serial correlation, not this topic.

Quickest way: Three-check shortcut

When to use it: When a vignette gives regression output and asks whether multicollinearity is present or what it does.

  1. Scan for high R² or significant F next to insignificant t-statistics.
  2. Look for a VIF above 5 (or 10) or a high correlation between two regressors.
  3. Pick the answer that says coefficients are still unbiased but standard errors are too high.
  4. For fixes, choose dropping or combining the correlated variables.

Common mistakes in Multicollinearity

  • Saying multicollinearity biases the coefficient estimates.

    Students link any regression problem to bias.

    Fix: Remember that the estimates stay unbiased and consistent. Only the standard errors are inflated.

  • Using the model's R² in the VIF formula.

    The same symbol R² appears in both places.

    Fix: Use Rj² from regressing variable j on the other independent variables, not on the dependent variable.

  • Confusing multicollinearity with heteroskedasticity.

    Both inflate or distort standard errors and both appear in the same chapter.

    Fix: Multicollinearity is about correlation among independent variables. Heteroskedasticity is about non-constant error variance. Different cause, different test, different fix.

  • Concluding that a significant F-test rules out multicollinearity.

    Students think a good model has no problems.

    Fix: A significant F with insignificant t-statistics is the very symptom of multicollinearity.

  • Treating VIF thresholds as exact tests.

    Rules of thumb are memorized as laws.

    Fix: Say a VIF above 5 warrants investigation and above 10 signals serious multicollinearity. Use the vignette's wording.

  • Relying on pairwise correlations alone.

    A low correlation matrix looks reassuring.

    Fix: Multicollinearity can involve a combination of three or more variables, which is why VIF is the better check.

Worked examples

Example 1

An analyst regresses a fund's excess return on four variables. The model has R² of 0.86 and a significant F-statistic. Three of the four slope t-statistics are below the critical value. The analyst regresses variable 2 on the other three variables and gets R² of 0.90. (1) What problem is suggested? (2) What is the VIF for variable 2? (3) What is the consequence for the slope estimates?

Show the solution
  1. The pattern of high R² and significant F with insignificant t-statistics suggests multicollinearity.
  2. VIF2 = 1 ÷ (1 − 0.90) = 1 ÷ 0.10 = 10.
  3. A VIF of 10 is at the level usually considered serious, and well above 5.
  4. Consequence: slopes remain unbiased, but standard errors are inflated, so t-statistics are understated.

Answer: (1) Multicollinearity. (2) VIF = 10. (3) Estimates are unbiased but imprecise, with inflated standard errors and a higher chance of Type II errors.

Example 2

A vignette shows a regression of a stock's return on market return, sector return and interest rate change. The correlation between market return and sector return is 0.95. The sector return coefficient has a t-statistic of 0.8, while the model R² is 0.74. The analyst proposes dropping sector return. (1) Is multicollinearity likely? (2) If sector return is the variable regressed on the other two and gives Rj² of 0.80, what is its VIF? (3) Is the proposed remedy appropriate?

Show the solution
  1. A 0.95 correlation between two regressors is high, and an insignificant t with a high R² fits the symptom pattern, so multicollinearity is likely.
  2. VIF = 1 ÷ (1 − 0.80) = 1 ÷ 0.20 = 5.
  3. A VIF of 5 is at the threshold to investigate, so it supports a concern but is not extreme.
  4. Dropping one of two highly correlated variables is a standard remedy, provided it is not an important variable on theoretical grounds. Dropping a needed variable could cause omitted variable bias.

Answer: (1) Yes, likely. (2) VIF = 5. (3) Appropriate if sector return is redundant with market return, but the analyst should check that dropping it does not cause omitted variable bias.

Exam tips

  • Questions are set inside a vignette, so find the exhibit values (R², F, t-statistics, VIF) before reading the answer options.
  • Memorize the contrast: multicollinearity leaves estimates unbiased but inflates standard errors, so tests become too conservative.
  • If an exhibit gives Rj², compute VIF = 1 ÷ (1 − Rj²) first. It takes seconds and is a common calculation.
  • When options mention remedies, prefer dropping or combining correlated variables, not robust standard errors, which address heteroskedasticity.
  • Read the diagnosis wording carefully. 'Significant F, insignificant t' points to multicollinearity, while a pattern in residuals points elsewhere.

Multicollinearity: frequently asked questions

What VIF indicates multicollinearity?

A common rule of thumb is that a VIF above 5 needs investigation and a VIF above 10 signals serious multicollinearity. These are conventions, not strict tests. A VIF of 1 means the variable is uncorrelated with the other regressors.

How do you detect and fix multicollinearity?

Detect it from a high R² with insignificant t-statistics, high pairwise correlations, or high VIFs. Fix it by dropping one of the correlated variables, combining variables into one, or using a different proxy. Methods such as principal components can also help.

What is the difference between heteroskedasticity and multicollinearity?

Heteroskedasticity means the variance of the regression errors is not constant. Multicollinearity means independent variables are highly correlated with each other. The first is tested with a Breusch-Pagan test and fixed with robust standard errors. The second is detected with VIF and fixed by changing the variables.

Does multicollinearity make the coefficients biased?

No. Coefficient estimates remain unbiased and consistent. The problem is that standard errors are inflated, so t-statistics are smaller and you may wrongly fail to reject a false null hypothesis.

Can the model still be used for prediction?

Often yes, because overall fit and prediction can remain good. The main problem is interpreting the individual coefficients. Prediction can suffer if the relationship between the correlated variables changes outside the sample.