FRM Exam Part I · Regression with Multiple Explanatory Variables
Multicollinearity in Multiple Regression: Detection and Fixes
Updated 11 October 2026 · Fact-checked
Multicollinearity means two or more explanatory variables in a regression are highly correlated. Perfect multicollinearity makes OLS impossible. Imperfect multicollinearity keeps OLS unbiased but inflates standard errors, so t-statistics fall. Detect it with a high R² but insignificant t-stats, high pairwise correlations, or a high VIF. Fix it by dropping or combining variables.
Understand Multicollinearity
In a multiple regression, each slope coefficient measures the effect of one variable while holding the others constant. That only works if each variable has some movement of its own. When two regressors move together, OLS cannot tell which one is driving the dependent variable.
Perfect multicollinearity means one regressor is an exact linear combination of the others. Example: including both X₂ and X₃ = 2 × X₂, or including a dummy for every category plus the intercept (the dummy variable trap). OLS cannot compute a unique solution. Software drops a variable or returns an error.
Imperfect multicollinearity means regressors are highly, but not exactly, correlated. OLS still runs. Estimators remain unbiased and consistent, and the model's forecasts can still be fine. The cost is precision: the variances and standard errors of the affected coefficients become large.
Large standard errors cause three symptoms. The t-statistics are small, so you fail to reject H₀ that a coefficient is zero. Confidence intervals are wide. Coefficients can be unstable, changing sign or size when you add or remove a variable or a few observations. A classic sign is a high R² and a significant F-test, while individual t-tests are insignificant.
To fix it, you can drop one of the redundant variables, combine correlated variables into one (for example an average or an index), use principal components, or collect more data. Dropping a variable that truly belongs in the model risks omitted variable bias, so you trade one problem for another. If your goal is prediction rather than interpreting individual coefficients, you may leave it alone.
Key formulas to remember
- Variance inflation factor
- VIFⱼ = 1 ÷ (1 − Rⱼ²)
- Rⱼ² comes from regressing Xⱼ on all the other regressors. VIF = 1 means no collinearity. A common rule of thumb flags VIF above 10 (some use 5). It is a guideline, not a strict test.
- Variance of a slope coefficient (two regressors)
- Var(β̂₁) = σ² ÷ [Σ(X₁ − X̄₁)² × (1 − r₁₂²)]
- As the correlation r₁₂ between regressors approaches 1, the variance grows without bound.
- Standard error inflation
- SE with collinearity = SE without × √VIF
- If VIF = 9, the standard error is 3 times larger than it would be with uncorrelated regressors.
- t-statistic
- t = (β̂ − β₀) ÷ SE(β̂)
- Larger SE gives smaller t, so true effects look insignificant.
- Perfect multicollinearity condition
- Rⱼ² = 1, so VIFⱼ is undefined (infinite)
- One regressor is an exact linear combination of others. OLS has no unique solution.
How to solve Multicollinearity questions
Use this sequence for any multicollinearity question, whether it is conceptual or numerical.
- 1Identify whether the relationship among regressors is exact (perfect) or high but not exact (imperfect).
- 2If perfect, state that OLS cannot produce unique estimates. Look for the cause, such as a duplicated variable or the dummy variable trap.
- 3If imperfect, recall that coefficients stay unbiased and consistent but standard errors are inflated.
- 4For a numerical question, compute VIF = 1 ÷ (1 − R²) using the R² from the auxiliary regression of that regressor on the others.
- 5Convert to the effect on precision: SE rises by √VIF; t-statistic falls by the same factor if the coefficient is unchanged.
- 6Check the symptoms in the question: high R², significant F, insignificant t-stats, unstable coefficients.
- 7Choose the remedy: drop or combine variables, use principal components, or add data. Note the omitted variable bias risk of dropping.
- 8Check that your answer does not claim bias, or invalid forecasts, from imperfect multicollinearity.
Quickest way: Symptom check and VIF shortcut
When to use it: Use when the question lists regression output or asks which statement is true about multicollinearity.
- High R² or significant F but insignificant individual t-stats points to multicollinearity.
- Imperfect multicollinearity: unbiased coefficients, larger standard errors. Eliminate any option saying bias or inconsistency.
- For VIF, compute 1 ÷ (1 − R²) directly. R² = 0.90 gives 10; R² = 0.75 gives 4; R² = 0.80 gives 5.
- For the SE effect, take √VIF.
- If a pairwise correlation is mentioned, remember that low pairwise correlations do not rule out multicollinearity involving three or more variables.
Common mistakes in Multicollinearity
Saying imperfect multicollinearity biases the OLS coefficients.
Students link unstable estimates with wrong estimates.
Fix: Remember OLS stays unbiased and consistent. Only the variances, and so standard errors, increase.
Thinking multicollinearity lowers the model's R² or ruins its fit.
It is confused with a poor model.
Fix: R² is often high. The problem is separating the individual effects, not explaining the dependent variable.
Using only pairwise correlations to rule it out.
Correlation matrices are easy to read.
Fix: One regressor can be a combination of several others without any large pairwise correlation. Use VIF to catch this.
Computing VIF with the R² of the main regression.
Both are called R².
Fix: Use the R² from regressing the regressor in question on the other regressors, not on Y.
Treating a VIF cutoff of 10 as a formal test.
Rules of thumb are memorised as laws.
Fix: Call it a guideline. Some practitioners use 5. Always use judgement.
Recommending dropping a variable without caveat.
It looks like the easiest fix.
Fix: Dropping a relevant variable can cause omitted variable bias. Mention this trade-off.
Worked examples
Example 1
In a regression of a fund's return on three factors, the auxiliary regression of factor 2 on factors 1 and 3 gives R² = 0.92. Compute the VIF for factor 2 and the factor by which its standard error is inflated.
Show the solution
- VIF = 1 ÷ (1 − R²) = 1 ÷ (1 − 0.92).
- 1 − 0.92 = 0.08.
- VIF = 1 ÷ 0.08 = 12.5.
- SE inflation = √12.5 ≈ 3.54.
Answer: VIF = 12.5, so the standard error is about 3.54 times larger than without collinearity. This is above the common cutoff of 10, indicating serious multicollinearity.
Example 2
A multiple regression has R² = 0.88 and a highly significant F-statistic, but each of the two slope coefficients has a t-statistic below 1.0. The correlation between the two regressors is 0.97. Which statement is most accurate: (A) the coefficients are biased; (B) the model has imperfect multicollinearity, so coefficients are unbiased but standard errors are inflated; (C) the regressors are perfectly collinear; (D) the F-test is invalid?
Show the solution
- Significant F with high R² shows the regressors jointly explain Y.
- Insignificant individual t-stats with a 0.97 correlation is the classic symptom of multicollinearity.
- The correlation is high but below 1, so the collinearity is imperfect. OLS runs.
- Imperfect multicollinearity does not cause bias, so A is wrong. C is wrong because correlation is not 1. D is wrong because the joint F-test remains valid.
Answer: B. Imperfect multicollinearity: unbiased coefficients with inflated standard errors.
Exam tips
- Questions often test the effect: pick unbiased coefficients, larger standard errors, lower t-stats, wider confidence intervals.
- Watch for the high R² with insignificant t-stats clue. It nearly always signals multicollinearity.
- Know perfect versus imperfect cleanly. Perfect means OLS fails; imperfect means OLS works but is imprecise.
- VIF calculations are quick. Learn that R² of 0.9 gives 10 and R² of 0.8 gives 5.
- Link the dummy variable trap to perfect multicollinearity when a question includes all categories plus an intercept.
Practice questions from Regression with Multiple Explanatory Variables
- The true model is Y = 2 + 1.5·X1 + 0.8·X2 + e. An analyst omits X2 and regresses Y on X1 only. Sample data show Cov(X1,X2)=0.6 and Var(X1)=2…
- The true model is Y = 1.0 + 2.0·X1 + 3.0·X2 + e. An analyst omits X2 and regresses Y on X1 only. In the sample, the regression of X2 on X1 g…
- A model of portfolio returns omits a relevant variable X2 that is positively correlated with the included variable X1, and X2 has a positive…
- In a multiple regression, a researcher finds that two explanatory variables have a sample correlation of 0.98, the overall F-test is highly …
- Which statement about the OLS estimators in a multiple regression is correct when the classical assumptions hold except that the error varia…
Multicollinearity: frequently asked questions
What is perfect multicollinearity?
It occurs when one explanatory variable is an exact linear combination of others, such as X₃ = 2X₂. OLS cannot compute unique coefficients. A common cause is including a dummy for every category along with the intercept.
How do you detect multicollinearity using VIF?
Regress each explanatory variable on all the others and take that regression's R². Then VIF = 1 ÷ (1 − R²). A VIF above 10 is a common warning sign, and some analysts use 5 as the cutoff.
Does multicollinearity bias regression coefficients?
No, not when it is imperfect. OLS estimators remain unbiased and consistent. The standard errors increase, which makes the estimates imprecise and the t-tests less powerful.
How can you fix multicollinearity?
You can drop one of the correlated variables, combine them into a single variable or index, use principal components, or obtain more data. Be careful, since dropping a relevant variable can cause omitted variable bias.