CFA Level II Exam · Evaluating Regression Model Fit and Interpreting Model Results
Multicollinearity in Regression: Detection and Remedies
Updated 7 October 2026 · Fact-checked
Multicollinearity means two or more independent variables in a multiple regression are highly correlated with each other. It inflates standard errors, so t-statistics fall and coefficients become unreliable, while R-squared stays high. Detect it with VIF (a value above 5 deserves investigation, above 10 is serious) and fix it by dropping or combining variables.
Understand Multicollinearity
In a multiple regression, each slope coefficient measures the effect of one variable holding the others constant. That only works if each variable moves with some independence. Multicollinearity occurs when two or more independent variables are highly correlated, or one is close to a linear combination of others.
Think of two variables that almost always move together, such as a company's sales and its total assets. The model cannot tell which one is driving the dependent variable. It sees the joint effect clearly but cannot split it between the two.
The result is that the coefficient estimates remain unbiased and consistent, but their standard errors are inflated. Larger standard errors give smaller t-statistics, so you fail to reject the null that a coefficient is zero (more Type II errors). Estimates also become unstable: small changes in the data can swing a coefficient a lot, even flip its sign.
The classic symptom is a high R-squared and a significant F-test, but individually insignificant t-statistics. The model as a whole explains the data, but no single variable looks important.
This is different from heteroskedasticity and serial correlation. Those are problems with the error term and usually lead to understated standard errors (and overstated t-statistics) when positive. Multicollinearity is a problem among the independent variables and overstates standard errors. Perfect multicollinearity, where one variable is an exact linear function of others, makes the regression impossible to estimate. Dummy variable traps are a common cause.
Key formulas to remember
- Variance inflation factor
- VIF_j = 1 ÷ (1 − R²_j)
- R²_j comes from regressing independent variable j on all the other independent variables. It is not the R² of the main model.
- VIF rule of thumb
- VIF > 5: investigate; VIF > 10: serious multicollinearity
- These are rules of thumb, not strict tests. VIF of 1 means no correlation with the other variables.
- Effect on standard error
- Higher VIF → larger standard error → smaller t-statistic
- Coefficients stay unbiased and consistent. Only precision is lost.
- Classic symptom
- High R² and significant F-test, but insignificant t-statistics
- Strong sign of multicollinearity, but its absence does not prove there is none.
- Pairwise correlation check
- Large |correlation| between two regressors
- Only catches pairs. With several regressors, multicollinearity can exist even when no pair is highly correlated, so use VIF.
How to solve Multicollinearity questions
Use this sequence for any multicollinearity question in an item set.
- 1Find the regression output in the exhibit: coefficients, standard errors, t-statistics, R-squared, F-statistic, and any VIFs or correlation matrix.
- 2Look for the symptoms: high R-squared or significant F-test combined with insignificant individual t-statistics, or very large standard errors, or unstable signs.
- 3If VIFs are given, compare each to 5 and 10. If an R²_j is given, compute VIF = 1 ÷ (1 − R²_j).
- 4Check the correlation matrix for large correlations between independent variables, remembering pairs alone can miss the problem.
- 5State the effect: coefficients unbiased, standard errors inflated, t-statistics too low, Type II errors more likely.
- 6Choose the remedy: drop one of the correlated variables, combine them (ratio or index), use a different proxy, or increase the sample. Principal components or penalized regression also appear in the machine learning topics.
- 7Confirm the answer does not confuse this with heteroskedasticity or serial correlation, which concern the error term.
Quickest way: Three-second symptom scan
When to use it: When the vignette gives R², F-test and t-statistics but no VIF, and you must decide fast whether multicollinearity is present.
- Ask: is R² high and F significant, yet t-stats weak? If yes, answer multicollinearity.
- If VIF is asked, use 1 ÷ (1 − R²_j). For R²_j of 0.80, VIF is 5; for 0.90, VIF is 10.
- Remember the direction: multicollinearity makes t-stats too small; positive serial correlation or heteroskedasticity usually makes them too large.
- For the remedy, pick drop or combine the redundant variable.
Common mistakes in Multicollinearity
Saying multicollinearity biases the coefficient estimates.
Students link any regression problem to bias.
Fix: Remember estimates stay unbiased and consistent. Only the standard errors are inflated and estimates become unstable.
Using the model's R-squared in the VIF formula.
The same symbol R² appears in both places.
Fix: VIF uses R²_j from regressing one independent variable on the other independent variables. The dependent variable is not involved.
Concluding that a significant F-test rules out multicollinearity.
Students treat a significant F-test as a sign of a healthy model.
Fix: A significant F with insignificant t-stats is exactly the symptom. Multicollinearity does not damage the joint test.
Mixing up multicollinearity and heteroskedasticity.
Both inflate or distort standard errors.
Fix: Multicollinearity concerns correlation between independent variables. Heteroskedasticity concerns non-constant error variance. Tests and remedies differ (VIF versus Breusch-Pagan and robust standard errors).
Relying only on pairwise correlations to rule it out.
A low correlation between each pair looks reassuring.
Fix: A variable can be a linear combination of several others without being highly correlated with any one. Use VIF.
Treating the VIF thresholds as exact cutoffs.
Students memorise 5 and 10 as laws.
Fix: They are rules of thumb. Say 'investigate' above 5 and 'serious' above 10, as the question frames it.
Worked examples
Example 1
An analyst regresses a stock's return on four variables: market return, size, book-to-market and a profitability score. The regression has R² of 0.78 and a significant F-statistic. Only market return has a significant t-statistic. Size and book-to-market have a correlation of 0.93. (1) What problem is suggested? (2) What is its effect on the coefficient estimates? (3) What is a sensible remedy?
Show the solution
- Symptoms: high R², significant F, mostly insignificant t-stats, and a 0.93 correlation between two regressors. This points to multicollinearity.
- Effect: coefficient estimates remain unbiased and consistent, but standard errors are inflated, so t-statistics are too low and Type II errors are more likely.
- Remedy: drop one of size or book-to-market, or combine them into a single measure, then re-estimate.
Answer: (1) Multicollinearity between size and book-to-market. (2) Estimates are unbiased but standard errors are inflated, making t-stats understated. (3) Remove or combine one of the two correlated variables.
Example 2
In a regression with three independent variables, the analyst regresses X1 on X2 and X3 and gets R² of 0.88. The model's own R² is 0.65. (1) Compute the VIF for X1. (2) Interpret it. (3) If the standard error of X1's coefficient is 0.40, would the VIF for X1 be consistent with a standard error that is understated or overstated relative to no collinearity?
Show the solution
- VIF uses the auxiliary R² of 0.88, not the model's 0.65.
- VIF = 1 ÷ (1 − 0.88) = 1 ÷ 0.12 = 8.33.
- Interpretation: above 5, so it warrants investigation, but below 10, so not yet in the serious range by the rule of thumb.
- Multicollinearity inflates the variance of the coefficient estimate, so the standard error is larger than it would be without collinearity. It is overstated relative to the uncorrelated case, which lowers the t-statistic.
Answer: (1) VIF = 8.33. (2) Moderately high: investigate, but under 10. (3) The standard error is inflated (overstated), so the t-statistic is too low.
Exam tips
- Questions usually give symptoms in the exhibit, not the word multicollinearity. Train yourself to spot high R² with weak t-stats.
- If an exhibit lists VIFs, compare every variable. The question may ask which variable is most affected.
- Be ready to rank effects: multicollinearity inflates standard errors, while positive serial correlation and conditional heteroskedasticity typically understate them.
- For remedy questions, drop or combine variables is the safest answer. Do not pick adding more variables or using robust standard errors, which address other problems.
- There is no penalty for wrong answers, so never leave an option blank.
Multicollinearity: frequently asked questions
How do you detect multicollinearity in CFA Level II?
Look for a high R² with a significant F-test but insignificant t-statistics. Compute or read VIF values, where above 5 needs investigation and above 10 is serious. A correlation matrix can help but misses cases involving several variables.
What does multicollinearity do to regression coefficients?
Coefficients stay unbiased and consistent, but their standard errors increase. That lowers t-statistics, makes variables look insignificant, and makes estimates unstable across samples.
What is the difference between multicollinearity and heteroskedasticity?
Multicollinearity is high correlation among independent variables. Heteroskedasticity is non-constant variance of the error term. Multicollinearity inflates standard errors, while conditional heteroskedasticity usually understates them and is fixed with robust standard errors.
How do you fix multicollinearity?
Drop one of the highly correlated variables, combine them into one variable such as a ratio or index, or use a different proxy. A larger sample can also help. The aim is to remove redundant information.