CFA Level II · CFA Level II Exam
Extensions of Multiple Regression for CFA Level II
Extensions of multiple regression covers what you do after the basic model is built: check for influential points, add dummy variables, detect and fix violations such as heteroskedasticity, serial correlation and multicollinearity, spot misspecification, and model yes/no outcomes with logit or probit. You solve questions by reading the exhibit and applying the right test.
What this chapter covers
This chapter builds on basic multiple regression. There, you learn to read coefficients, t-statistics, F-tests and R². Here, you ask whether the model can be trusted. Real data break the assumptions. The chapter teaches you to diagnose each problem, name its effect on coefficients and standard errors, and choose a remedy.
The topics fall into three groups. First, data and inputs: outliers and leverage, and dummy variables. Second, assumption violations: heteroskedasticity, serial correlation and multicollinearity, plus model misspecification. Third, a different type of model: qualitative dependent variables, where the outcome is a category such as default or no default.
The chapter links to the rest of the paper. Regression output appears in item sets on equities, fixed income, corporate finance and portfolio construction, as well as in Quantitative Methods. Level II item sets give you an exhibit of regression results. You must interpret it and apply a model, not recite a definition. Time-series topics later in Quantitative Methods reuse the serial correlation ideas from this chapter.
Quantitative Methods carries a modest topic weight, but this chapter is highly testable because every question reduces to a pattern: read the exhibit, name the problem, state the consequence, pick the fix. Once you learn the patterns, the marks are quick and reliable. The same regression reading skills also help you in other topics where a vignette shows regression output. Since all questions come from a vignette, you also save time by recognising which test the exhibit is pointing to.
Extensions of Multiple Regression: topics in the order to study them
- 1Influence Analysis: Outliers and LeverageIt is about the data going into the model, so it comes first. It also introduces residual-based diagnostics you will reuse.
- 2Dummy Variables in Multiple RegressionThis is a simple tool for building models, and interpreting intercept and slope shifts is needed before the problems that follow.
- 3HeteroskedasticityThis is the first assumption violation. It introduces the pattern of detect, consequence, then correct that every later violation follows.
- 4Serial CorrelationIt follows the same pattern as heteroskedasticity, so you can compare the two, including the tests and the robust standard error fix.
- 5MulticollinearityIt differs from the previous two because it affects the standard errors of correlated variables rather than the residual pattern. Learn it after you have the other two clear.
- 6Model MisspecificationIt pulls the earlier problems together and asks what went wrong in model design, so it is best studied once you know each individual violation.
- 7Qualitative Dependent Variables: Probit, Logit, DiscriminantIt changes the type of dependent variable, so it comes last as a separate extension of the framework.
How to prepare Extensions of Multiple Regression
Treat the chapter as a set of diagnostic patterns. Your goal is to look at an exhibit and quickly say what is wrong, what it does, and what to do.
- Refresh the basic regression output: coefficient, standard error, t-statistic, p-value, F-test and adjusted R². Everything later depends on reading these quickly.
- For each violation, write a three-line card: how to detect it, what it does to coefficients and standard errors, and how to correct it. Keep the cards short enough to read on your phone.
- Build a comparison table in your notes for heteroskedasticity, serial correlation and multicollinearity. Include the effect on the t-statistics and whether the coefficient estimates stay unbiased.
- Practise dummy variables by writing the fitted equation for each category. Check how many dummies you need for n categories, and that the interpretation is relative to the omitted category.
- Work through vignette-style questions. Underline the test statistic, the critical value or p-value, and the decision in the exhibit before you read the answer choices.
- Finish with logit, probit and discriminant analysis. Focus on what the dependent variable is, what the output means, and how they differ from ordinary regression.
- In the last week, redo your missed questions and recite your comparison table from memory.
Common mistakes in Extensions of Multiple Regression
Saying heteroskedasticity or serial correlation biases the coefficient estimates.
Fix: Remember that these two leave the coefficients unbiased in the usual setup but make the standard errors, and hence the t-tests, unreliable.
Using n dummy variables for n categories.
Fix: Use n − 1 dummies and read each coefficient relative to the omitted base category.
Diagnosing multicollinearity from a single insignificant t-statistic.
Fix: Look for a significant F-test and high R² alongside insignificant individual coefficients, or a high variance inflation factor.
Confusing outliers with high-leverage observations.
Fix: Tie outliers to the dependent variable and leverage to the independent variables. Then ask whether the point actually changes the fitted model.
Matching the wrong remedy to the problem, such as robust errors for multicollinearity.
Fix: Keep the three-line cards and the comparison table, and always read the fix next to its problem.
Treating logit or probit output as a linear prediction of the outcome.
Fix: Remember that these models estimate the probability of an event, and the coefficients are not read as simple changes in the outcome.
Last-day revision: Extensions of Multiple Regression
- Outliers are extreme values of the dependent variable. High-leverage points are extreme values of the independent variables. Either can be influential.
- With n categories, use n − 1 dummy variables. The omitted category is the base, and each dummy coefficient is a difference from it.
- Heteroskedasticity means the error variance is not constant. Conditional heteroskedasticity is the problematic type, since it relates to the independent variables.
- Heteroskedasticity leaves coefficient estimates unbiased but makes standard errors unreliable, so t-tests mislead.
- The Breusch–Pagan test detects heteroskedasticity. A fix is to use robust (White-corrected) standard errors.
- Serial correlation means errors are correlated across observations. Positive serial correlation typically understates standard errors and overstates t-statistics.
- The Durbin–Watson test and the Breusch–Godfrey test detect serial correlation. Use Newey-West-type adjusted standard errors to correct it.
- Multicollinearity means independent variables are highly correlated. The classic sign is a high R² and a significant F-test with insignificant individual t-statistics.
- A variance inflation factor above a commonly used threshold (such as 5 or 10) signals concern about multicollinearity. Fixes include dropping or combining variables.
- Misspecification includes omitted variables, wrong functional form, inappropriate scaling, and pooling data that should not be pooled. It can bias coefficients.
- Logit and probit are used when the dependent variable is binary. Logit uses the logistic distribution and probit uses the normal distribution.
- Discriminant analysis produces a score that classifies observations into categories.
Extensions of Multiple Regression in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Extensions of Multiple Regression: frequently asked questions
What is the difference between heteroskedasticity and serial correlation?
Heteroskedasticity means the variance of the errors is not constant across observations. Serial correlation means the errors are correlated with each other, usually across time. Both leave coefficients unbiased in the standard setup but make the standard errors unreliable.
How do I spot multicollinearity in a Level II vignette?
Look for a high R² and a significant F-statistic while individual t-statistics are insignificant. A high variance inflation factor or very high correlations between independent variables also point to it.
When do I use logit or probit instead of ordinary regression?
Use them when the dependent variable is qualitative, such as default versus no default. They estimate the probability of the outcome, so the predicted value stays between 0 and 1.
How many dummy variables do I need?
You need one fewer than the number of categories. The omitted category acts as the base, and each dummy coefficient shows the difference from that base.
Do I need to memorise test critical values for this chapter?
Usually the vignette supplies the test statistic and often the critical value or p-value. Focus on knowing what each test detects, how to read the decision, and which correction to apply.