CFA Level II · CFA Level II Exam
Basics of Multiple Regression and Underlying Assumptions
Multiple regression models a dependent variable as a linear function of two or more independent variables. To solve questions, read the coefficients and standard errors from the vignette exhibit, compute t-statistics or F-statistics, compare with critical values, and interpret each slope as the effect holding the other variables constant.
What this chapter covers
This chapter introduces the multiple linear regression model: Y = b0 + b1X1 + b2X2 + ... + bkXk + ε. You learn what each coefficient means, which assumptions the model needs, how to test coefficients and the whole model, and how to use dummy variables and make predictions.
On the exam, you will rarely run a regression. You will read a regression output table in a vignette exhibit and answer questions on it. That means reading coefficients, standard errors, t-statistics, ANOVA values, R-squared and adjusted R-squared, then deciding what they imply.
The chapter is the base for the next quantitative chapters on violations of assumptions (heteroskedasticity, serial correlation, multicollinearity), model misspecification and time-series analysis. It also supports the other topics that use regression: equity valuation and factor models, portfolio construction, and economics. If this chapter is solid, those chapters become much easier.
Quantitative Methods carries 5-10% of the exam, and regression is its core. Every item set is a vignette with exhibits, and regression output is one of the most common exhibit formats. The calculations are short: a t-statistic, an F-statistic, a prediction. The marks go to candidates who read the table quickly and know the interpretation rules. There is also no penalty for wrong answers, so you should always answer, but a clear method gets you the correct answer much more often than guessing.
Basics of Multiple Regression and Underlying Assumptions: topics in the order to study them
- 1Multiple Linear Regression Model and CoefficientsStart here because every later topic uses the model equation and the meaning of slope coefficients held constant.
- 2Assumptions of the Multiple Linear Regression ModelNext, learn what the model needs to hold, since hypothesis tests are valid only if these assumptions are reasonably met.
- 3Hypothesis Testing of Regression CoefficientsWith the model and assumptions clear, you can learn the t-test on a single coefficient, the most frequent exam calculation.
- 4ANOVA, F-Test, R-squared and Adjusted R-squaredThis moves from single coefficients to the whole model, using sums of squares from the ANOVA table.
- 5Dummy Variables and Predicting with RegressionFinish with applications: encoding categories and plugging values into the fitted equation, which builds on everything before.
How to prepare Basics of Multiple Regression and Underlying Assumptions
Treat this chapter as a table-reading skill backed by a few short formulas. Practise with exhibits, not just with theory.
- Write the model equation and explain in your own words why a slope is the change in Y for a one-unit change in that X, holding the other variables constant.
- Memorise the assumptions as a list and, for each, note what it means in plain words and what goes wrong if it fails.
- Practise the t-test: t = (estimated coefficient − hypothesised value) ÷ standard error. Compare with the critical value using degrees of freedom n − k − 1, where k is the number of independent variables.
- Learn the ANOVA layout: regression sum of squares (SSR), sum of squared errors (SSE), total (SST). Practise F = (SSR ÷ k) ÷ (SSE ÷ (n − k − 1)) and R² = SSR ÷ SST until you can do them from an exhibit quickly.
- Work with dummy variables: with n categories, use n − 1 dummies, and interpret each coefficient as the difference from the omitted category.
- Solve full item sets under time pressure. For each, first scan the exhibit and note n, k, coefficients and standard errors, then answer the questions from those numbers.
- Keep an error log of every wrong answer, tagged by cause (misread table, wrong degrees of freedom, wrong tail), and review it before the exam.
Common mistakes in Basics of Multiple Regression and Underlying Assumptions
Using n − k instead of n − k − 1 as degrees of freedom for the coefficient t-test.
Fix: Write n − k − 1 at the top of your rough work and count k as the number of independent variables, excluding the intercept.
Interpreting a slope as the effect of X on Y without saying other variables are held constant.
Fix: Always read a slope as the expected change in Y for a one-unit change in that X, with the other independent variables unchanged.
Treating a high R² as proof of a good model.
Fix: Use adjusted R² to compare models with different numbers of variables, and check the assumptions and significance tests as well.
Using the F-test to conclude that every coefficient is significant.
Fix: Remember the F-test only shows that at least one slope coefficient is not zero. Use individual t-tests to judge each variable.
Using n dummy variables for n categories.
Fix: Use n − 1 dummies. Including all n with an intercept makes the variables perfectly collinear, violating an assumption.
Misreading the exhibit, such as taking a t-statistic for a coefficient or mixing up SSR and SSE.
Fix: Label n, k, coefficients, standard errors and sums of squares in the margin before answering, and check that each number matches its heading.
Last-day revision: Basics of Multiple Regression and Underlying Assumptions
- Model: Y = b0 + b1X1 + ... + bkXk + ε; each slope holds the other variables constant.
- Degrees of freedom for coefficient t-tests: n − k − 1.
- t-statistic = (estimated coefficient − hypothesised value) ÷ standard error; the usual null value is 0.
- Reject the null if the absolute t-statistic exceeds the critical value, or if the p-value is below the significance level.
- F-statistic = (SSR ÷ k) ÷ (SSE ÷ (n − k − 1)); it tests whether all slope coefficients are jointly zero.
- The F-test for this null is a one-tailed test, with the rejection region in the upper tail.
- R² = SSR ÷ SST = 1 − SSE ÷ SST; it never falls when you add a variable.
- Adjusted R² = 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²); it can fall when a variable adds little.
- Adjusted R² is less than or equal to R² for models with at least one independent variable.
- Assumptions include linearity, homoskedastic errors, no serial correlation of errors, normally distributed errors, and independent variables not perfectly linearly related.
- With n categories, use n − 1 dummy variables; the omitted category is the benchmark.
- Prediction: substitute the given X values into the fitted equation; check that units match.
Basics of Multiple Regression and Underlying Assumptions in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Basics of Multiple Regression and Underlying Assumptions: frequently asked questions
Do I need to calculate a regression by hand for the CFA Level II exam?
No. Questions give you the regression output in an exhibit. You work from coefficients, standard errors and ANOVA values to compute test statistics, interpret results and make predictions.
What is the difference between R-squared and adjusted R-squared?
R-squared is the share of variation in Y explained by the model, and it never falls when you add variables. Adjusted R-squared penalises extra variables, so it can fall when a new variable adds little. Use it to compare models with different numbers of variables.
When do I use the t-test and when the F-test?
Use the t-test to test one coefficient, for example whether a slope differs from zero. Use the F-test to test whether all slope coefficients are jointly equal to zero, which tests the model as a whole.
How many dummy variables do I need for a categorical variable?
Use one fewer than the number of categories. The omitted category is the benchmark, and each dummy coefficient measures the difference from that benchmark, holding the other variables constant.