CFA Level II · CFA Level II Exam
Extensions of Multiple Regression: formula sheet
Key formulas
- Leverage rule of thumb
- Potentially high leverage if hᵢᵢ > 3 × (k + 1) ÷ n
- k is the number of independent variables and n is the number of observations. (k + 1) ÷ n is the average leverage. Leverage values lie between 0 and 1.
- Studentized residual
- tᵢ* = eᵢ ÷ (standard error of the residual estimated with observation i deleted)
- Compare |tᵢ*| with the critical t-value from a t-distribution with n − k − 2 degrees of freedom. A value beyond the critical value flags an outlier. A cutoff of |tᵢ*| > 3 is also used as a common rule of thumb.
- Cook's D rules of thumb
- Dᵢ > 0.5: possibly influential. Dᵢ > 1: very likely influential. Another rule: Dᵢ > √(k ÷ n)
- Values above 0.5 (possibly influential) and above 1 (very likely influential) are the main guides. The curriculum also gives √(k ÷ n) as another cutoff, where k is the number of independent variables and n is the number of observations. It is not the standard for every sample size. Cook's D measures the overall change in fitted values when observation i is dropped. Use the cutoff the question supplies.
- Classification rule
- High leverage = unusual X; outlier = unusual Y given X; influential = changes the results when removed
- The cutoffs are guides, not hard laws. Always say 'flag for investigation'.
- Number of dummies
- Dummies needed = n − 1, for n mutually exclusive categories
- Use n − 1 when the model has an intercept. The omitted category is the base.
- Intercept dummy model
- Y = b0 + b1D + b2X + ε
- Base group intercept = b0. D = 1 group intercept = b0 + b1. Slope is b2 for both.
- Slope dummy (interaction) model
- Y = b0 + b1X + b2(D × X) + ε
- Base slope = b1. D = 1 slope = b1 + b2. Intercept is b0 for both.
- Combined model
- Y = b0 + b1D + b2X + b3(D × X) + ε
- D = 1 line: intercept b0 + b1, slope b2 + b3.
- Test of a dummy coefficient
- t = (b̂ − 0) ÷ s(b̂), with n − k − 1 degrees of freedom
- A significant coefficient means that category differs from the base category.
- Breusch-Pagan test statistic
- BP = n × R²(resid)
- n = number of observations. R²(resid) is from the regression of squared residuals on the independent variables.
- Breusch-Pagan decision rule
- Chi-square with k degrees of freedom; reject H0 (no conditional heteroskedasticity) if BP > critical value
- k = number of independent variables in the original regression. It is a one-tailed test.
- Null and alternative
- H0: no conditional heteroskedasticity; Ha: conditional heteroskedasticity
- Rejecting H0 means you should use robust standard errors.
- t-statistic
- t = (b̂ − b hypothesized) ÷ standard error of b̂
- If the standard error is too small, t is too large. Recompute t using the robust standard error.
- Durbin-Watson approximation
- DW ≈ 2(1 − r)
- r is the sample correlation between residuals and their first lag. DW ranges from 0 to 4. DW = 2 means no first-order correlation; below 2 points to positive; above 2 points to negative.
- Durbin-Watson decision rule (positive correlation)
- H0: no positive serial correlation. DW < dL: reject. dL ≤ DW ≤ dU: inconclusive. DW > dU: fail to reject
- dL and dU come from a table and depend on n and the number of independent variables k. The exam normally gives them in the vignette.
- Durbin-Watson decision rule (negative correlation)
- DW > 4 − dL: reject. 4 − dU ≤ DW ≤ 4 − dL: inconclusive. DW < 4 − dU: fail to reject
- Mirror image of the positive test. Only use it if the question asks about negative correlation.
- Breusch-Godfrey test statistic
- BG = n × R² (from the auxiliary regression), compared with χ² with p degrees of freedom
- Regress the original residuals on the original independent variables plus p lagged residuals. H0: no serial correlation up to lag p. Reject if the statistic exceeds the critical value. This is a one-tailed test.
- Direction of effect on standard errors
- Positive serial correlation → SE too small, t too large. Negative → SE tends to be too large, t tends to be too small
- The negative case is a tendency (when regressors are positively correlated). Coefficient estimates are unaffected when there is no lagged dependent variable.
- Variance inflation factor
- VIFj = 1 ÷ (1 − Rj²)
- Rj² comes from regressing independent variable j on all the other independent variables. It is not the R² of the main model.
- VIF rules of thumb
- VIF > 5: investigate. VIF > 10: serious multicollinearity
- These are conventions, not exact tests. A VIF of 1 means no collinearity with the other variables.
- Classic symptom pattern
- High R², significant F-test, insignificant individual t-statistics
- Suggests multicollinearity but does not prove it. Confirm with VIF or pairwise correlations.
- Effect on estimates
- Slopes: unbiased and consistent. Standard errors: inflated. Type II errors: more likely
- You fail to reject a false null more often because t-statistics are too small.
- Principles of a good model
- Economic reasoning + parsimony + good out-of-sample performance + appropriate functional form + no violated assumptions
- Use this as a checklist when a vignette asks whether a model is well specified.
- Omitted variable bias (direction)
- Bias in included coefficient has the sign of [corr(omitted, included) × true effect of omitted]
- Bias exists only if the omitted variable affects Y and is correlated with an included regressor. If uncorrelated, slope estimates stay unbiased.
- Log-linear (log-log) form
- ln(Y) = b0 + b1·ln(X) + ε
- Slope b1 is an elasticity: a 1% change in X is associated with about b1% change in Y.
- Quadratic term
- Y = b0 + b1·X + b2·X² + ε
- Use when residuals against X show a curve. The marginal effect of X is b1 + 2·b2·X, so it changes with X.
- Interaction term
- Y = b0 + b1·X1 + b2·X2 + b3·(X1·X2) + ε
- Marginal effect of X1 is b1 + b3·X2. Omitting a real interaction is a form error.
- Remedy for a unit root: first differences
- ΔY = b0 + b1·ΔX + ε
- Differencing removes a unit root. It is a remedy, not a test. Diagnose with unit root tests and a cointegration test (Engle-Granger). If the levels are cointegrated, the levels regression can still be valid.
- Linear probability model
- Y = b0 + b1X1 + ... + bkXk + ε, where Y is 0 or 1
- Fitted values can fall outside 0 to 1. Errors are heteroskedastic.
- Logit model
- p = 1 ÷ (1 + e^−(b0 + b1X1 + ... + bkXk))
- Uses the logistic CDF. The probability always lies between 0 and 1.
- Log odds
- ln[p ÷ (1 − p)] = b0 + b1X1 + ... + bkXk
- Each slope is the change in log odds for a one-unit change in that X, holding others constant.
- Odds
- odds = p ÷ (1 − p); p = odds ÷ (1 + odds)
- Use to convert between probability and odds.
- Probit model
- p = N(b0 + b1X1 + ... + bkXk), where N is the standard normal CDF
- Uses the cumulative normal. The coefficient sign shows direction only.
- Discriminant function
- Score = a0 + a1X1 + ... + akXk; classify by comparing the score with a cutoff
- Produces a classification, not a probability.
Quick revision
- Outliers are extreme values of the dependent variable. High-leverage points are extreme values of the independent variables. Either can be influential.
- With n categories, use n − 1 dummy variables. The omitted category is the base, and each dummy coefficient is a difference from it.
- Heteroskedasticity means the error variance is not constant. Conditional heteroskedasticity is the problematic type, since it relates to the independent variables.
- Heteroskedasticity leaves coefficient estimates unbiased but makes standard errors unreliable, so t-tests mislead.
- The Breusch–Pagan test detects heteroskedasticity. A fix is to use robust (White-corrected) standard errors.
- Serial correlation means errors are correlated across observations. Positive serial correlation typically understates standard errors and overstates t-statistics.
- The Durbin–Watson test and the Breusch–Godfrey test detect serial correlation. Use Newey-West-type adjusted standard errors to correct it.
- Multicollinearity means independent variables are highly correlated. The classic sign is a high R² and a significant F-test with insignificant individual t-statistics.
- A variance inflation factor above a commonly used threshold (such as 5 or 10) signals concern about multicollinearity. Fixes include dropping or combining variables.
- Misspecification includes omitted variables, wrong functional form, inappropriate scaling, and pooling data that should not be pooled. It can bias coefficients.
- Logit and probit are used when the dependent variable is binary. Logit uses the logistic distribution and probit uses the normal distribution.
- Discriminant analysis produces a score that classifies observations into categories.
Common mistakes
- Treating high leverage and outlier as the same thing. Fix: Leverage is about X values. An outlier is about Y given X, shown by a large residual.
- Using k instead of k + 1 in the leverage cutoff. Fix: Use 3 × (k + 1) ÷ n. For 2 independent variables, use 3 × 3 ÷ n. Note that the Cook's D cutoff √(k ÷ n) uses k.
- Including a dummy for every category along with the intercept. Fix: Use n − 1 dummies. The base category is absorbed in the intercept. Otherwise you get perfect multicollinearity.
- Reading a dummy coefficient as an absolute level. Fix: Always say the coefficient is the difference from the base category, holding other variables constant.
- Saying heteroskedasticity biases the coefficient estimates. Fix: Remember that coefficients stay unbiased and consistent. Only the standard errors, and so the tests, are affected.
- Using the R² of the original regression in the Breusch-Pagan statistic. Fix: Use the R² from the regression of squared residuals on the independent variables.
- Saying serial correlation biases the regression coefficients. Fix: Remember: coefficients stay unbiased and consistent (without a lagged dependent variable). Only the standard errors, t-tests and F-test are affected.
- Getting the direction of the standard error effect backwards. Fix: Positive correlation makes the data look more informative than it is, so standard errors are too small and t-statistics too large. Type I errors rise. Negative correlation tends to do the opposite (when regressors are positively correlated).
- Saying multicollinearity biases the coefficient estimates. Fix: Remember that the estimates stay unbiased and consistent. Only the standard errors are inflated.
- Using the model's R² in the VIF formula. Fix: Use Rj² from regressing variable j on the other independent variables, not on the dependent variable.
Exam tips
- Write the leverage cutoff first. It is the one number you must calculate yourself.
- Read option wording closely. Examiners swap 'outlier' and 'high leverage' to build wrong answers.
- If the vignette gives its own cutoff, use it instead of a rule of thumb.
- When asked what the analyst should do, choose investigate and compare results with and without the point, not delete.
- Remember k excludes the intercept but the leverage cutoff uses k + 1.
- Find the coding in the vignette first. A dummy that is 1 for the base group flips every interpretation.
- Count dummies against categories. A model with all n dummies and an intercept signals a trap question.
- When asked for a forecast, write the full equation for the group, then plug in X. Do not skip the interaction term.