Skip to content

CFA Level II · CFA Level II Exam

Extensions of Multiple Regression: formula sheet

Full chapter guide

Key formulas

Leverage rule of thumb
Potentially high leverage if hᵢᵢ > 3 × (k + 1) ÷ n
k is the number of independent variables and n is the number of observations. (k + 1) ÷ n is the average leverage. Leverage values lie between 0 and 1.
Studentized residual
tᵢ* = eᵢ ÷ (standard error of the residual estimated with observation i deleted)
Compare |tᵢ*| with the critical t-value from a t-distribution with n − k − 2 degrees of freedom. A value beyond the critical value flags an outlier. A cutoff of |tᵢ*| > 3 is also used as a common rule of thumb.
Cook's D rules of thumb
Dᵢ > 0.5: possibly influential. Dᵢ > 1: very likely influential. Another rule: Dᵢ > √(k ÷ n)
Values above 0.5 (possibly influential) and above 1 (very likely influential) are the main guides. The curriculum also gives √(k ÷ n) as another cutoff, where k is the number of independent variables and n is the number of observations. It is not the standard for every sample size. Cook's D measures the overall change in fitted values when observation i is dropped. Use the cutoff the question supplies.
Classification rule
High leverage = unusual X; outlier = unusual Y given X; influential = changes the results when removed
The cutoffs are guides, not hard laws. Always say 'flag for investigation'.
Number of dummies
Dummies needed = n − 1, for n mutually exclusive categories
Use n − 1 when the model has an intercept. The omitted category is the base.
Intercept dummy model
Y = b0 + b1D + b2X + ε
Base group intercept = b0. D = 1 group intercept = b0 + b1. Slope is b2 for both.
Slope dummy (interaction) model
Y = b0 + b1X + b2(D × X) + ε
Base slope = b1. D = 1 slope = b1 + b2. Intercept is b0 for both.
Combined model
Y = b0 + b1D + b2X + b3(D × X) + ε
D = 1 line: intercept b0 + b1, slope b2 + b3.
Test of a dummy coefficient
t = (b̂ − 0) ÷ s(b̂), with n − k − 1 degrees of freedom
A significant coefficient means that category differs from the base category.
Breusch-Pagan test statistic
BP = n × R²(resid)
n = number of observations. R²(resid) is from the regression of squared residuals on the independent variables.
Breusch-Pagan decision rule
Chi-square with k degrees of freedom; reject H0 (no conditional heteroskedasticity) if BP > critical value
k = number of independent variables in the original regression. It is a one-tailed test.
Null and alternative
H0: no conditional heteroskedasticity; Ha: conditional heteroskedasticity
Rejecting H0 means you should use robust standard errors.
t-statistic
t = (b̂ − b hypothesized) ÷ standard error of b̂
If the standard error is too small, t is too large. Recompute t using the robust standard error.
Durbin-Watson approximation
DW ≈ 2(1 − r)
r is the sample correlation between residuals and their first lag. DW ranges from 0 to 4. DW = 2 means no first-order correlation; below 2 points to positive; above 2 points to negative.
Durbin-Watson decision rule (positive correlation)
H0: no positive serial correlation. DW < dL: reject. dL ≤ DW ≤ dU: inconclusive. DW > dU: fail to reject
dL and dU come from a table and depend on n and the number of independent variables k. The exam normally gives them in the vignette.
Durbin-Watson decision rule (negative correlation)
DW > 4 − dL: reject. 4 − dU ≤ DW ≤ 4 − dL: inconclusive. DW < 4 − dU: fail to reject
Mirror image of the positive test. Only use it if the question asks about negative correlation.
Breusch-Godfrey test statistic
BG = n × R² (from the auxiliary regression), compared with χ² with p degrees of freedom
Regress the original residuals on the original independent variables plus p lagged residuals. H0: no serial correlation up to lag p. Reject if the statistic exceeds the critical value. This is a one-tailed test.
Direction of effect on standard errors
Positive serial correlation → SE too small, t too large. Negative → SE tends to be too large, t tends to be too small
The negative case is a tendency (when regressors are positively correlated). Coefficient estimates are unaffected when there is no lagged dependent variable.
Variance inflation factor
VIFj = 1 ÷ (1 − Rj²)
Rj² comes from regressing independent variable j on all the other independent variables. It is not the R² of the main model.
VIF rules of thumb
VIF > 5: investigate. VIF > 10: serious multicollinearity
These are conventions, not exact tests. A VIF of 1 means no collinearity with the other variables.
Classic symptom pattern
High R², significant F-test, insignificant individual t-statistics
Suggests multicollinearity but does not prove it. Confirm with VIF or pairwise correlations.
Effect on estimates
Slopes: unbiased and consistent. Standard errors: inflated. Type II errors: more likely
You fail to reject a false null more often because t-statistics are too small.
Principles of a good model
Economic reasoning + parsimony + good out-of-sample performance + appropriate functional form + no violated assumptions
Use this as a checklist when a vignette asks whether a model is well specified.
Omitted variable bias (direction)
Bias in included coefficient has the sign of [corr(omitted, included) × true effect of omitted]
Bias exists only if the omitted variable affects Y and is correlated with an included regressor. If uncorrelated, slope estimates stay unbiased.
Log-linear (log-log) form
ln(Y) = b0 + b1·ln(X) + ε
Slope b1 is an elasticity: a 1% change in X is associated with about b1% change in Y.
Quadratic term
Y = b0 + b1·X + b2·X² + ε
Use when residuals against X show a curve. The marginal effect of X is b1 + 2·b2·X, so it changes with X.
Interaction term
Y = b0 + b1·X1 + b2·X2 + b3·(X1·X2) + ε
Marginal effect of X1 is b1 + b3·X2. Omitting a real interaction is a form error.
Remedy for a unit root: first differences
ΔY = b0 + b1·ΔX + ε
Differencing removes a unit root. It is a remedy, not a test. Diagnose with unit root tests and a cointegration test (Engle-Granger). If the levels are cointegrated, the levels regression can still be valid.
Linear probability model
Y = b0 + b1X1 + ... + bkXk + ε, where Y is 0 or 1
Fitted values can fall outside 0 to 1. Errors are heteroskedastic.
Logit model
p = 1 ÷ (1 + e^−(b0 + b1X1 + ... + bkXk))
Uses the logistic CDF. The probability always lies between 0 and 1.
Log odds
ln[p ÷ (1 − p)] = b0 + b1X1 + ... + bkXk
Each slope is the change in log odds for a one-unit change in that X, holding others constant.
Odds
odds = p ÷ (1 − p); p = odds ÷ (1 + odds)
Use to convert between probability and odds.
Probit model
p = N(b0 + b1X1 + ... + bkXk), where N is the standard normal CDF
Uses the cumulative normal. The coefficient sign shows direction only.
Discriminant function
Score = a0 + a1X1 + ... + akXk; classify by comparing the score with a cutoff
Produces a classification, not a probability.

Quick revision

  • Outliers are extreme values of the dependent variable. High-leverage points are extreme values of the independent variables. Either can be influential.
  • With n categories, use n − 1 dummy variables. The omitted category is the base, and each dummy coefficient is a difference from it.
  • Heteroskedasticity means the error variance is not constant. Conditional heteroskedasticity is the problematic type, since it relates to the independent variables.
  • Heteroskedasticity leaves coefficient estimates unbiased but makes standard errors unreliable, so t-tests mislead.
  • The Breusch–Pagan test detects heteroskedasticity. A fix is to use robust (White-corrected) standard errors.
  • Serial correlation means errors are correlated across observations. Positive serial correlation typically understates standard errors and overstates t-statistics.
  • The Durbin–Watson test and the Breusch–Godfrey test detect serial correlation. Use Newey-West-type adjusted standard errors to correct it.
  • Multicollinearity means independent variables are highly correlated. The classic sign is a high R² and a significant F-test with insignificant individual t-statistics.
  • A variance inflation factor above a commonly used threshold (such as 5 or 10) signals concern about multicollinearity. Fixes include dropping or combining variables.
  • Misspecification includes omitted variables, wrong functional form, inappropriate scaling, and pooling data that should not be pooled. It can bias coefficients.
  • Logit and probit are used when the dependent variable is binary. Logit uses the logistic distribution and probit uses the normal distribution.
  • Discriminant analysis produces a score that classifies observations into categories.

Common mistakes

  • Treating high leverage and outlier as the same thing. Fix: Leverage is about X values. An outlier is about Y given X, shown by a large residual.
  • Using k instead of k + 1 in the leverage cutoff. Fix: Use 3 × (k + 1) ÷ n. For 2 independent variables, use 3 × 3 ÷ n. Note that the Cook's D cutoff √(k ÷ n) uses k.
  • Including a dummy for every category along with the intercept. Fix: Use n − 1 dummies. The base category is absorbed in the intercept. Otherwise you get perfect multicollinearity.
  • Reading a dummy coefficient as an absolute level. Fix: Always say the coefficient is the difference from the base category, holding other variables constant.
  • Saying heteroskedasticity biases the coefficient estimates. Fix: Remember that coefficients stay unbiased and consistent. Only the standard errors, and so the tests, are affected.
  • Using the R² of the original regression in the Breusch-Pagan statistic. Fix: Use the R² from the regression of squared residuals on the independent variables.
  • Saying serial correlation biases the regression coefficients. Fix: Remember: coefficients stay unbiased and consistent (without a lagged dependent variable). Only the standard errors, t-tests and F-test are affected.
  • Getting the direction of the standard error effect backwards. Fix: Positive correlation makes the data look more informative than it is, so standard errors are too small and t-statistics too large. Type I errors rise. Negative correlation tends to do the opposite (when regressors are positively correlated).
  • Saying multicollinearity biases the coefficient estimates. Fix: Remember that the estimates stay unbiased and consistent. Only the standard errors are inflated.
  • Using the model's R² in the VIF formula. Fix: Use Rj² from regressing variable j on the other independent variables, not on the dependent variable.

Exam tips

  • Write the leverage cutoff first. It is the one number you must calculate yourself.
  • Read option wording closely. Examiners swap 'outlier' and 'high leverage' to build wrong answers.
  • If the vignette gives its own cutoff, use it instead of a rule of thumb.
  • When asked what the analyst should do, choose investigate and compare results with and without the point, not delete.
  • Remember k excludes the intercept but the leverage cutoff uses k + 1.
  • Find the coding in the vignette first. A dummy that is 1 for the base group flips every interpretation.
  • Count dummies against categories. A model with all n dummies and an intercept signals a trap question.
  • When asked for a forecast, write the full equation for the group, then plug in X. Do not skip the interaction term.