Skip to content

CFA Level II · CFA Level II Exam

Basics of Multiple Regression and Underlying Assumptions: formula sheet

Full chapter guide

Key formulas

Multiple regression model
Yi = b0 + b1X1i + b2X2i + ... + bkXki + εi
i indexes the observation. k is the number of independent variables.
Predicted (fitted) value
Ŷ = b̂0 + b̂1X1 + b̂2X2 + ... + b̂kXk
Use the estimated coefficients and drop the error term, which is expected to be zero.
Residual
ei = Yi − Ŷi
OLS minimises Σei², the sum of squared residuals.
Interpretation of a partial slope
bj = expected change in Y for a one-unit change in Xj, other X variables held constant
Always state the units of Y and Xj.
Degrees of freedom of the regression
Residual df = n − k − 1
n is observations and k is independent variables. The extra 1 is for the intercept.
Multiple regression model
Yi = b0 + b1X1i + b2X2i + ... + bkXki + εi
Linear in the coefficients. Variables can be transformed (for example, a log) and still satisfy linearity.
Zero conditional mean of errors
E(ε | X1, ..., Xk) = 0
If errors are related to an X variable, coefficient estimates are biased. This often signals an omitted variable or misspecification.
Homoskedasticity
Var(εi) = σ² for all i
Constant error variance. Violation is heteroskedasticity.
Independence of errors
Cov(εi, εj) = 0 for i ≠ j
Violation is serial correlation, common in time series.
Normality of errors
ε ~ N(0, σ²)
Matters most for small samples.
No exact collinearity
No Xj is an exact linear combination of the other X variables
Exact collinearity makes estimation impossible. High but imperfect correlation is multicollinearity.
t-statistic for a coefficient
t = (b̂ⱼ − bⱼ,H0) ÷ s(b̂ⱼ)
b̂ⱼ is the estimated slope, bⱼ,H0 is the hypothesized value (often 0), s(b̂ⱼ) is its standard error.
Degrees of freedom
df = n − k − 1
n observations, k independent variables. The +1 accounts for the intercept.
Confidence interval for a coefficient
b̂ⱼ ± t(critical) × s(b̂ⱼ)
Use the two-tailed critical value for the chosen confidence level, with n − k − 1 degrees of freedom.
Decision rule (t-test)
Reject H0 if |t| > t(critical)
For a one-tailed test, use the one-tailed critical value and check the sign of t.
Decision rule (p-value)
Reject H0 if p-value < α
α is the significance level, such as 0.05.
Variation decomposition
SST = SSR + SSE
SST is total, SSR is explained (regression), SSE is unexplained (residual).
Degrees of freedom
Regression = k; Error = n − k − 1; Total = n − 1
k is the number of independent variables, n the number of observations.
Mean squares
MSR = SSR ÷ k; MSE = SSE ÷ (n − k − 1)
Mean square = SS ÷ df.
F-statistic
F = MSR ÷ MSE, df = k and n − k − 1
One-tailed right-tail test of H0: all slopes = 0.
R-squared
R² = SSR ÷ SST = 1 − SSE ÷ SST
Never decreases when a variable is added.
Adjusted R-squared
Adj R² = 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²)
Can fall when a variable is added. Not above R² when k ≥ 1.
Standard error of estimate
SEE = √MSE = √[SSE ÷ (n − k − 1)]
In the units of the dependent variable.
AIC
AIC = n × ln(SSE ÷ n) + 2(k + 1)
Lower is better. Compare models on the same data and dependent variable.
BIC
BIC = n × ln(SSE ÷ n) + ln(n) × (k + 1)
Lower is better. Heavier penalty than AIC when n ≥ 8.
Regression with a dummy
Y = b0 + b1X1 + b2D + ε
D = 1 if the condition is true, 0 otherwise. b2 is the shift in the intercept.
Intercept by group
D = 0: intercept = b0; D = 1: intercept = b0 + b2
Slope b1 is the same in both groups unless you add an interaction term.
Number of dummies
Dummies needed = n - 1 (for n categories, with an intercept)
The omitted category is the base. Using n dummies causes perfect multicollinearity.
Predicted value
Ŷ = b0 + b1X1 + b2X2 + ... + bkXk
Substitute the given values. Use 1 or 0 for each dummy.
t-test for a coefficient
t = (bj - 0) ÷ SE(bj), with n - k - 1 degrees of freedom
k = number of independent variables. Tests whether the group difference is zero.
Interaction dummy (slope shift)
Y = b0 + b1X + b2D + b3(D × X) + ε
When D = 1 the slope is b1 + b3 and the intercept is b0 + b2.

Quick revision

  • Model: Y = b0 + b1X1 + ... + bkXk + ε; each slope holds the other variables constant.
  • Degrees of freedom for coefficient t-tests: n − k − 1.
  • t-statistic = (estimated coefficient − hypothesised value) ÷ standard error; the usual null value is 0.
  • Reject the null if the absolute t-statistic exceeds the critical value, or if the p-value is below the significance level.
  • F-statistic = (SSR ÷ k) ÷ (SSE ÷ (n − k − 1)); it tests whether all slope coefficients are jointly zero.
  • The F-test for this null is a one-tailed test, with the rejection region in the upper tail.
  • R² = SSR ÷ SST = 1 − SSE ÷ SST; it never falls when you add a variable.
  • Adjusted R² = 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²); it can fall when a variable adds little.
  • Adjusted R² is less than or equal to R² for models with at least one independent variable.
  • Assumptions include linearity, homoskedastic errors, no serial correlation of errors, normally distributed errors, and independent variables not perfectly linearly related.
  • With n categories, use n − 1 dummy variables; the omitted category is the benchmark.
  • Prediction: substitute the given X values into the fitted equation; check that units match.

Common mistakes

  • Interpreting a slope without 'holding other variables constant'. Fix: Always add that the other independent variables are held constant. This is what makes the slope partial.
  • Assuming the multiple regression slope equals the simple regression slope. Fix: The slopes differ when X variables are correlated with each other. Use the slope from the model in the question.
  • Saying linearity means the X variables must enter without transformation. Fix: The model must be linear in the coefficients. A log of X or X squared can be used and the model is still linear in the parameters.
  • Treating multicollinearity as an exact-collinearity problem that stops estimation. Fix: Exact collinearity means estimation fails. Multicollinearity means high but imperfect correlation, which inflates standard errors but still gives estimates.
  • Using the wrong degrees of freedom, such as n − k instead of n − k − 1. Fix: Always write df = n − k − 1 and count k as the number of slope coefficients only.
  • Dividing the estimate by its standard error when H0 is not zero. Fix: Subtract the hypothesized value from the estimate first, then divide.
  • Using n − k instead of n − k − 1 as the error degrees of freedom. Fix: Always count the intercept. Error df = n − (k + 1). Check that regression df + error df = n − 1.
  • Treating F as a test of each coefficient. Fix: F tests all slopes jointly. Use t-statistics for individual slopes. A significant F does not mean every slope is significant.
  • Using n dummies for n categories while keeping the intercept. Fix: Use n - 1 dummies. The omitted category is the base and is captured by the intercept.
  • Reading a dummy coefficient as the absolute level of the group rather than a difference. Fix: State it as: the group's value differs from the base by the coefficient, holding other variables constant.

Exam tips

  • Level II gives you regression output inside a vignette. Find the coefficient table first, then read the question.
  • Interpretation questions test the phrase 'holding other variables constant'. Pick the option that includes it.
  • For prediction, watch the units of each variable and any negative values in the vignette.
  • When asked about a change in one variable, multiply only that slope by the change. Do not add the intercept.
  • If a question shifts to significance or fit, the same output also holds standard errors, t-statistics and R-squared. Read what is asked before using them.
  • Expect to name the violation from a symptom, then state its effect on coefficients or standard errors.
  • Learn the pairs: heteroskedasticity and serial correlation affect standard errors, multicollinearity inflates them, and omitted variables or wrong functional form can bias coefficients.
  • Watch for the phrase linear in the coefficients when an option mentions transformed variables.