Skip to content

CFA Level II · CFA Level II Exam

Evaluating Regression Model Fit and Interpreting Model Results: formula sheet

Full chapter guide

Key formulas

Sum of squares identity
SST = RSS + SSE
RSS is explained variation; SSE is unexplained variation.
R-squared
R² = RSS ÷ SST = 1 − SSE ÷ SST
Never falls when a variable is added.
Adjusted R-squared
Adjusted R² = 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²)
n = observations, k = number of independent variables.
Standard error of estimate
SEE = √[SSE ÷ (n − k − 1)]
Also √MSE. Lower means a tighter fit.
Mean square error
MSE = SSE ÷ (n − k − 1)
Degrees of freedom for SSE are n − k − 1.
AIC
AIC = n × ln(SSE ÷ n) + 2(k + 1)
Lower is better. Preferred for prediction.
BIC
BIC = n × ln(SSE ÷ n) + ln(n) × (k + 1)
Lower is better. Penalty is heavier when n ≥ 8, so it favors simpler models.
t-statistic for a coefficient
t = (b̂ − b₀) ÷ s(b̂)
b̂ is the estimated coefficient, b₀ the hypothesized value (usually 0), s(b̂) the standard error. Exhibits often print t for b₀ = 0 only.
Degrees of freedom
df = n − k − 1
k is the number of independent variables, excluding the intercept.
Confidence interval
b̂ ± t(critical) × s(b̂)
Use the two-sided critical value for the stated confidence level and df.
Decision rule (two-sided)
Reject H0 if |t| > t(critical), or if p-value < α
For a one-sided test, use the one-tail critical value and check the sign of t.
Rearranged standard error
s(b̂) = b̂ ÷ t
Use when the exhibit gives the coefficient and its t-statistic for b₀ = 0 but not the standard error.
F-statistic from ANOVA (all slopes zero)
F = MSR ÷ MSE = (RSS ÷ k) ÷ (SSE ÷ (n − k − 1))
H0: all slope coefficients = 0. Degrees of freedom: k and n − k − 1. One-tailed, right tail.
ANOVA components
SST = RSS + SSE; MSR = RSS ÷ k; MSE = SSE ÷ (n − k − 1)
SST is total variation, RSS is explained, SSE is unexplained. df: SST n − 1, RSS k, SSE n − k − 1.
F-test for a subset (restricted vs unrestricted)
F = [(SSE_R − SSE_U) ÷ q] ÷ [SSE_U ÷ (n − k − 1)]
q = number of restrictions (slopes set to zero). k = slopes in the unrestricted model. df: q and n − k − 1.
F in terms of R-squared
F = [(R²_U − R²_R) ÷ q] ÷ [(1 − R²_U) ÷ (n − k − 1)]
Valid when both models use the same dependent variable. For the all-slopes test, R²_R = 0.
Decision rule
Reject H0 if F > critical F(q, n − k − 1)
Rejecting means at least one tested slope is not zero.
Link between t and F with one restriction
F = t²
Holds when q = 1 and the test is two-sided.
Breusch-Pagan test statistic
BP = n × R² (from regressing squared residuals on the independent variables)
Chi-square with k degrees of freedom, one-tailed (right tail). H0: no conditional heteroskedasticity. Reject if BP exceeds the critical value.
Durbin-Watson statistic
DW ≈ 2 × (1 − r), where r is the correlation between residuals and lagged residuals
DW = 2 means no serial correlation. DW below 2 suggests positive, above 2 suggests negative.
Durbin-Watson decision rule (positive serial correlation)
DW < dl: reject H0; DW > du: fail to reject; dl ≤ DW ≤ du: inconclusive
H0: no positive serial correlation. The table gives dl and du using n and k.
Effect on standard errors
Conditional heteroskedasticity → standard errors unreliable (often understated in finance data, but the direction depends on how the variance relates to the regressors). Positive serial correlation → standard errors typically understated → t-statistics overstated
The F-test is also unreliable. When standard errors are understated, you get too many Type I errors. When they are overstated, the reverse holds. Coefficient estimates are not biased in a correct specification.
Corrections
Heteroskedasticity: White (robust) standard errors. Serial correlation: Newey-West (also robust to heteroskedasticity)
Newey-West is the safe choice when both problems may be present.
Variance inflation factor
VIF_j = 1 ÷ (1 − R²_j)
R²_j comes from regressing independent variable j on all the other independent variables. It is not the R² of the main model.
VIF rule of thumb
VIF > 5: investigate; VIF > 10: serious multicollinearity
These are rules of thumb, not strict tests. VIF of 1 means no correlation with the other variables.
Effect on standard error
Higher VIF → larger standard error → smaller t-statistic
Coefficients stay unbiased and consistent. Only precision is lost.
Classic symptom
High R² and significant F-test, but insignificant t-statistics
Strong sign of multicollinearity, but its absence does not prove there is none.
Pairwise correlation check
Large |correlation| between two regressors
Only catches pairs. With several regressors, multicollinearity can exist even when no pair is highly correlated, so use VIF.
Average leverage
average h = (k + 1) ÷ n
k = number of independent variables, n = number of observations. The hat values sum to k + 1.
Leverage rule of thumb
Potentially influential if h > 3 × (k + 1) ÷ n
Flags unusual X values. It is a screening rule, not proof of influence.
Studentized residual
t_i* = e_i* ÷ s_e*, where e_i* = Y_i − Ŷ_i(i) is the residual of observation i computed from the model fitted without it, and s_e* is its standard error
Flags Y outliers. Rule of thumb: |t*| > 3, or compare with the critical t-value with n − k − 2 degrees of freedom.
Cook's distance
D_i = Σ(Ŷ_j − Ŷ_j(i))² ÷ ((k + 1) × MSE), summed over all observations j, where Ŷ_j(i) is the fitted value for j from the model refit without observation i
Combines residual size and leverage. Larger D means greater influence on the fitted values. An equivalent form is D_i = (e_i² ÷ ((k + 1) × MSE)) × (h_i ÷ (1 − h_i)²), where e_i is the ordinary residual from the full model.
Cook's distance rules of thumb
D > √(k ÷ n) is the curriculum's common flag; other conventions: D > 0.5 worth a look, D > 1 likely influential
These are conventions. The exam will tell you which threshold to use or give a clear gap.
Number of dummies
Dummies needed = n − 1 (n = number of categories)
Using n dummies plus an intercept causes perfect multicollinearity.
Intercept dummy
Y = b0 + b1D + b2X + ε
Base group intercept = b0. D = 1 group intercept = b0 + b1. Slope on X is the same.
Slope dummy (interaction)
Y = b0 + b1D + b2X + b3(D × X) + ε
Slope when D = 0 is b2. Slope when D = 1 is b2 + b3.
Logit model
ln[P ÷ (1 − P)] = b0 + b1X1 + … + bkXk
Left side is the log odds. Estimated by maximum likelihood.
Probability from log odds
P = 1 ÷ (1 + e^(−z)), where z = b0 + b1X1 + …
Gives a probability between 0 and 1. Odds = P ÷ (1 − P) = e^z.

Quick revision

  • R² = explained variation ÷ total variation; it never falls when you add a variable.
  • Adjusted R² penalises extra variables and can fall; it is always at most R².
  • t-statistic = (estimated coefficient − hypothesised value) ÷ standard error, with n − k − 1 degrees of freedom.
  • F = MSR ÷ MSE tests whether all slope coefficients are jointly zero; it is a one-tailed test with k and n − k − 1 degrees of freedom.
  • In a multiple regression, a significant F-test only tells you at least one slope is non-zero, not which one.
  • Heteroskedasticity: error variance changes with the independent variables; coefficients stay consistent, but standard errors are unreliable. Use the Breusch-Pagan test and robust (White) standard errors.
  • Serial correlation: errors are correlated across time; positive serial correlation usually makes standard errors too small and t-statistics too large. Use the Durbin-Watson test and Newey-West standard errors.
  • Multicollinearity: highly correlated independent variables inflate standard errors; the signal is a high R² and significant F with insignificant t-statistics. Check variance inflation factors.
  • Outliers and high-leverage points can drive results; leverage and Cook's distance help you find influential observations.
  • With a dummy variable for n categories, use n − 1 dummies; the omitted category is the benchmark.
  • Misspecification, such as an omitted variable or wrong functional form, can make coefficients biased and inconsistent.

Common mistakes

  • Choosing the model with the highest R² when models have different numbers of variables. Fix: R² cannot fall when variables are added. Compare adjusted R², AIC or BIC.
  • Picking the highest AIC or BIC as best. Fix: Both are penalized error measures. Lower is better.
  • Using the printed t-statistic to test a null value other than zero. Fix: Recompute t = (b̂ − b₀) ÷ s(b̂) whenever the null is not zero.
  • Using n − 1 or n − k as degrees of freedom. Fix: Use n − k − 1. The extra 1 is for the intercept.
  • Using n − k instead of n − k − 1 in the denominator degrees of freedom. Fix: Always write n − k − 1, where k counts slope coefficients only.
  • Taking SSE from the restricted model in the denominator of the subset F-test. Fix: The denominator is always the unrestricted SSE divided by its degrees of freedom.
  • Saying the coefficients are biased because of heteroskedasticity or serial correlation. Fix: In a correctly specified model the coefficients stay unbiased. The damage is to the standard errors and tests.
  • Using a two-tailed test for Breusch-Pagan. Fix: BP is a one-tailed chi-square test. A large statistic indicates heteroskedasticity.
  • Saying multicollinearity biases the coefficient estimates. Fix: Remember estimates stay unbiased and consistent. Only the standard errors are inflated and estimates become unstable.
  • Using the model's R-squared in the VIF formula. Fix: VIF uses R²_j from regressing one independent variable on the other independent variables. The dependent variable is not involved.

Exam tips

  • Expect a vignette with an ANOVA table. Find RSS, SSE and the degrees of freedom before reading the question.
  • Reading the question's purpose is the key. Forecasting means AIC; description means BIC.
  • Questions often test the direction: highest adjusted R², lowest AIC and BIC. Check this before computing.
  • Watch for a conclusion that treats R² as proof of a valid model. It is usually the wrong answer.
  • There is no penalty for wrong answers, so always answer every question.
  • Check the null value first. If it is not zero, the printed t-statistic does not answer the question.
  • Compute df as n − k − 1 with k as slope count; vignettes often give n and the variables in different places.
  • If a p-value is printed, comparing it with α is the fastest route. Use it before any calculation.