CFA Level II · CFA Level II Exam
Model Misspecification: formula sheet
Key formulas
- Features of a correctly specified model
- Economic reasoning + parsimony + good out-of-sample fit + appropriate functional form + no assumption violations
- Use this as a checklist when a vignette asks whether a model is well specified.
- Log transformation for a non-linear relationship
- ln(Y) = b0 + b1X + ε or Y = b0 + b1 ln(X) + ε
- Use when the relationship is curved or when the variable grows proportionally. Check the residuals for a pattern first.
- Consequence of regressors correlated with the error
- Cov(Xj, ε) ≠ 0 ⇒ coefficient estimates are biased and inconsistent
- Hypothesis tests and confidence intervals are then unreliable. This is the most serious type of misspecification.
- Adjusted R-squared
- Adjusted R² = 1 − [(n − 1) ÷ (n − k − 1)] × (1 − R²)
- Penalises extra variables, so it supports parsimony. R² never falls when you add a variable.
- Omitted variable bias condition
- Bias in b₁ ≠ 0 when the omitted variable affects Y and is correlated with X₁
- Both conditions are needed. If the omitted variable is uncorrelated with the included regressors, no bias in their slopes arises from it.
- Direction of bias (general econometric rule of thumb)
- Sign of bias = sign(effect of omitted variable on Y) × sign(correlation of omitted variable with X₁)
- This is a general econometric rule of thumb for one omitted variable, not a formula from the CFA curriculum. Same signs overstate the included coefficient; opposite signs understate it. It does not carry over cleanly when several variables are omitted.
- Four types of functional form error
- Omitted variable | Wrong transformation | Inappropriate pooling | Inappropriate scaling
- Learn these as a checklist. Most questions ask you to name one.
- Typical consequence
- Misspecification → biased, inconsistent coefficients; unreliable standard errors and tests
- Adding more observations does not cure it.
- Lagged dependent variable model
- Y(t) = b0 + b1·X(t) + b2·Y(t-1) + ε(t)
- Problem arises when ε(t) is serially correlated: Y(t-1) is then correlated with ε(t), so estimates are inconsistent.
- Measurement error in one regressor (simple regression)
- Estimated slope is biased toward zero (attenuation)
- Holds for one regressor measured with classical noise. With multiple regressors the direction of bias is not predictable.
- Random walk (unit root)
- x(t) = x(t-1) + ε(t), so b1 = 1 in x(t) = b0 + b1·x(t-1) + ε(t)
- Nonstationary: mean-reversion level undefined, variance grows over time. First difference is stationary.
- First difference
- Δx(t) = x(t) - x(t-1)
- Common fix for a unit root before regressing.
- Breusch-Pagan test statistic
- BP = n × R² (from regressing squared residuals on the independent variables); chi-square with k degrees of freedom
- One-tailed, right-tail test. H0: no conditional heteroskedasticity. Reject if BP exceeds the critical value. R² is from the auxiliary regression, not the original model.
- Durbin-Watson statistic
- DW ≈ 2(1 − r), where r is the first-order correlation of residuals
- DW ≈ 2 means no serial correlation. DW below 2 suggests positive correlation; above 2 suggests negative. Compare with lower and upper critical values dl and du: below dl reject H0; between dl and du inconclusive; above du fail to reject (for positive correlation).
- Breusch-Godfrey test
- BG = (n − p) × R² from regressing residuals on the independent variables and p lagged residuals; chi-square with p degrees of freedom
- Tests serial correlation at higher orders. Using it, you can test more than one lag. Check the exact statistic form given in the question.
- Variance inflation factor
- VIF_j = 1 ÷ (1 − R_j²)
- R_j² is from regressing variable j on the other independent variables. A VIF above 5 warrants investigation and above 10 is a serious concern. These are rules of thumb, not strict laws.
- Corrections
- Heteroskedasticity: robust (White) standard errors. Serial correlation: Newey-West (HAC) standard errors. Multicollinearity: remove or combine variables, or use more data.
- Newey-West also corrects for heteroskedasticity. Robust errors change standard errors, not coefficients.
Quick revision
- A well-specified model is based on economic reasoning, has a parsimonious set of variables, and performs well out of sample.
- Omitted variables can bias coefficients and make estimates inconsistent when the omitted variable is correlated with included ones.
- Wrong functional form, such as not transforming a variable when the relationship is nonlinear, shows up as patterns in the residuals.
- Heteroskedasticity means error variance is not constant; coefficient estimates stay consistent, but standard errors are unreliable.
- Conditional heteroskedasticity is the form that matters most, because it is related to the independent variables; the Breusch-Pagan test checks for it.
- Fix heteroskedasticity with robust (White-corrected) standard errors, or with generalised least squares.
- Serial correlation means errors are correlated across observations; the Durbin-Watson test and the Breusch-Godfrey test detect it.
- Positive serial correlation makes standard errors too small, so t-statistics are too large and Type I errors are more likely. This holds when no regressor is a lagged dependent variable. If a lagged dependent variable is a regressor, serial correlation makes the coefficient estimates inconsistent, not just the standard errors.
- Fix serial correlation with Newey-West (HAC) standard errors or by modifying the model.
- Multicollinearity means independent variables are highly correlated: a high R² with insignificant individual t-statistics is the classic sign.
- A VIF above 5 warrants investigation and above 10 is a serious concern; remedies include dropping or combining variables.
- Nonstationary time series with a unit root can produce spurious regressions; test with a Dickey-Fuller test and consider differencing.
Common mistakes
- Choosing the model with the highest R-squared as the best specified. Fix: Remember that a good model also needs economic logic, parsimony and out-of-sample performance. Use adjusted R-squared and out-of-sample results to judge.
- Treating all misspecification as having the same consequence. Fix: Link each error to its effect. Any misspecification that makes a regressor correlated with the error, including omitted variables, wrong form or scale, and pooling, gives biased and inconsistent estimates, so you cannot trust tests.
- Saying any omitted variable causes bias. Fix: Check that the omitted variable affects Y and is correlated with an included regressor. If not, slope bias from it is not expected.
- Confusing pooling with omission. Fix: Pooling is about combining different regimes or groups in one equation. Omission is about a missing explanatory variable.
- Saying a lagged dependent variable always causes misspecification. Fix: Remember the trouble comes from the lag combined with serially correlated errors. If the errors are not serially correlated, including a lag does not create this problem.
- Saying measurement error in a regressor only inflates standard errors. Fix: Measurement error makes the regressor correlated with the error, so the coefficient itself is biased and inconsistent.
- Saying multicollinearity biases the coefficients. Fix: Coefficients remain unbiased and consistent. Only standard errors inflate, so significance is understated.
- Using the R² of the original regression in the Breusch-Pagan statistic. Fix: Use the R² from the regression of squared residuals on the independent variables.
Exam tips
- Read each analyst statement in the vignette as a possible specification claim and test it against the checklist.
- Do not choose an option just because it cites a higher R-squared; look for out-of-sample or economic support.
- Know which failures make estimates biased and inconsistent, since options often hinge on that wording.
- Keep specification errors and error-assumption violations distinct, but remember they can interact; options often blur them deliberately.
- Remember that perfect multicollinearity violates the assumption and prevents estimation. High multicollinearity leaves estimates unbiased but inflates standard errors, so t-statistics can be insignificant despite a high R-squared.
- If a question asks for a fix, match it: transform for curvature, drop or justify weak variables for parsimony, and remove the correlation between regressor and error.
- Most questions ask you to name the error from a short description, so memorize the four types and their clues.
- For omitted variable questions, always test the correlation condition before claiming bias.