CFA Level II · CFA Level II Exam
Model Misspecification for CFA Level II: Chapter Guide
Model misspecification means a regression model is built wrongly, so its coefficients or test statistics cannot be trusted. On the exam, you read the vignette, find the symptom (wrong form, nonstationary series, unequal variance, correlated errors, correlated regressors), name the problem, and state its effect and fix.
What this chapter covers
This chapter asks one question: can you trust a regression? You have already learned how to estimate and test a multiple regression. Here you learn how it goes wrong. A useful way to group the material is into four areas: the principles of good specification, errors in functional form, problems that appear in time series, and the three classic violations of regression assumptions: heteroskedasticity, serial correlation and multicollinearity. This is a suggested study grouping, not the official curriculum layout. In the curriculum, these ideas are spread across readings such as Multiple Regression and Time-Series Analysis.
The skill tested is diagnosis. A vignette gives you regression output, a plot description, a test statistic such as a Breusch-Pagan or Durbin-Watson result, or a variance inflation factor. You must recognise which problem it points to, say what it does to coefficients and standard errors, and pick the remedy.
Regression output may appear in item sets in various topics, and each question must be answered from the vignette. Interpreting a regression critically is a habit Level II rewards in many item sets.
Regression output appears inside item sets across several topics, and each question must be answered from the vignette. If you can quickly say whether a result is reliable, you pick up marks that candidates who only memorise formulas miss. The chapter is also compact and rule-based. Once you know each problem's symptom, effect and fix, the questions become pattern recognition. There is no penalty for wrong answers, so a clear decision rule lets you answer every question even when the numbers are heavy.
Model Misspecification: topics in the order to study them
- 1Principles of Model SpecificationAs a suggested starting point, it gives the framework for what a well-specified model looks like, so every later error has something to be compared against.
- 2Misspecified Functional FormIt is the most basic specification error, covering omitted variables, wrong transformations, pooled data and similar issues, and it uses ideas from the principles directly.
- 3Heteroskedasticity, Serial Correlation and MulticollinearityThese three violations share a structure (symptom, test, effect, fix), so it helps to learn them together before applying them to time series. In the curriculum they sit in the Multiple Regression reading.
- 4Time-Series MisspecificationIt extends serial correlation and the other violations to time-series data, adding issues such as trends, unit roots and nonstationarity, so we suggest taking it last. In the curriculum it is covered in the Time-Series Analysis reading.
How to prepare Model Misspecification
Treat this chapter as a diagnostic checklist. Your aim is to match a symptom in the vignette to a problem, then to its effect and fix.
- Read the principles of specification and write a one-page checklist of what a good model needs.
- For functional form errors, list each type, what it looks like in the residuals or output, and its consequence.
- Build a table on paper for heteroskedasticity, serial correlation and multicollinearity with four columns: symptom, detection test, effect on estimates and standard errors, and remedy.
- Learn what each test statistic tells you and the direction of the decision, not just its name. Know when a result signals a problem and when it does not.
- Study time-series misspecification by asking whether the series is stationary, whether errors are correlated, and whether a relationship is spurious.
- Practise item sets on a phone or paper. Underline the symptom in the vignette first, then name the problem, then answer the question.
- In the last week, redo only the questions you missed and recite your checklist from memory.
Common mistakes in Model Misspecification
Saying heteroskedasticity biases the regression coefficients.
Fix: Remember that the slope estimates stay consistent. The damage is to standard errors, so t-tests and F-tests become unreliable.
Getting the direction of the serial correlation effect wrong.
Fix: With positive serial correlation, standard errors are understated, t-statistics are overstated and you reject too often. Write this once and review it.
Reading a significant F-statistic as proof that there is no multicollinearity.
Fix: Multicollinearity often shows as a significant F-test and high R² alongside insignificant t-statistics. Check VIFs and correlations among regressors.
Applying the wrong remedy for the problem named in the vignette.
Fix: Keep the table: robust standard errors for heteroskedasticity, HAC standard errors for serial correlation, and dropping or combining variables for multicollinearity.
Running a regression on nonstationary series and trusting a high R².
Fix: Ask first whether the series is stationary. If a unit root is present, the relationship may be spurious, so difference the data or test properly.
Memorising test names without knowing what a result means.
Fix: For each test, know the null hypothesis, what a small p-value means, and which problem it indicates.
Last-day revision: Model Misspecification
- A well-specified model is based on economic reasoning, has a parsimonious set of variables, and performs well out of sample.
- Omitted variables can bias coefficients and make estimates inconsistent when the omitted variable is correlated with included ones.
- Wrong functional form, such as not transforming a variable when the relationship is nonlinear, shows up as patterns in the residuals.
- Heteroskedasticity means error variance is not constant; coefficient estimates stay consistent, but standard errors are unreliable.
- Conditional heteroskedasticity is the form that matters most, because it is related to the independent variables; the Breusch-Pagan test checks for it.
- Fix heteroskedasticity with robust (White-corrected) standard errors, or with generalised least squares.
- Serial correlation means errors are correlated across observations; the Durbin-Watson test and the Breusch-Godfrey test detect it.
- Positive serial correlation makes standard errors too small, so t-statistics are too large and Type I errors are more likely. This holds when no regressor is a lagged dependent variable. If a lagged dependent variable is a regressor, serial correlation makes the coefficient estimates inconsistent, not just the standard errors.
- Fix serial correlation with Newey-West (HAC) standard errors or by modifying the model.
- Multicollinearity means independent variables are highly correlated: a high R² with insignificant individual t-statistics is the classic sign.
- A VIF above 5 warrants investigation and above 10 is a serious concern; remedies include dropping or combining variables.
- Nonstationary time series with a unit root can produce spurious regressions; test with a Dickey-Fuller test and consider differencing.
Model Misspecification in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Model Misspecification: frequently asked questions
What is model misspecification in CFA Level II?
It is a flaw in how a regression model is built, such as a wrong form, missing variables, or violated assumptions. The result is unreliable coefficients or test statistics. The exam asks you to spot the flaw from the vignette and state its effect and fix.
Do I need to memorise test names like Breusch-Pagan and Durbin-Watson?
Yes, you should know which problem each test detects and how to read the outcome. The vignette will often give a statistic or p-value. You then decide whether it signals a problem and what follows from it.
How do I tell heteroskedasticity from serial correlation in a vignette?
Heteroskedasticity is about error variance changing with the independent variables or across observations. Serial correlation is about errors being related to each other across time or order. Look at which test or description the vignette gives.
Is this chapter worth studying if I am short of time?
Yes. It is compact, rule-based and regression output can appear inside other topics' item sets. A clear checklist of symptom, effect and fix gets you marks quickly.